How to Choose an Incrementality Testing Tool: 32 Evaluation Criteria

16 min read
Published May 29, 2026
Updated

Incrementality testing has moved from a niche capability to a mainstream requirement for serious marketing measurement programs. Yet most organizations struggle with the same problem when evaluating tools: vendor comparison guides are shallow, RFP templates are vague, and it's genuinely hard to know what good looks like across geo tests, A/B tests, conversion lift tests, and the platforms that unify them.

This article presents a research-backed evaluation framework of 32 criteria across 7 categories for assessing incrementality testing tools. The framework was developed by Sellforte from marketer discussions, enterprise RFPs, practitioner interviews, and desk research. It uses the same criteria as the September 2026 vendor comparison to help marketing analytics teams and procurement professionals evaluate tools. Sellforte is also one of the vendors assessed in that comparison.

If you're looking for a vendor comparison applying this framework, see our separate article: Best Incrementality Testing Tools in 2026: In-depth Vendor Comparison.

How to Choose an Incrementality Testing Tool 32 Evaluation Criteria

Table of Contents

  1. What is incrementality testing?
  2. The three primary incrementality test types
  3. How we developed the evaluation criteria
  4. The 32 evaluation criteria across 7 categories
  5. How the criteria are scored
  6. How to use this framework in your own evaluation
  7. Frequently asked questions
  8. Further reading

What is Incrementality Testing?

Incrementality testing is a method for measuring the true causal impact of advertising: the additional sales, conversions, or revenue that would not have happened without a specific marketing activity.

Compared to Marketing Mix Modeling (MMM), which estimates incrementality by analyzing historical time-series data, incrementality testing is an active method. A marketing intervention is designed (such as stopping spend on a channel in one geography), then executed, then analyzed. The objective is a clean, causal read of the true incremental lift from a specific marketing activity.

Incrementality testing is the most accurate approach for estimating the true incremental sales impact of a channel at a specific point in time and spend level. But it has limitations: the estimate applies only to the tested channel, at that specific moment, and at that specific spend level. It does not provide continuous marketing measurement or show how Incremental ROAS changes as spend changes. This is why the primary use case for incrementality testing is to calibrate a Marketing Mix Model, which provides continuous, cross-channel measurement of Incremental ROAS (iROAS) and Marginal Incremental ROAS (miROAS).

The Three Primary Incrementality Test Types

Modern incrementality programs rely on three distinct test designs, each suited to different questions and contexts. Understanding all three is essential to evaluating whether a tool genuinely covers your measurement program or only part of it.

Geo Tests

Geo tests compare treatment and control geographies to derive causal estimates of incremental lift. They are the workhorse of modern incrementality measurement and are well-suited for any channel where spend can be varied by region. A geo test produces an iROAS estimate with a confidence interval for the tested channel, geography, and spend period.

A/B Tests for Owned Media

A/B tests randomize at the user or audience level, making them best suited for owned media such as catalog mailings, email campaigns, or push notifications. They produce the same core output as geo tests (an iROAS estimate with a confidence interval), but at the individual exposure level rather than the geographic level.

Conversion Lift Tests

Conversion Lift tests are run inside ad platforms (Meta, Google, TikTok). The platform randomly withholds ads from a portion of the target audience, then compares conversion rates between the exposed and withheld groups. Conversion Lift tests are widely used by marketing teams but treated very differently by incrementality testing vendors: some dismiss them entirely, others natively ingest and normalize results across platforms.

Each test type produces a point estimate of iROAS with a confidence interval. A mature incrementality testing program typically runs all three, and the best tools support all three in a unified library.

How We Developed the Evaluation Criteria

The 32 criteria in this framework were built from primary research across four sources.

1. 700+ discussions with marketers and analytics professionals

We analyzed incrementality testing-related comments and requirements from more than 700 discussions with Sellforte customers and prospects, including marketers, marketing analytics leads, and data scientists working in advertising-heavy industries such as retail, ecommerce, DTC, travel and hospitality, and restaurants.

2. Enterprise RFP documentation

We reviewed the requirements documentation from more than ten enterprise RFPs explicitly specifying requirements for incrementality testing platforms. Enterprise RFPs tend to be more precise than vendor marketing materials about what actually matters in procurement.

3. Internal practitioner interviews

We interviewed Customer Success and Data Science team members who work with incrementality testing in production across dozens of enterprise advertiser implementations, giving us ground-level insight into what differentiates tools in real use versus on paper.

4. Desk research and LLM-assisted analysis

We complemented primary research with desk research and LLM-assisted investigation to identify gaps and pressure-test assumptions against publicly available product documentation, technical specifications, and analyst coverage.

The framework contains 32 evaluation criteria across 7 categories. The definitions below match the current evaluation criteria in the vendor comparison.

The 32 Evaluation Criteria Across 7 Categories

The 7 categories reflect the full scope of what a mature incrementality testing platform needs to do: from analyzing individual test types to unifying experiments, integrating with MMM, and meeting enterprise requirements.

Category 1: Geo Test Analysis

Geo testing is the workhorse of modern incrementality measurement. This category measures the depth and quality of a tool's geo test analysis capabilities, from the statistical methodology behind the analysis to the user interface that makes it accessible without analyst support.

Category 1: Geo Test Analysis - 6 criteria
IDCriterionWhat it means
1.1Analyzes geo tests with synthetic control method, providing iROAS and confidence intervalUses synthetic control methodology to compare treatment vs. matched control geos, outputting incremental ROAS with statistical confidence intervals to quantify causal lift.
1.2Self-serve UI for analyzing & reviewing geo test resultsMarketers can upload data, run analyses, and review geo test results through a web interface without needing a data scientist or analyst to write code.
1.3Estimates media counterfactual for lost/incremental spendModels what media spend would have been in the absence of the test, so iROAS reflects actual incremental spend rather than nominal budget changes.
1.4Configurable default post-test treatment / measurement windowUser can set default treatment and measurement windows (e.g., test duration, post-test cooldown) that apply across tests, with the option to override per test.
1.5User can launch geo tests from the platformUser can launch geo experiments directly from the tool.
1.6Automatically detects geo tests from media and sales dataIdentifies likely geo experiments from observed media and sales patterns automatically, without requiring users to manually flag test periods or geos.

 

Why it matters: Criteria 1.1 to 1.4 cover geo analysis and its settings. Criterion 1.5 asks whether users can launch geo tests directly from the tool; it does not require a particular ad-platform API implementation. Criterion 1.6 separately assesses whether the tool detects tests automatically from media and sales data.

Category 2: A/B Test Analysis for Owned Media

A/B tests at the user or audience level are common among large ecommerce businesses testing owned media such as catalogs and email. From an analysis perspective they mirror geo tests, with the same expectations around iROAS estimation and confidence intervals, but dedicated A/B test analysis is rare among incrementality testing tools, which makes this category a meaningful differentiator.

Category 2: A/B Test Analysis for Owned Media - 4 criteria
IDCriterionWhat it means
2.1Analyzes own media A/B tests (such as catalog or email tests), providing iROAS and confidence intervalAnalyzes own media A/B tests (such as Leaflet tests), producing incremental ROAS with confidence intervals.
2.2Self-serve UI for analyzing & reviewing own media A/B test (such as catalog or email tests) resultsMarketers can upload data, run analyses, and review A/B test results through a web interface without analyst support.
2.3Estimates media counterfactual for own media A/B (such as catalog or email tests) testsModels counterfactual media spend so iROAS reflects true incremental investment, not just budget delta between cells.
2.4Configurable default post-test treatment / measurement window for own A/B tests (such as catalog or email tests)User can set default treatment and measurement windows that apply across A/B tests, with per-test overrides.

 

Why it matters: Organizations with catalog or email programs need analysis of owned-media experiments. Check that the tool reports iROAS with confidence intervals and supports the relevant spend counterfactual and measurement windows. Synthetic control is specified for geo tests in criterion 1.1; it is not a requirement for owned-media A/B tests in criterion 2.1.

Category 3: Conversion Lift Test Analysis

Conversion Lift tests are platform-native experiments run inside Meta, Google, TikTok, and other ad platforms. They are widely run by marketing teams, often as a standard practice on major paid media channels, but they receive very different treatment from incrementality testing vendors. Some dismiss them citing platform bias; others natively ingest, normalize, and feed them into MMM calibration. This category separates the two approaches.

Category 3: Conversion Lift Test Analysis - 5 criteria
IDCriterionWhat it means
3.1Ingests Conversion Lift test (e.g. Meta Conversion Lift) results, providing iROAS and confidence interval through the platformImports Conversion Lift test results from ad platforms (Meta, Google, etc.) and reports iROAS with confidence intervals on the platform.
3.2Self-serve UI dashboard for analyzing Conversion Lift test (e.g. Meta Conversion Lift) resultsMarketers can review Conversion Lift test results in a web interface without needing to pull raw data from each ad platform.
3.3API connectors that can ingest Conversion Lift test results directly from the platform (instead of manual uploads)Purpose-built API connectors that can automatically pull Conversion Lift results from ad platforms (not generic connectors), eliminating manual exports.
3.4Conversion Lift test results made comparable to ad platform data on campaign and ad set levelNormalizes Conversion Lift outputs so iROAS and lift can be compared at the campaign and ad set level across platforms on a like-for-like basis.
3.5Daily snapshot of Conversion Lift Test progress, including iROAS and confidence intervalProvides daily updated views of in-flight tests, including running iROAS and confidence interval estimates, so users can monitor progress before completion.

 

Why it matters: Ingesting conversion lift results into the same platform helps teams review them alongside other experiments. Criterion 3.3 requires connectors specifically able to retrieve conversion lift results; a generic advertising-data connector is insufficient evidence. Criterion 3.5 assesses daily progress snapshots that include both iROAS and its confidence interval.

Category 4: Experiment Recommendations & Insights

Beyond analyzing experiments you've already run, the best incrementality testing tools actively help you get more value from your measurement program by recommending what to test next, designing statistically rigorous experiments, and translating technical outputs into narratives that non-technical stakeholders can act on.

Category 4: Experiment Recommendations & Insights - 6 criteria
IDCriterionWhat it means
4.1Platform recommends what to test nextPlatform recommends what to test next.
4.2Platform recommends control & test groupsRecommends which geos, audiences, or users to assign to control vs. test based on similarity, balance, and statistical power.
4.3Platform recommends test design and predicts test successPlatform recommends test design and predicts test success.
4.4AI-generated plain-language summariesGenerates plain-language summaries of test results.
4.5Conversational AI for discussing experimentsBuilt-in AI assistant (not an MCP) lets users ask natural-language questions about experiments, results, and learnings across the library.
4.6MCP that can access experiment resultsExperiment results can be interacted with through an MCP that can be connected to Claude, ChatGPT or Gemini.

 

Why it matters: Criteria 4.1 to 4.3 assess recommendations about what to test, group selection, and test design with success prediction. Criterion 4.4 covers plain-language summaries. Criterion 4.5 assesses a built-in conversational AI assistant, while criterion 4.6 separately assesses access to experiment results through Model Context Protocol (MCP), which can connect to tools such as Claude, ChatGPT, or Gemini. Evidence of MCP access alone does not establish built-in conversational AI.

Category 5: Unified Experiment Library

For organizations with mature incrementality programs, experiment management becomes a significant challenge. Large advertisers may run dozens or hundreds of experiments annually across channels, geographies, teams, and test types. A unified library that stores all of them in one searchable place prevents duplicate testing, compounds organizational learning, and enables governance at scale.

Category 5: Unified Experiment Library - 3 criteria
IDCriterionWhat it means
5.1Central library that covers geo experiments, A/B tests for own media and Conversion Lift testsSingle repository stores results from all experiments (geo, A/B, conversion lift) across channels, teams, and methodologies in one searchable place.
5.2Library is filterable, for example by channel and platformLibrary supports filtering of experiments by attributes, for example channel and platform.
5.3Role-based access & governance for the experiment librarySupports role-based access control and governance features so different users see appropriate experiments and have appropriate edit rights.

 

Why it matters: A shared library helps teams find earlier experiments and reuse their findings. Criterion 5.1 requires coverage of geo experiments, owned-media A/B tests, and conversion lift tests. Criterion 5.2 asks whether users can filter the library, with channel and platform as examples; it does not require every country, brand, campaign, team, and date filter. Criterion 5.3 assesses permissions and governance.

Category 6: MMM Integration

Incrementality testing and Marketing Mix Modeling are complements, not substitutes. Experiments provide ground-truth point estimates for specific channels and time windows. MMM provides continuous, cross-channel measurement of incremental ROAS over time. The integration between the two is what makes each more valuable. This category assesses how deeply a tool supports that integration.

Category 6: MMM Integration - 3 criteria
IDCriterionWhat it means
6.1Has MMM that includes prior-based calibration by the userIncludes a Marketing Mix Model that users can calibrate with priors.
6.2UI to connect experiment results to MMMUser interface for connecting experiment results into the MMM as calibration inputs, without requiring custom code.
6.3Experiment-based priors comparable to attribution-based priorsAllows side-by-side comparison of priors derived from experiments vs. priors derived from attribution data.

 

Why it matters: Criterion 6.1 asks whether the tool includes MMM that users can calibrate with priors. Criterion 6.2 assesses the interface for connecting experiment results to that model. Criterion 6.3 adds a separate requirement: comparing experiment-based and attribution-based priors side by side.

Category 7: Enterprise-Grade Platform

The final category assesses whether the platform can operate in an enterprise environment. These criteria appear repeatedly in RFPs from large advertisers and often act as hard filters in procurement processes.

Category 7: Enterprise-Grade Platform - 5 criteria
IDCriterionWhat it means
7.1At least 10 public reference customers from $1B+ revenue brandsHas at least 10 publicly named reference customers among brands with $1B+ in annual revenue, demonstrating enterprise-scale adoption.
7.2SOC 2, ISO 27001, or audited IT security by a third-party cyber security auditorHolds SOC 2, ISO 27001, or equivalent third-party-audited security certification demonstrating mature security controls.
7.3Data residency: US and EU optionsCustomer can choose whether their data is stored and processed in US or EU regions to meet data residency and regulatory requirements.
7.4Dual-cloud option between AWS, GCP, and AzureCustomer can choose at least two cloud providers from AWS, GCP, and Azure to align with their IT requirements.
7.5Single sign-on (SSO) for enterprisesSupports enterprise authentication with single sign-on (SSO).

 

Why it matters: Check enterprise requirements against your organization’s IT policies. The framework assesses publicly verifiable enterprise references, independent security assurance, US and EU data-residency options, a choice of at least two cloud providers from AWS, GCP, and Azure, and SSO. Criterion 7.4 does not require support for all three cloud providers.

How the Criteria Are Scored

Use public information from each vendor’s website and its technical documentation domain. Prioritize technical documentation, then product pages, then marketing collateral such as product launch posts. Record a score, a two-sentence rationale, the source URL, and the source type for each criterion.

Evidence-based scoring scale
ScoreDefinition
1Strong evidence that the platform supports the capability.
0.5Partial evidence that the platform supports the capability.
0No evidence found that the platform supports the capability.

A zero records a gap in the evidence found. It does not establish that the capability is absent. Use the score and rationale to identify what to verify with the vendor.

Each criterion has equal weight, for a maximum of 32 points. The seven category maximums are 6, 4, 5, 6, 3, 3, and 5 points, respectively. Categories with more criteria contribute more points to the total.

In the September 2026 comparison, Claude and ChatGPT separately evaluated eight tools using public information. Each criterion’s reported score is the average of the two assessments. Category and total scores sum the unrounded averages; displayed scores are rounded to one decimal. For example, scores of 1 and 0.5 average to 0.75, displayed as 0.8, while 0.75 is used in totals. This research assesses public evidence and was not a hands-on product test. See the full scoring methodology and research limitations.

How to Use This Framework in Your Own Evaluation

The published comparison weights every criterion equally. For your own purchasing decision, identify the requirements that matter most to your organization and document any different weighting separately.

If you run geo tests only, start with Category 1. But plan for Categories 2 and 3 as well, as most mature measurement programs add conversion lift tests and owned-media A/B tests over time. Choosing a platform that only covers geo tests now may force a migration later.

If you run conversion lift tests (Meta, Google, TikTok), review Category 3 closely. Ask vendors to demonstrate how they ingest and display results, and explain how they handle methodological limitations. A low public-evidence score is a reason to request documentation and verify the workflow.

If you have a mature, multi-team incrementality program, Category 5 (Unified Experiment Library) becomes critical. At scale, an experiment library without role-based access, cross-type coverage, and robust filtering creates organizational risk: teams duplicate tests, learnings fragment, and governance breaks down.

If your incrementality program is primarily designed to calibrate an MMM, Category 6 is the integration you should scrutinize most carefully. A UI-based workflow for connecting experiment results to model priors (criterion 6.2) versus a manual spreadsheet process is a substantial difference in operational overhead at scale.

If you are in enterprise procurement, check Category 7 against your IT requirements early. Confirm which security assurances, data-residency regions, cloud-provider options, and authentication methods your organization needs.

For any evaluation: ask shortlisted vendors to demonstrate the relevant workflows using realistic scenarios from your measurement program. Use public documentation to prepare your questions and record what the demonstration establishes separately from the public-evidence score.

To see how eight incrementality testing tools score against all 32 criteria, see the full vendor comparison: Best Incrementality Testing Tools in 2026: In-depth Vendor Comparison.

Frequently Asked Questions

How do I choose an incrementality testing tool?

Evaluate candidates against seven dimensions: geo test analysis, A/B test analysis for owned media, conversion lift test analysis, experiment recommendations and insights, unified experiment library, MMM integration, and enterprise-grade platform requirements. The 32 specific criteria in this article define what good looks like in each dimension. Prioritize based on which test types you run today and plan to run in the future. A platform that covers only geo tests may require a migration if your program expands to conversion lift tests or owned-media A/B tests.

What is the difference between geo tests, A/B tests, and conversion lift tests?

Geo tests compare treatment and control geographies to estimate incremental lift: they work for any paid channel where spend can vary by region. A/B tests randomize at the user or audience level, making them best suited for owned media like catalogs and email. Conversion Lift tests are run inside ad platforms (Meta, Google, TikTok) and compare conversion rates between users who saw an ad and those who didn't. All three produce an iROAS estimate with a confidence interval, but they differ in methodology, use case, and the vendor support they receive.

What is the difference between built-in conversational AI and MCP access?

Criterion 4.5 covers an assistant inside the platform that can discuss experiments and their results. Criterion 4.6 covers access to experiment results through an MCP connection to an external AI tool. Score these separately: evidence for one does not establish the other.

Does a score of zero mean a capability is missing?

No. It means the evaluation found no public evidence that the tool supports the capability. Ask the vendor for documentation and a demonstration before treating it as a product limitation.

Why does the unified experiment library matter?

Large organizations run dozens or hundreds of experiments per year across channels, regions, and teams. Without a unified library, results scatter across spreadsheets and platform dashboards. Teams re-run experiments already completed elsewhere, learnings don't compound over time, and governance becomes impossible. A unified library that covers all test types (geo, A/B, and conversion lift) in one searchable, filterable, role-controlled repository is the infrastructure that makes an incrementality program scale.

How do incrementality testing tools integrate with Marketing Mix Modeling?

Experiments provide point estimates of iROAS for specific channels, time windows, and spend levels. These estimates serve as calibration inputs (called priors) that inform a Bayesian MMM's estimates of incremental ROAS across all channels continuously. The best tools provide a UI-based workflow for connecting experiment results directly to model priors, with the ability to compare experiment-derived priors against attribution-derived priors side by side. Tools without this integration require manual, code-based calibration processes that are slow and error-prone.

Should I include conversion lift tests in my evaluation, even if I'm skeptical of their accuracy?

Yes. Even if you discount conversion lift test results or treat them as directional rather than definitive, having them in a unified experiment library alongside geo tests and A/B tests gives you a more complete picture of incrementality across your media mix. The appropriate response to platform bias concerns is methodological awareness and careful interpretation, not excluding a widely run test type from your measurement framework entirely. The best platforms let you decide how much weight to put on conversion lift results when feeding them into MMM calibration.

Further Reading

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte, with over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world’s largest advertisers on data-driven marketing optimization. Follow Lauri in LinkedIn, where he is one of the leading voices in MMM and marketing measurement.

Kacper Solarski

Kacper Solarski is a Lead Data Scientist at Sellforte, focused on developing Sellforte's Experiments product. Kacper is one of the most senior data scientists and developers at Sellforte, where he has implemented Marketing Mix Models and incrementality testing solutions to Sellforte customers, while at the same time developing Sellforte's platform. Follow Kacper in LinkedIn.

Juha Nuutinen

Juha Nuutinen is the Chief Executive Officer and co-founder at Sellforte, with over 15 years of experience in optimizing marketing spend and promotional activity for the largest advertisers in the world. Before co-founding Sellforte, he worked as a management consultant at the Boston Consulting Group, specializing in promotion optimization. Follow Juha in LinkedIn, where he is actively sharing his views on marketing measurement.