How to Choose an Incrementality Testing Tool: 32 Evaluation Criteria
Incrementality testing has moved from a niche capability to a mainstream requirement for serious marketing measurement programs. Yet most organizations struggle with the same problem when evaluating tools: vendor comparison guides are shallow, RFP templates are vague, and it's genuinely hard to know what good looks like across geo tests, A/B tests, conversion lift tests, and the platforms that unify them.
This article presents a research-backed evaluation framework of 32 criteria across 7 categories for assessing incrementality testing tools. The framework was developed by Sellforte from marketer discussions, enterprise RFPs, practitioner interviews, and desk research. It uses the same criteria as the September 2026 vendor comparison to help marketing analytics teams and procurement professionals evaluate tools. Sellforte is also one of the vendors assessed in that comparison.
If you're looking for a vendor comparison applying this framework, see our separate article: Best Incrementality Testing Tools in 2026: In-depth Vendor Comparison.

Table of Contents
- What is incrementality testing?
- The three primary incrementality test types
- How we developed the evaluation criteria
- The 32 evaluation criteria across 7 categories
- How the criteria are scored
- How to use this framework in your own evaluation
- Frequently asked questions
- Further reading
What is Incrementality Testing?
Incrementality testing is a method for measuring the true causal impact of advertising: the additional sales, conversions, or revenue that would not have happened without a specific marketing activity.
Compared to Marketing Mix Modeling (MMM), which estimates incrementality by analyzing historical time-series data, incrementality testing is an active method. A marketing intervention is designed (such as stopping spend on a channel in one geography), then executed, then analyzed. The objective is a clean, causal read of the true incremental lift from a specific marketing activity.
Incrementality testing is the most accurate approach for estimating the true incremental sales impact of a channel at a specific point in time and spend level. But it has limitations: the estimate applies only to the tested channel, at that specific moment, and at that specific spend level. It does not provide continuous marketing measurement or show how Incremental ROAS changes as spend changes. This is why the primary use case for incrementality testing is to calibrate a Marketing Mix Model, which provides continuous, cross-channel measurement of Incremental ROAS (iROAS) and Marginal Incremental ROAS (miROAS).
The Three Primary Incrementality Test Types
Modern incrementality programs rely on three distinct test designs, each suited to different questions and contexts. Understanding all three is essential to evaluating whether a tool genuinely covers your measurement program or only part of it.
Geo Tests
Geo tests compare treatment and control geographies to derive causal estimates of incremental lift. They are the workhorse of modern incrementality measurement and are well-suited for any channel where spend can be varied by region. A geo test produces an iROAS estimate with a confidence interval for the tested channel, geography, and spend period.
A/B Tests for Owned Media
A/B tests randomize at the user or audience level, making them best suited for owned media such as catalog mailings, email campaigns, or push notifications. They produce the same core output as geo tests (an iROAS estimate with a confidence interval), but at the individual exposure level rather than the geographic level.
Conversion Lift Tests
Conversion Lift tests are run inside ad platforms (Meta, Google, TikTok). The platform randomly withholds ads from a portion of the target audience, then compares conversion rates between the exposed and withheld groups. Conversion Lift tests are widely used by marketing teams but treated very differently by incrementality testing vendors: some dismiss them entirely, others natively ingest and normalize results across platforms.
Each test type produces a point estimate of iROAS with a confidence interval. A mature incrementality testing program typically runs all three, and the best tools support all three in a unified library.
How We Developed the Evaluation Criteria
The 32 criteria in this framework were built from primary research across four sources.
1. 700+ discussions with marketers and analytics professionals
We analyzed incrementality testing-related comments and requirements from more than 700 discussions with Sellforte customers and prospects, including marketers, marketing analytics leads, and data scientists working in advertising-heavy industries such as retail, ecommerce, DTC, travel and hospitality, and restaurants.
2. Enterprise RFP documentation
We reviewed the requirements documentation from more than ten enterprise RFPs explicitly specifying requirements for incrementality testing platforms. Enterprise RFPs tend to be more precise than vendor marketing materials about what actually matters in procurement.
3. Internal practitioner interviews
We interviewed Customer Success and Data Science team members who work with incrementality testing in production across dozens of enterprise advertiser implementations, giving us ground-level insight into what differentiates tools in real use versus on paper.
4. Desk research and LLM-assisted analysis
We complemented primary research with desk research and LLM-assisted investigation to identify gaps and pressure-test assumptions against publicly available product documentation, technical specifications, and analyst coverage.
The framework contains 32 evaluation criteria across 7 categories. The definitions below match the current evaluation criteria in the vendor comparison.
The 32 Evaluation Criteria Across 7 Categories
The 7 categories reflect the full scope of what a mature incrementality testing platform needs to do: from analyzing individual test types to unifying experiments, integrating with MMM, and meeting enterprise requirements.
Category 1: Geo Test Analysis
Geo testing is the workhorse of modern incrementality measurement. This category measures the depth and quality of a tool's geo test analysis capabilities, from the statistical methodology behind the analysis to the user interface that makes it accessible without analyst support.
| ID | Criterion | What it means |
|---|---|---|
| 1.1 | Analyzes geo tests with synthetic control method, providing iROAS and confidence interval | Uses synthetic control methodology to compare treatment vs. matched control geos, outputting incremental ROAS with statistical confidence intervals to quantify causal lift. |
| 1.2 | Self-serve UI for analyzing & reviewing geo test results | Marketers can upload data, run analyses, and review geo test results through a web interface without needing a data scientist or analyst to write code. |
| 1.3 | Estimates media counterfactual for lost/incremental spend | Models what media spend would have been in the absence of the test, so iROAS reflects actual incremental spend rather than nominal budget changes. |
| 1.4 | Configurable default post-test treatment / measurement window | User can set default treatment and measurement windows (e.g., test duration, post-test cooldown) that apply across tests, with the option to override per test. |
| 1.5 | User can launch geo tests from the platform | User can launch geo experiments directly from the tool. |
| 1.6 | Automatically detects geo tests from media and sales data | Identifies likely geo experiments from observed media and sales patterns automatically, without requiring users to manually flag test periods or geos. |
Why it matters: Criteria 1.1 to 1.4 cover geo analysis and its settings. Criterion 1.5 asks whether users can launch geo tests directly from the tool; it does not require a particular ad-platform API implementation. Criterion 1.6 separately assesses whether the tool detects tests automatically from media and sales data.
Category 2: A/B Test Analysis for Owned Media
A/B tests at the user or audience level are common among large ecommerce businesses testing owned media such as catalogs and email. From an analysis perspective they mirror geo tests, with the same expectations around iROAS estimation and confidence intervals, but dedicated A/B test analysis is rare among incrementality testing tools, which makes this category a meaningful differentiator.
| ID | Criterion | What it means |
|---|---|---|
| 2.1 | Analyzes own media A/B tests (such as catalog or email tests), providing iROAS and confidence interval | Analyzes own media A/B tests (such as Leaflet tests), producing incremental ROAS with confidence intervals. |
| 2.2 | Self-serve UI for analyzing & reviewing own media A/B test (such as catalog or email tests) results | Marketers can upload data, run analyses, and review A/B test results through a web interface without analyst support. |
| 2.3 | Estimates media counterfactual for own media A/B (such as catalog or email tests) tests | Models counterfactual media spend so iROAS reflects true incremental investment, not just budget delta between cells. |
| 2.4 | Configurable default post-test treatment / measurement window for own A/B tests (such as catalog or email tests) | User can set default treatment and measurement windows that apply across A/B tests, with per-test overrides. |
Why it matters: Organizations with catalog or email programs need analysis of owned-media experiments. Check that the tool reports iROAS with confidence intervals and supports the relevant spend counterfactual and measurement windows. Synthetic control is specified for geo tests in criterion 1.1; it is not a requirement for owned-media A/B tests in criterion 2.1.
Category 3: Conversion Lift Test Analysis
Conversion Lift tests are platform-native experiments run inside Meta, Google, TikTok, and other ad platforms. They are widely run by marketing teams, often as a standard practice on major paid media channels, but they receive very different treatment from incrementality testing vendors. Some dismiss them citing platform bias; others natively ingest, normalize, and feed them into MMM calibration. This category separates the two approaches.
| ID | Criterion | What it means |
|---|---|---|
| 3.1 | Ingests Conversion Lift test (e.g. Meta Conversion Lift) results, providing iROAS and confidence interval through the platform | Imports Conversion Lift test results from ad platforms (Meta, Google, etc.) and reports iROAS with confidence intervals on the platform. |
| 3.2 | Self-serve UI dashboard for analyzing Conversion Lift test (e.g. Meta Conversion Lift) results | Marketers can review Conversion Lift test results in a web interface without needing to pull raw data from each ad platform. |
| 3.3 | API connectors that can ingest Conversion Lift test results directly from the platform (instead of manual uploads) | Purpose-built API connectors that can automatically pull Conversion Lift results from ad platforms (not generic connectors), eliminating manual exports. |
| 3.4 | Conversion Lift test results made comparable to ad platform data on campaign and ad set level | Normalizes Conversion Lift outputs so iROAS and lift can be compared at the campaign and ad set level across platforms on a like-for-like basis. |
| 3.5 | Daily snapshot of Conversion Lift Test progress, including iROAS and confidence interval | Provides daily updated views of in-flight tests, including running iROAS and confidence interval estimates, so users can monitor progress before completion. |
Why it matters: Ingesting conversion lift results into the same platform helps teams review them alongside other experiments. Criterion 3.3 requires connectors specifically able to retrieve conversion lift results; a generic advertising-data connector is insufficient evidence. Criterion 3.5 assesses daily progress snapshots that include both iROAS and its confidence interval.
Category 4: Experiment Recommendations & Insights
Beyond analyzing experiments you've already run, the best incrementality testing tools actively help you get more value from your measurement program by recommending what to test next, designing statistically rigorous experiments, and translating technical outputs into narratives that non-technical stakeholders can act on.
| ID | Criterion | What it means |
|---|---|---|
| 4.1 | Platform recommends what to test next | Platform recommends what to test next. |
| 4.2 | Platform recommends control & test groups | Recommends which geos, audiences, or users to assign to control vs. test based on similarity, balance, and statistical power. |
| 4.3 | Platform recommends test design and predicts test success | Platform recommends test design and predicts test success. |
| 4.4 | AI-generated plain-language summaries | Generates plain-language summaries of test results. |
| 4.5 | Conversational AI for discussing experiments | Built-in AI assistant (not an MCP) lets users ask natural-language questions about experiments, results, and learnings across the library. |
| 4.6 | MCP that can access experiment results | Experiment results can be interacted with through an MCP that can be connected to Claude, ChatGPT or Gemini. |
Why it matters: Criteria 4.1 to 4.3 assess recommendations about what to test, group selection, and test design with success prediction. Criterion 4.4 covers plain-language summaries. Criterion 4.5 assesses a built-in conversational AI assistant, while criterion 4.6 separately assesses access to experiment results through Model Context Protocol (MCP), which can connect to tools such as Claude, ChatGPT, or Gemini. Evidence of MCP access alone does not establish built-in conversational AI.
Category 5: Unified Experiment Library
For organizations with mature incrementality programs, experiment management becomes a significant challenge. Large advertisers may run dozens or hundreds of experiments annually across channels, geographies, teams, and test types. A unified library that stores all of them in one searchable place prevents duplicate testing, compounds organizational learning, and enables governance at scale.
| ID | Criterion | What it means |
|---|---|---|
| 5.1 | Central library that covers geo experiments, A/B tests for own media and Conversion Lift tests | Single repository stores results from all experiments (geo, A/B, conversion lift) across channels, teams, and methodologies in one searchable place. |
| 5.2 | Library is filterable, for example by channel and platform | Library supports filtering of experiments by attributes, for example channel and platform. |
| 5.3 | Role-based access & governance for the experiment library | Supports role-based access control and governance features so different users see appropriate experiments and have appropriate edit rights. |
Why it matters: A shared library helps teams find earlier experiments and reuse their findings. Criterion 5.1 requires coverage of geo experiments, owned-media A/B tests, and conversion lift tests. Criterion 5.2 asks whether users can filter the library, with channel and platform as examples; it does not require every country, brand, campaign, team, and date filter. Criterion 5.3 assesses permissions and governance.
Category 6: MMM Integration
Incrementality testing and Marketing Mix Modeling are complements, not substitutes. Experiments provide ground-truth point estimates for specific channels and time windows. MMM provides continuous, cross-channel measurement of incremental ROAS over time. The integration between the two is what makes each more valuable. This category assesses how deeply a tool supports that integration.
| ID | Criterion | What it means |
|---|---|---|
| 6.1 | Has MMM that includes prior-based calibration by the user | Includes a Marketing Mix Model that users can calibrate with priors. |
| 6.2 | UI to connect experiment results to MMM | User interface for connecting experiment results into the MMM as calibration inputs, without requiring custom code. |
| 6.3 | Experiment-based priors comparable to attribution-based priors | Allows side-by-side comparison of priors derived from experiments vs. priors derived from attribution data. |
Why it matters: Criterion 6.1 asks whether the tool includes MMM that users can calibrate with priors. Criterion 6.2 assesses the interface for connecting experiment results to that model. Criterion 6.3 adds a separate requirement: comparing experiment-based and attribution-based priors side by side.
Category 7: Enterprise-Grade Platform
The final category assesses whether the platform can operate in an enterprise environment. These criteria appear repeatedly in RFPs from large advertisers and often act as hard filters in procurement processes.
| ID | Criterion | What it means |
|---|---|---|
| 7.1 | At least 10 public reference customers from $1B+ revenue brands | Has at least 10 publicly named reference customers among brands with $1B+ in annual revenue, demonstrating enterprise-scale adoption. |
| 7.2 | SOC 2, ISO 27001, or audited IT security by a third-party cyber security auditor | Holds SOC 2, ISO 27001, or equivalent third-party-audited security certification demonstrating mature security controls. |
| 7.3 | Data residency: US and EU options | Customer can choose whether their data is stored and processed in US or EU regions to meet data residency and regulatory requirements. |
| 7.4 | Dual-cloud option between AWS, GCP, and Azure | Customer can choose at least two cloud providers from AWS, GCP, and Azure to align with their IT requirements. |
| 7.5 | Single sign-on (SSO) for enterprises | Supports enterprise authentication with single sign-on (SSO). |
Why it matters: Check enterprise requirements against your organization’s IT policies. The framework assesses publicly verifiable enterprise references, independent security assurance, US and EU data-residency options, a choice of at least two cloud providers from AWS, GCP, and Azure, and SSO. Criterion 7.4 does not require support for all three cloud providers.
How the Criteria Are Scored
Use public information from each vendor’s website and its technical documentation domain. Prioritize technical documentation, then product pages, then marketing collateral such as product launch posts. Record a score, a two-sentence rationale, the source URL, and the source type for each criterion.
| Score | Definition |
|---|---|
| 1 | Strong evidence that the platform supports the capability. |
| 0.5 | Partial evidence that the platform supports the capability. |
| 0 | No evidence found that the platform supports the capability. |
A zero records a gap in the evidence found. It does not establish that the capability is absent. Use the score and rationale to identify what to verify with the vendor.
Each criterion has equal weight, for a maximum of 32 points. The seven category maximums are 6, 4, 5, 6, 3, 3, and 5 points, respectively. Categories with more criteria contribute more points to the total.
In the September 2026 comparison, Claude and ChatGPT separately evaluated eight tools using public information. Each criterion’s reported score is the average of the two assessments. Category and total scores sum the unrounded averages; displayed scores are rounded to one decimal. For example, scores of 1 and 0.5 average to 0.75, displayed as 0.8, while 0.75 is used in totals. This research assesses public evidence and was not a hands-on product test. See the full scoring methodology and research limitations.
How to Use This Framework in Your Own Evaluation
The published comparison weights every criterion equally. For your own purchasing decision, identify the requirements that matter most to your organization and document any different weighting separately.
If you run geo tests only, start with Category 1. But plan for Categories 2 and 3 as well, as most mature measurement programs add conversion lift tests and owned-media A/B tests over time. Choosing a platform that only covers geo tests now may force a migration later.
If you run conversion lift tests (Meta, Google, TikTok), review Category 3 closely. Ask vendors to demonstrate how they ingest and display results, and explain how they handle methodological limitations. A low public-evidence score is a reason to request documentation and verify the workflow.
If you have a mature, multi-team incrementality program, Category 5 (Unified Experiment Library) becomes critical. At scale, an experiment library without role-based access, cross-type coverage, and robust filtering creates organizational risk: teams duplicate tests, learnings fragment, and governance breaks down.
If your incrementality program is primarily designed to calibrate an MMM, Category 6 is the integration you should scrutinize most carefully. A UI-based workflow for connecting experiment results to model priors (criterion 6.2) versus a manual spreadsheet process is a substantial difference in operational overhead at scale.
If you are in enterprise procurement, check Category 7 against your IT requirements early. Confirm which security assurances, data-residency regions, cloud-provider options, and authentication methods your organization needs.
For any evaluation: ask shortlisted vendors to demonstrate the relevant workflows using realistic scenarios from your measurement program. Use public documentation to prepare your questions and record what the demonstration establishes separately from the public-evidence score.
To see how eight incrementality testing tools score against all 32 criteria, see the full vendor comparison: Best Incrementality Testing Tools in 2026: In-depth Vendor Comparison.
Frequently Asked Questions
How do I choose an incrementality testing tool?
Evaluate candidates against seven dimensions: geo test analysis, A/B test analysis for owned media, conversion lift test analysis, experiment recommendations and insights, unified experiment library, MMM integration, and enterprise-grade platform requirements. The 32 specific criteria in this article define what good looks like in each dimension. Prioritize based on which test types you run today and plan to run in the future. A platform that covers only geo tests may require a migration if your program expands to conversion lift tests or owned-media A/B tests.
What is the difference between geo tests, A/B tests, and conversion lift tests?
Geo tests compare treatment and control geographies to estimate incremental lift: they work for any paid channel where spend can vary by region. A/B tests randomize at the user or audience level, making them best suited for owned media like catalogs and email. Conversion Lift tests are run inside ad platforms (Meta, Google, TikTok) and compare conversion rates between users who saw an ad and those who didn't. All three produce an iROAS estimate with a confidence interval, but they differ in methodology, use case, and the vendor support they receive.
What is the difference between built-in conversational AI and MCP access?
Criterion 4.5 covers an assistant inside the platform that can discuss experiments and their results. Criterion 4.6 covers access to experiment results through an MCP connection to an external AI tool. Score these separately: evidence for one does not establish the other.
Does a score of zero mean a capability is missing?
No. It means the evaluation found no public evidence that the tool supports the capability. Ask the vendor for documentation and a demonstration before treating it as a product limitation.
Why does the unified experiment library matter?
Large organizations run dozens or hundreds of experiments per year across channels, regions, and teams. Without a unified library, results scatter across spreadsheets and platform dashboards. Teams re-run experiments already completed elsewhere, learnings don't compound over time, and governance becomes impossible. A unified library that covers all test types (geo, A/B, and conversion lift) in one searchable, filterable, role-controlled repository is the infrastructure that makes an incrementality program scale.
How do incrementality testing tools integrate with Marketing Mix Modeling?
Experiments provide point estimates of iROAS for specific channels, time windows, and spend levels. These estimates serve as calibration inputs (called priors) that inform a Bayesian MMM's estimates of incremental ROAS across all channels continuously. The best tools provide a UI-based workflow for connecting experiment results directly to model priors, with the ability to compare experiment-derived priors against attribution-derived priors side by side. Tools without this integration require manual, code-based calibration processes that are slow and error-prone.
Should I include conversion lift tests in my evaluation, even if I'm skeptical of their accuracy?
Yes. Even if you discount conversion lift test results or treat them as directional rather than definitive, having them in a unified experiment library alongside geo tests and A/B tests gives you a more complete picture of incrementality across your media mix. The appropriate response to platform bias concerns is methodological awareness and careful interpretation, not excluding a widely run test type from your measurement framework entirely. The best platforms let you decide how much weight to put on conversion lift results when feeding them into MMM calibration.
Further Reading
- Best Incrementality Testing Tools in 2026: In-depth Vendor Comparison
- What is Incrementality Testing? Guide for Marketers
- What is Marketing Mix Modeling?
- Calibrating Marketing Mix Models with Experiments and Attribution Data
- Marginal Incremental ROAS (miROAS) explained
- ROAS, iROAS, miROAS: Choosing the Right KPI for Optimizing Media Spend
- How to Integrate Experiments Into an MMM Platform: A Practical Guide
- 7 Best AI Tools for MMM and Incrementality Testing in 2026
Authors

Lauri Potka is the Chief Operating Officer at Sellforte, with over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world’s largest advertisers on data-driven marketing optimization. Follow Lauri in LinkedIn, where he is one of the leading voices in MMM and marketing measurement.

Kacper Solarski is a Lead Data Scientist at Sellforte, focused on developing Sellforte's Experiments product. Kacper is one of the most senior data scientists and developers at Sellforte, where he has implemented Marketing Mix Models and incrementality testing solutions to Sellforte customers, while at the same time developing Sellforte's platform. Follow Kacper in LinkedIn.
.png?width=701&height=132&name=Juha%20Nuutinen%20(701%20x%20132%20px).png)
Juha Nuutinen is the Chief Executive Officer and co-founder at Sellforte, with over 15 years of experience in optimizing marketing spend and promotional activity for the largest advertisers in the world. Before co-founding Sellforte, he worked as a management consultant at the Boston Consulting Group, specializing in promotion optimization. Follow Juha in LinkedIn, where he is actively sharing his views on marketing measurement.
You May Also Like
These Related Stories

Can you trust Meta Conversion Lift tests for MMM calibration?
What is the best incrementality test?

