How do you know if an incrementality test is credible?

6 min read
Published Sep 10, 2026

Context: A marketing analytics professional at a large ecommerce company needed to decide which incrementality tests were reliable enough to use in marketing decision making and calibrating MMM.

The short answer

Check an incrementality test's setup, uncertainty, and consistency with comparable experiments. The test and control groups must support a fair comparison, and the confidence interval must be precise enough for the decision. Review sales before, during, and after the test, investigate unexplained differences, and document why you accept or reject the result.

Here's a summary video from Sellforte CEO, Juha Nuutinen:

 
 

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why marketers ask this

An incrementality report can put a precise-looking number in front of a difficult budget decision. An incremental return on ad spend, or iROAS, of 4 suggests that every dollar of advertising generated four dollars of additional revenue. Before you use that number, you need to know how much confidence to place in it.

The difficulty is that an unexpected result can have several explanations. The campaign may have performed differently, the experiment may have been too small to measure its effect, or the test and control groups may have been a poor match. Each explanation calls for a different response. A credible test can also show that a campaign produced little incremental value.

Are the test and control groups comparable?

Start by asking how each group was selected and why the control represents what would have happened without the intervention. Incrementality testing depends on that comparison. If the groups would have behaved differently anyway, some of the reported lift may come from those differences.

For a randomized audience holdout, ask how people were assigned and whether the holdout actually missed the advertising being tested. For a geo test, inspect the regions and the historical sales patterns used to build the comparison. A synthetic control combines untreated regions to estimate how sales in the test region would have developed without the spending change.

The groups do not need identical population sizes or raw sales totals. They do need a defensible comparison after the method accounts for differences in scale. Check purchasing power, demographics, and customer mix alongside the statistical fit. A group of large cities and a group of smaller towns may respond differently to seasonal demand, even if their sales matched over a short period.

Ask the analyst to show you:

  • How closely the test group and its control tracked before the intervention, including relevant seasonal periods.
  • Whether the planned spending change actually happened, on the documented dates and in the intended campaigns.
  • Whether advertising reached the control group through geographic spillover or another delivery route.
  • Whether a promotion, stock change, holiday, or tracking issue affected one group differently.

A narrow confidence interval cannot repair a biased comparison. If the groups were already drifting apart before the test, the analyst needs to explain that drift before you can attribute the later gap to advertising.

Is the confidence interval narrow enough for the decision?

Read the interval alongside the iROAS estimate. The central estimate gives you the return; the interval shows how precisely the experiment measured it.

Consider an illustrative example. Two tests both report an iROAS of 4. One has a 90% confidence interval from 3.5 to 4.5. The other has a 90% confidence interval from 0 to 8. The first is much more precise. The second leaves you unable to distinguish between no incremental return and a very strong return.

You should reject the second result as a basis for a confident iROAS-based budget decision. Record it as inconclusive and investigate whether the design had enough statistical power to detect the effect you needed to measure. It does not establish that the channel has zero impact, and a wide interval alone does not prove that the experiment was executed incorrectly.

Precision must also be useful for your business. Compare the full interval with the return your company needs after accounting for margins and other relevant costs. Even a statistically positive result may leave the budget decision unresolved if its interval spans that threshold.

Confirm the stated confidence level and how the interval was calculated. A range of 3.5 to 4.5 does not, by itself, establish 90% confidence. Frequentist confidence intervals and Bayesian credible intervals have different interpretations; neither measures the probability that the entire test setup is sound. NIST's explanation of confidence intervals provides the statistical background.

What happens to sales after the test ends?

For a test with daily sales and a synthetic control, inspect the post-treatment period as well as the test window. Look at the following 30 days to see whether sales keep falling behind the control, recover, or jump above it. Agree on the observation window when designing the test, taking the buying cycle and any overlapping activity into account.

Look at both daily sales and the cumulative sales difference. In a spending-cut test, the daily sales gap may close once advertising resumes. The cumulative gap can remain below zero because it still includes sales lost earlier. A flat cumulative line means the gap has stopped growing; it does not have to return to zero for the experiment to be credible.

If the cumulative shortfall keeps growing, the intervention may still be affecting purchases. If it starts to close, some sales may have been delayed and recovered later. A sudden jump deserves investigation: it could reflect recovery, but it could also come from a promotion, a stock change, or a problem with the control.

Eyeballing the charts is a useful diagnostic. It helps you identify patterns that a headline iROAS hides, but the analyst still needs to check their cause. Do not keep extending or shortening the window until the result looks favorable. If a documented event makes the comparison unreliable, record the reason for changing the analysis and show how that change affects the conclusion.

This check depends on the available data. A platform report with only summary statistics may not let you inspect the same daily pattern. In that case, verify its post-test conversion window and keep that visibility limitation in the assessment.

Do comparable tests tell a consistent story?

Start with earlier tests on the same platform in the same market. Then look at the same channel in other geographies and other channels in the same market. The further the comparison moves from the original conditions, the more explanation you need for any difference.

Before comparing the numbers, check that the studies measured the same outcome. Revenue before returns and net revenue are different measures. Web-only purchases and combined web and app purchases cover different sales. Campaign objectives, audiences, spending levels, seasonality, and post-test windows can also change the result.

Repeated tests under similar conditions should produce broadly compatible evidence after accounting for uncertainty. You should not expect identical point estimates. Different channels or markets can have different iROAS values.

A large, unexplained disagreement is a reason to investigate both the latest test and the earlier evidence. Ask what changed, whether the comparison is fair, and whether the original setup can be reconstructed. An earlier result should face the same scrutiny as a new one.

Keep every test in the evidence library, including weak and negative results. If you exclude a study from a decision, state whether the reason is a design flaw, a data problem, insufficient precision, or a mismatch with the question you are answering. An unfavorable result is not an exclusion criterion.

How this looks in practice

Suppose your team is deciding whether a campaign's tested spending level meets its return requirement. The following are hypothetical examples. Assume the same revenue definition and a business iROAS threshold of 3, set by the team before reviewing the tests.

Evidence Assessment Decision
iROAS 4; 90% interval 3.5 to 4.5. Sound setup and no unexplained sales pattern. The full interval exceeds the assumed threshold of 3. Use the result as evidence that the tested spending level meets the requirement in those conditions.
iROAS 4; 90% interval 0 to 8. No identified design flaw. The interval includes outcomes below and above the threshold. Mark the result inconclusive. Review test power before planning another experiment.
iROAS 4; 90% interval 3.5 to 4.5. A promotion affected only the test group. The reported precision does not resolve the alternative explanation for the sales gap. Withhold acceptance until the analyst establishes whether the advertising effect can be isolated.

For each result, save the setup, exact dates, included campaigns, KPI definition, interval, post-test observations, and comparison with earlier studies. Add the decision and its reason. That gives the next reviewer enough information to understand why the evidence was accepted, left inconclusive, or rejected.

How Sellforte helps

Sellforte brings Geo Lift, Conversion Lift, and uploaded A/B test results into one Experiments Hub, with iROAS and uncertainty visible for review. For Geo Lift and A/B analyses, the results dashboard also shows actual performance against the counterfactual and the cumulative effect, helping teams inspect the pattern behind the headline number. See how to read the results dashboard or explore the Sellforte demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.