Why can two analyses of the same holdout test produce different lift?

4 min read
Published Sep 22, 2026
Updated

The short answer

Two analyses of the same holdout test can produce different lift because they use different sales definitions, time windows, or estimates of what the treatment group would have bought without the intervention. Reconcile the underlying data first, then compare group scaling, pretreatment history, estimation methods, and uncertainty before choosing a result.

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why marketers ask this

Your analyst subtracts holdout sales from treatment-group sales. Another team analyzes the same experiment and reports a different incremental-sales total. The difference could change whether you continue the campaign, so you need to trace it back to the data before using either number to set a budget.

What should we reconcile before comparing the methods?

Ask both teams to reproduce the same daily sales totals for the same customers or regions. A shared experiment name, spreadsheet, or source table does not guarantee that the analyses include the same transactions.

In incrementality testing, lift is the difference between an observed outcome and an estimate of what would have happened without the intervention. Differences in either part change the answer. Use this checklist in order:

Checks to complete before debating which lift estimate is better
Check Ask both teams to show
Sales definition and extract The same data snapshot and query logic: orders or revenue, gross or net of returns, included sales channels, currency, and handling of duplicates or missing records.
Customers or regions The original assignment, eligibility rules, exclusions, and group sizes. Include assigned customers who made no purchase. Check whether either analysis keeps only people who actually received or engaged with the advertising.
Dates The same treatment start and end, sales cutoff, time zone, and post-treatment observation window. Separate the date a sale happened from the date it reached the report.
Units and scaling Whether the result is per customer, the total for the treatment group, or an extrapolation to a larger population. Show any scaling for unequal group sizes.
Calculation and reporting The counterfactual calculation, pretreatment dates, percentage-lift denominator, and uncertainty method. If reporting incremental ROAS, also reconcile the incremental spending denominator.

Start with a few individual days, then reconcile the total. If observed sales already differ, resolve that discrepancy before comparing statistical models. Keep the original group assignment visible: selecting only people who engaged with an ad can undo the comparability created by random assignment.

How do we decide which result to use?

Use the prespecified primary analysis if it remains valid after the data and execution checks. If you discover an error or an assumption that failed, document the correction and retain the original result alongside the revised one.

Ask for a reconciliation that changes one choice at a time: first the extract, then the scope or dates, then the estimation method. Record the lift after each change and why the change is justified. This shows which choices explain the gap, although the size of each step can depend on the order of the changes.

Compare uncertainty on the same basis. A 90% interval and a 95% interval use different confidence levels, and a geo analysis must respect regional assignment and dependence over time. One result being statistically significant while another is not does not establish that the estimates disagree. Because both use the same experiment, their errors are related; ask the analyst to assess the difference directly if it matters to the decision.

If plausible, defensible analyses lead to different budget decisions, report the result as sensitive to the analysis choices. Keep both estimates visible and identify the evidence needed to resolve the uncertainty. Averaging them would conceal the disagreement.

How this looks in practice

Consider a hypothetical customer holdout with 10,000 assigned customers in each group. Both teams use the same net-revenue definition and observation window, including customers who bought nothing. Treatment-group sales are €110,000 and holdout sales are €100,000. These figures are illustrative.

Team A uses the unadjusted difference. Team B assumes, based on pretreatment data, that the treatment group would have generated 2% more revenue than the holdout even without the campaign. For clarity, its adjustment below is a simple multiplier; this is not a recommendation to use that estimator.

The same observed sales, with two different baselines
Calculation Team A Team B
Observed treatment sales €110,000 €110,000
Estimated sales without campaign €100,000 €100,000 × 1.02 = €102,000
Incremental sales €10,000 €8,000
Lift relative to estimated sales without campaign 10.0% 7.8%

The €2,000 gap comes entirely from the baseline. To choose between the estimates, examine whether the 2% adjustment is justified by the design and pretreatment evidence, and compare uncertainty. The table explains the arithmetic, but choosing the more accurate estimate requires those checks. A small percentage adjustment to total sales can be large relative to the incremental effect.

Should separate customer-segment estimates add up to the overall result?

Observed sales should reconcile when segments are mutually exclusive and cover the same population and dates. Separately fitted counterfactuals need not add up to a model fitted to the total. Compare the summed incremental-sales estimates and document any difference; do not take a simple average of segment ROAS values.

Can MMM still differ after we reconcile the holdout analysis?

Yes. The experiment measures a particular intervention and time period, while the model may use a different outcome scope or translate that evidence into a broader estimate. Our article on why calibrated MMM and conversion-lift ROAS can differ explains that separate comparison. For the process of bringing experiment evidence into the model, see our guide to MMM calibration.

How Sellforte helps

Sellforte's experiment analysis shows observed sales against a counterfactual, cumulative treatment effects, and the associated spending change. Your team can use these views to trace a disputed result to the baseline and inspect the pattern behind the headline lift. Book a demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.