What is the right holdout-group size for an incrementality test?

7 min read
Aug 27, 2026

Context: A marketing analytics professional at a large ecommerce company needed to size country-level catalog holdouts without sacrificing more demand than the evidence was worth.

The short answer

Choose the smallest holdout group that gives the incrementality test enough power to detect a commercially meaningful lift with a useful confidence interval. Calculate it from baseline demand, expected lift, natural variance, test duration, and allocation design, then cap it where the revenue or margin withheld would cost more than the decision is worth.

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why do marketers ask this?

The statistical answer and the commercial answer pull in opposite directions. A larger holdout can make a small effect easier to detect, but it also withholds advertising, catalogs, or other marketing from more customers or regions. That can mean lost revenue, margin, or reach before the team has learned anything.

The percentage also looks deceptively portable. One large market may have enough customers and conversions for a 5% control group. A smaller market running the same campaign may need 10% or more and still produce a wide interval. A geo test can require a much more balanced split because it has far fewer experimental units and depends on the quality of the regional counterfactual.

This is why incrementality testing needs a design decision before it becomes a measurement exercise. The holdout must be large enough to answer a question that could change spend, but no larger than the value of that answer justifies. If the result will also calibrate a Marketing Mix Model, the KPI and test scope must match the model feature the evidence will inform.

What determines the right holdout-group size?

Five inputs determine the right holdout-group size. None can be replaced by a universal percentage.

Input Question to answer Effect on the holdout
Expected lift How large a causal effect is plausible at the planned spend level? Smaller effects usually require more information
Baseline and variance How often does the outcome occur, and how much does it fluctuate without treatment? Rare or noisy outcomes need a larger sample or longer test
Decision threshold What lift or iROAS would change the budget decision? The design should resolve this boundary, not any positive effect
Test duration and design How long will the test run, and how are customers or regions assigned? More observations can reduce the group size, if the business remains comparable
Commercial cost How much demand, margin, or reach could withholding marketing put at risk? The business may set a hard cap before the statistical optimum

Start with the decision threshold. If the business would keep spending whenever iROAS could plausibly exceed 1.5, the test needs enough precision to distinguish that boundary. Designing only to detect whether lift is greater than zero can produce a statistically positive result that still cannot answer the budget question.

Use historical data to estimate the baseline, variance, seasonality, and expected lift. For geo tests, also test whether the untreated regions can reproduce the treated region before the intervention. A weak donor pool can make a large holdout look powerful on paper while leaving the counterfactual unreliable.

Why do customer and geo tests need different holdout splits?

Customer-level and geo tests use different experimental units, so the same percentage does not carry the same amount of information.

In a high-volume customer-level test with a strong expected effect, a 5% to 10% holdout can be enough. The absolute number matters: 5% of several million eligible customers is still a substantial control group. Before approving it, calculate the minimum detectable effect for the actual baseline conversion rate, treatment size, control size, and planned duration.

Geo tests usually have far fewer units. A country might contain millions of customers but only dozens of usable regions. In many practical designs, holding out roughly 30% to 50% of eligible geographies creates a stronger comparison. The best split still depends on regional size, similarity, spillover, media delivery, and whether the remaining regions form a credible control. Going beyond a balanced split can reduce the amount of treated signal without improving the counterfactual.

Platform conversion-lift studies are different again because the platform controls random assignment and reach. Marketers should review the planned holdout, expected conversions, number of test cells, and reported power, but they may not be able to choose every allocation detail themselves.

How should marketers use MDE and confidence intervals?

Use minimum detectable effect, or MDE, to check the design before the test and the confidence interval to make the decision after it.

MDE is the smallest effect the planned design can detect with the chosen power and error thresholds. Calculate it for the real test duration, not an arbitrary fraction of the available history. A power curve is more useful than one MDE number because it shows the tradeoff among holdout size, duration, and detectable lift.

Then compare MDE with the effect that matters. Sellforte practitioners sometimes use expected lift of at least twice the MDE as a conservative screening rule for customer holdouts with strong historical signals. It is not a statistical law. It is a way to avoid launching a costly test that only works if every assumption is favorable.

After the test, keep the full interval. A result does not become useless simply because the interval includes zero. Suppose the commercial rule is to continue only when iROAS could plausibly exceed 1.5. An estimated iROAS of 0.7 with a 90% interval from 0.1 to 1.3 would still provide evidence against continued investment because even the upper bound falls below the threshold. An interval from -0.5 to 2.2 would not resolve the decision.

What if the required holdout is too costly?

Do not automatically enlarge the holdout until the power calculation passes. First check whether a different design can answer the same commercial question at lower cost.

A longer test may add information without withholding marketing from more units, provided seasonality and campaign conditions remain stable. A better-matched set of control regions can improve geo-test precision. Several comparable markets can sometimes be pooled, but only when the KPI, channel mechanics, campaign objective, and market conditions support that assumption.

The KPI also matters. Orders may provide a stronger statistical signal than revenue in some designs, but the team should not switch outcomes merely to achieve significance. The test must still measure the result that finance and marketing will use.

If a low-lift market would require most of the customer base to sit in the holdout, set a commercial cap and be explicit about the remaining uncertainty. A capped test can still be worthwhile when its interval can rule out an acceptable return. If it cannot, defer the experiment, borrow evidence cautiously from a comparable market, or ask whether an effect too small to measure is material enough to justify the current spend.

How this looks in practice

Consider two hypothetical catalog tests. The figures below illustrate the design logic and do not reproduce customer data. Expected lift and MDE use the same relative sales-lift definition.

Design input Market A Market B
Eligible customers 4,000,000 600,000
Expected relative sales lift 9% 2%
Candidate holdout 5%, or 200,000 customers 10%, or 60,000 customers
Modeled MDE 4% 6%
Design decision Proceed. Expected lift is more than twice the MDE, and the holdout remains within the commercial limit. Do not enlarge automatically. The expected lift is below the MDE, and the 10% cost cap has been reached.

Market A can use a smaller percentage because its customer base and expected signal are larger. Market B has the opposite problem. Even doubling its holdout may destroy more value without making a 2% lift reliably detectable. The useful options are to extend the test, improve the design, accept a wider interval tied to a specific decision rule, or not run it.

A pre-test design view should show this tradeoff as a power curve or an MDE-by-duration curve, with the commercial cap visible beside it. Once the test ends, the result view should keep the point estimate and confidence interval together so the team can compare both with the decision threshold.

Is 10% a good holdout-group size?

Ten percent can be a sensible starting cap for a large customer-level test, but it is not a default answer. Calculate the MDE in absolute customers and conversions, then compare the expected lift and commercial exposure. A smaller group may be enough in a large market, while 10% may still be underpowered in a small one.

Should geo tests use a 50/50 split?

A balanced split often improves statistical efficiency, but geo quality matters more than symmetry. Region size, pre-test similarity, spillover, media delivery, and donor-pool quality can justify a different allocation. The design should be backtested on historical data before any region is withheld.

Can an inconclusive incrementality test still change a decision?

Yes. A wide interval can still rule out the return required to justify the investment, even when it does not prove a positive effect. Read the interval against the commercial threshold instead of treating statistical significance as a pass or fail label.

How does holdout size affect the test roadmap?

Commercially expensive or underpowered tests should fall behind questions that can be answered with useful precision at reasonable cost. Recalculate the roadmap when expected spend, lift, or test feasibility changes. See How should we prioritize incrementality tests across markets and channels?.

How Sellforte helps

Sellforte helps teams design geo-lift and other incrementality tests around a decision-relevant MDE, a credible counterfactual, and a realistic test duration. Results can then be analyzed in Experiments Hub and connected to MMM calibration with the uncertainty kept visible. Book a demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.