Should we run a full marketing blackout or a partial holdout test?
Context: When finance required an immediate, material reduction in media spend, a senior marketing leader at a large ecommerce company had to decide whether to stop marketing across whole markets or keep some regions active as a control.
The short answer
Run a partial geo holdout in most cases: keep comparable regions at normal spend and pause the selected marketing in the remaining regions. This preserves a local counterfactual, limits lost demand, and isolates the decision better. Use a broad blackout only when the strongest possible total-media signal or an unavoidable budget cut outweighs the loss of diagnostic detail.
This article is part of Asked by Marketers, a series answering real questions from marketing leaders.
Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →
Why do marketers ask this?
A blackout promises clarity. If marketing stops and sales fall, the result is easy to explain to finance and senior leadership. It can also create a large enough change to measure when smaller adjustments would disappear inside normal demand variation.
But a blackout can mean two very different things. One version stops all marketing everywhere in a market, leaving no local control. Another switches marketing off completely in selected regions while matched regions continue at normal spend. Both create a sharp exposure contrast, but only the second preserves a within-market counterfactual.
This distinction matters because incrementality testing is not just about making spend change. It is about estimating what would have happened without that change. A dramatic before-and-after chart is still weak causal evidence when seasonality, promotions, stock, competitors, or the economy moved at the same time.
The decision also has a commercial cost. A test designed to prove that marketing works can destroy more demand than the answer is worth. The right design balances signal strength, diagnostic value, implementation speed, and the sales or margin put at risk.
What is the real choice between a blackout and a holdout?
The useful choice is not simply full reduction versus partial reduction. It is whether the business can preserve a credible untreated comparison while creating a strong enough treatment.
| Design | Counterfactual | What it can estimate | Main tradeoff |
|---|---|---|---|
| Market-wide blackout All paid media stops everywhere |
Historical baseline, synthetic control, or another market | Combined effect of the media that stopped | Maximum disruption, weak local comparison, no channel isolation |
| Regional blackout Selected regions go to zero; others stay normal |
Normal-spend regions in the same market | Combined effect of the media that stopped in treatment regions | Strong signal, but several channels can still be bundled together |
| Channel-specific geo holdout One channel stops in selected regions |
Regions where that channel continues normally | Incremental effect of the selected channel | More diagnostic, but the smaller effect may need more time or units |
| Uniform partial cut Every region and campaign is reduced by the same percentage |
No untreated group | Usually a modeled or historical estimate, not a clean experiment | Easy to execute, but often creates little causal learning |
Reducing every region by 20% is not the same as keeping 20% of regions at normal spend and switching the rest off. The first design changes everyone and leaves no untreated comparison. The second creates a clear 100-versus-0 contrast between matched geographic groups.
Why is a partial holdout usually the stronger design?
A partial holdout usually gives the business a better answer because it preserves a local counterfactual. Test and control regions share more of the same economy, calendar, brand position, and operating conditions than two different countries usually do. That does not make them identical, but it reduces the amount of reconstruction needed after the fact.
A common failure mode is an 80% cut everywhere that leaves the remaining 20% working disproportionately well. The business believes it ran close to a blackout, but the residual campaigns keep reaching the best audiences and inventory. A cleaner intervention is often binary inside each region: normal spend in control regions and zero spend for the selected activity in treatment regions.
A channel-specific holdout is also easier to use. If paid search alone changes, the result can inform the paid-search decision and the corresponding feature in a Marketing Mix Model. When paid search, paid social, display, and app campaigns all stop together, the experiment estimates their combined effect. It cannot allocate that effect across the channels without additional evidence.
Partial does not automatically mean low risk. A large treatment group at zero spend can still put substantial demand at risk. Size the experiment from baseline variance, expected lift, minimum detectable effect, test duration, and the commercial cost of withholding marketing. There is no universal holdout percentage.
When does a full marketing blackout make sense?
A broad blackout is most defensible when the business needs an unusually strong combined-media signal, a mandatory cut has already created the treatment, or there is not enough time to arrange channel-level studies. It can answer whether the marketing system as a whole produces measurable incremental demand.
It may also be useful as the first stage of a learning plan. A period with zero media creates a clean baseline. The team can then reactivate one channel or spend level at a time and measure the lift from turning activity back on. Those activation tests can be more diagnostic than the initial blackout.
Senior stakeholder skepticism is another legitimate reason, but it should be stated plainly. A combined blackout may provide visible evidence that marketing drives sales. That does not make it the best design for allocating next quarter's budget. The proof question and the optimization question are different.
Even when a blackout is justified, avoid switching off every geography if a normal-spend control can be preserved. Keep the duration no longer than the power and outcome window require, define conditions for restoring spend, and record which channels and outcomes the result can support. Treat one market-wide result as one data point, not a universal verdict on marketing.
How should you design the partial holdout?
Start with the decision, then build the smallest experiment that can change it.
- Define one intervention. Specify the channels, campaigns, spend level, regions, dates, and audience rules that will change. If several channels must stop together, label the result as combined media.
- Choose the business outcome. Use the sales, profit, orders, leads, or customer value measure that matches the decision. Confirm that it is available consistently by region and over the required conversion window.
- Create comparable groups. Match regions on historical outcome patterns and important business drivers. Check that targeting boundaries mean the same thing across platforms and that sales data maps to those boundaries.
- Run the power and cost analysis together. Estimate the treatment size and duration needed to detect a decision-relevant effect, then quantify the demand or margin that may be withheld. Redesign the test if the commercial exposure is disproportionate to the learning.
- Predefine the analysis. Record the pre-period, test window, confidence level, exclusion rules, spillover assumptions, stopping rule, and restoration plan before the result is visible.
- Monitor comparability, not just significance. Promotions, stock quality, prices, holidays, competitor moves, and local operations can make test and control regions diverge. If that happens, use the last defensible cutoff or treat the result as inconclusive rather than choosing a favorable window.
- Allow for delayed effects. Continue the outcome analysis through the relevant post-test conversion or carryover window even after media is restored.
Operational details can decide whether the design survives contact with reality. Confirm region names with maps, keep a shared activation sheet, verify platform exclusions after launch, and assign one owner for the analysis and one for campaign execution.
How this looks in practice
Consider a hypothetical ecommerce company with 24 usable regions and a finance-mandated media reduction. The team matches the regions on two years of weekly sales, then keeps six regions at normal spend and pauses all paid media in the other 18 for six weeks. The figures below are illustrative and do not reproduce customer data.
| Element | Illustrative setup or result | What the team can conclude |
|---|---|---|
| Control | 6 matched regions at normal spend | Provides the within-market counterfactual |
| Treatment | 18 regions at zero paid-media spend for 6 weeks | Creates a strong combined-media contrast |
| Spend avoided | $1.2M | The planned saving was delivered |
| Incremental sales foregone | $4.8M, with a 90% interval from $3.2M to $6.4M | The removed media had a material causal effect |
| Implied incremental ROAS | 4.0, with an interval from 2.7 to 5.3 | The combined return is decision-useful, but not a channel allocation |
The company now has credible evidence that the combined media caused incremental sales. It still cannot tell how much came from paid search, paid social, display, or app campaigns. In a second stage, it can split the blackout regions again, reactivate one channel in one matched group, keep another group dark, and compare the slopes. That follow-up needs a fresh power check because it is estimating a smaller effect.
The result should also have a clear validity window. If stock, promotions, school holidays, or another demand driver causes the groups to stop moving together, the team should use the last predeclared defensible date or report that the later period is contaminated.
Related questions
Is zero versus normal spend better than a small cut everywhere?
Usually, yes, when the commercial risk is acceptable. Zero spend in selected treatment regions and normal spend in control regions creates a cleaner exposure contrast than reducing every region by the same percentage. A shallow cut everywhere may be operationally easy, but it often leaves no untreated comparison and can produce little causal learning.
Can a marketing blackout identify which channel works?
Not when several channels stop together. The result estimates their combined effect, including any interactions among them. To isolate a channel, change that channel while keeping other important activity consistent, or follow the blackout with staged channel reactivation.
How large should the holdout group be?
There is no universal percentage. Use baseline demand, regional similarity, expected lift, natural variance, test duration, and the commercial cost of withholding marketing to calculate a defensible design. A statistically neat split is not useful if the expected effect is too small to detect or the lost demand is too costly.
Can a mandatory budget cut become an incrementality test?
Yes, if the team can concentrate part of the reduction into a planned treatment and preserve a credible counterfactual. Define the KPI, power, test window, and restoration rule before the cut starts. See How can you turn a mandatory budget cut into an incrementality testing opportunity?.
How Sellforte helps
Sellforte helps teams design and analyze geo-lift experiments, compare observed outcomes with a counterfactual, and bring the resulting evidence into MMM calibration. Teams can see the estimated lift, uncertainty, and model comparison before deciding whether to restore, reduce, or reallocate spend. Book a demo.
Authors

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.
You May Also Like
These Related Stories

How should we cut our marketing budget while protecting growth?

Can you trust Meta Conversion Lift tests for MMM calibration?

