When is a market or channel too small to justify an incrementality test?

8 min read
Sep 3, 2026

Context: A marketing analytics professional at a large ecommerce company needed to decide whether low-uplift market-level catalog tests were worth the customers withheld and the effort required.

The short answer

A market or channel is too small to justify an incrementality test when no commercially acceptable design can distinguish the effect that would change the budget decision, and the decision is not material enough to justify the test's cost. Check power and confidence intervals. If neither resolves a meaningful threshold, redesign, combine comparable evidence, wait, or defer the test.

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why do marketers ask this?

International marketers rarely face one clean testing decision. They face a grid of countries, channels, campaign objectives, and customer groups. The largest cells can produce clear sales signals. In smaller cells, the expected effect may disappear inside normal weekly variation even when the advertising works.

That creates two expensive mistakes. One is to keep enlarging a holdout until the test works on paper, even if the withheld demand costs more than the answer is worth. The other is to cancel every test that might miss a conventional significance threshold, leaving low-return spend untouched because nobody can produce a perfect local estimate.

Incrementality testing should resolve a decision, not fill every empty cell in a measurement matrix. A small market or channel can still justify a test when the result will determine whether the business stops, scales, or reallocates meaningful spend. A larger one may not justify another test when recent evidence already answers the question.

What does “too small” actually mean?

“Too small” is not a fixed spend, audience, or market-size cutoff. It means the candidate fails both a statistical gate and a commercial gate.

Gate Question Reason to defer
Statistical Can a feasible design distinguish an effect large enough to change the decision? The minimum detectable effect exceeds the largest credible effect, and the expected interval will not resolve the decision threshold
Commercial Is the budget decision valuable enough to justify withheld demand, setup work, and analysis? Even a perfect answer would not change a material allocation, or the test costs almost as much as the available upside

Market size matters because it affects the amount and stability of outcome data. Channel size matters because it limits the maximum signal the test can create. If a channel could plausibly move total market sales by 0.5%, but the best feasible geo design can detect only changes above 2%, that design cannot answer the question. A larger control group cannot manufacture treatment signal that does not exist.

This does not prove the channel has no effect. It says the proposed experiment cannot measure that effect precisely enough for the intended decision. Keep those statements separate.

How should you test feasibility before launch?

Start with the business threshold, then work backward to the design. Do not begin with a familiar holdout percentage or the effect needed for a p-value.

  1. Define the decision. State what will happen if incremental ROAS or lift falls below, reaches, or exceeds a specific level.
  2. Choose the right experimental unit. Decide whether customers, platform users, stores, or geographies can be assigned cleanly without unacceptable spillover.
  3. Estimate the plausible effect. Use current spend, reach, historical tests, baseline demand, and existing model evidence. Keep the KPI consistent with the eventual budget decision.
  4. Simulate the feasible design. Calculate power, minimum detectable effect, and the expected interval across realistic durations and treatment splits. For geo tests, backtest whether the control regions can reproduce the treated regions before the test.
  5. Price the learning. Estimate margin or demand at risk, media changes, analyst effort, operational lead time, and the cost of a delayed decision.

The useful comparison is not expected lift versus zero. It is the range of results the test is likely to produce versus the threshold that changes the action. A test can have adequate power to show a positive effect and still be too imprecise to distinguish an incremental ROAS of 1.5 from 3.0. If finance requires 2.5, that design has not solved the problem.

If the evidence will also calibrate a Marketing Mix Model, match the test to the market, channel, campaign objective, KPI, and spend range the model feature represents. A precise result for a different scope is not a precise answer to the local question.

When can an underpowered test still be worth running?

An underpowered test can be worth running when its likely confidence interval can still rule out a commercially acceptable return. “Not statistically significant” and “not decision-useful” are not synonyms.

Suppose the business requires an incremental ROAS of 2.5. A test estimates 0.7 with a 90% interval from minus 0.2 to 1.9. The interval includes zero, so the team cannot make a strong claim that the effect is positive. Yet the upper bound remains below 2.5, which is useful evidence against continuing the investment at its current terms.

Now take an estimate of 1.1 with an interval from minus 1.0 to 4.2. That result supports several opposing decisions. The channel could destroy value or comfortably exceed the hurdle. Unless the test will narrow the range enough to separate those actions, it is unlikely to justify much commercial sacrifice.

The design review should therefore forecast the interval, not only the probability of rejecting zero. After the test, keep the point estimate and full interval together. Do not relabel a noisy result as proof of no effect.

How can you make a small market or channel testable?

When the first design fails, change the design before abandoning the question. The best option depends on what makes the signal weak.

Adjustment When it helps Limit to watch
Run longer More observations improve precision while market conditions remain comparable Seasonality, promotions, creative changes, and fatigue can make a long test answer a different question
Create a stronger contrast A full pause in selected units produces more signal than a shallow cut everywhere Lost demand and stakeholder risk can become unacceptable
Change the test type A user-level conversion-lift study may work when a channel is too small relative to total sales for geo lift Platform tests cover only eligible activity and observable outcomes
Improve the counterfactual Better-matched regions or variance reduction can tighten the interval without a larger holdout A weak donor pool or unstable pre-period fit remains a validity problem
Pool comparable cells Several small markets share the same channel mechanics, KPI, objective, and decision Brand maturity, customer mix, promotions, or local execution can make the pooled answer misleading

Do not switch to a more frequent KPI merely because it makes the test significant. Orders may be less noisy than revenue, for example, but only use them when the resulting decision can be translated back to profit or customer value. Precision on the wrong outcome is false comfort.

What should you do when a local test is not justified?

Use the strongest broader evidence available and keep its uncertainty visible. The usual fallback order is a closely related channel test in the same market, the same channel in a structurally comparable market, a relevant industry benchmark, and finally a weak prior. None should be presented as a precise local fact.

Pooling or borrowing evidence works best when platform mechanics, campaign objective, KPI, brand maturity, audience, and trading conditions are close. Print, promotions, and other locally executed activity often transfer less cleanly than large digital platforms. A calibrated MMM can use broader evidence as a cautious prior while letting local time-series data update the estimate.

Waiting is also a valid choice. A growth channel that is too small today may become testable after its spend and reach increase. Set a trigger, such as a spend level, number of conversions, or feasible MDE, so “later” becomes a deliberate rule rather than a permanent blind spot.

Finally, challenge the investment itself. If even an optimistic effect would not clear the commercial hurdle, the more useful question may be why the company continues to spend there. This conclusion must come from the plausible effect range and the business threshold, not from the bare fact that one design was underpowered.

How this looks in practice

Consider four hypothetical candidates in an international retailer's test roadmap. The figures are illustrative and do not reproduce customer data.

Candidate Feasibility and decision Recommendation
Market A, brand search
$1.5M annual spend
The feasible geo design has a 1.2% sales-lift MDE. Existing evidence puts plausible lift between 1.8% and 3.0%, and the result could change a material budget. Run the local test.
Market B, short-form video
$250K annual spend
The geo design can detect a 2.5% change in total sales, but even the optimistic channel effect is 0.7%. Do not run geo lift. Check a user-level study or pool genuinely comparable markets.
Market C, catalog
$450K annual spend
A customer holdout capped at 10% is unlikely to prove positive lift, but the expected iROAS interval can rule out the 2.5 hurdle if performance is weak. Run the capped test. Judge the interval against the hurdle.
Market D, niche affiliate
$80K annual spend
The setup and likely margin at risk approach the maximum annual budget that a different answer could redirect. There is no plan to scale the channel. Defer. Use broader evidence and revisit only if the decision grows.

Market C is the important exception. Its test may be underpowered against zero and still be worth running because it can answer the keep-or-stop question. Market B needs a different method, while Market D fails the commercial gate even if a valid design exists.

Is an underpowered incrementality test always worthless?

No. It can still be useful when the confidence interval rules out the return needed to continue or scale the investment. It is worthless for a particular decision when the plausible range still supports opposing actions and the design cannot narrow it enough.

What is the right holdout-group size?

Choose the smallest holdout that can resolve a commercially meaningful effect, then apply a limit for demand or margin at risk. There is no universal percentage because customer-level and geo tests have different experimental units. See What is the right holdout-group size for an incrementality test?

Can incrementality evidence transfer across markets?

Yes, as a cautious prior when the channel mechanics, KPI, objective, customer behavior, brand position, and market conditions are sufficiently similar. Local evidence should update or replace it when available. See Can we use incrementality evidence from one country to calibrate an MMM in another?

How should we prioritize the candidates that are testable?

Rank them by business impact, current evidence gap, feasibility, and the cost of learning. A large uncertainty is not urgent when every plausible answer leads to the same action. See How should we prioritize incrementality tests across markets and channels?

How Sellforte helps

Sellforte helps teams design and analyze geo-lift, conversion-lift, and other incrementality tests, keep confidence intervals visible, and connect the resulting evidence to MMM calibration. Teams can focus testing on market and channel decisions where better causal evidence is both feasible and commercially valuable. Book a demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.