How should we prioritize incrementality tests across markets and channels?

7 min read
Sep 1, 2026

A senior marketing leader at a large ecommerce company needed to turn a long list of market and channel test ideas into one roadmap without spending time and budget on experiments that would add little new evidence.

The short answer

Prioritize incrementality tests by scoring each market and channel combination on business impact, uncertainty, feasibility, and the commercial cost of the test. Test high-spend decisions with weak or conflicting evidence first. Once a result is credible and current, lower that item's priority and move the roadmap toward the next material uncertainty.

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why do marketers ask this?

An experiment roadmap becomes crowded as soon as a marketer operates across several countries. Ten markets and eight channel groups already create 80 possible cells before the team separates campaign objectives, customer segments, or seasonal periods. No team can run all of those tests at once.

The usual shortcuts break down quickly. Testing the largest channel every quarter can repeat what the team already knows. Testing the easiest platform cell can produce a precise answer to a small question. Filling every blank in the matrix can consume years while material budget decisions still rest on weak assumptions.

Treat incrementality testing as a portfolio of learning investments. A useful test should reduce uncertainty around a decision that matters. Its expected value is the chance that better evidence changes future budget allocation, minus the sales, margin, media, and analyst time consumed by the test.

What belongs in the priority score?

Start with one row for every market and channel combination. Add campaign objective when objectives behave differently enough to support separate tests. Then score four clusters: impact, uncertainty, feasibility, and cost. The score should explain the order of the roadmap, not replace judgment.

Score cluster Questions to answer Higher priority when
Business impact How much is spent? How large is the budget decision? Could a different answer change profit or growth? The spend or steering opportunity is material
Evidence gap Is there a direct local test? Is the evidence recent, precise, relevant, and consistent? Evidence is borrowed, stale, imprecise, or conflicting
Feasibility Can the treatment be isolated? Is the expected effect detectable? Is the KPI available at the right level? The design can answer the question with useful precision
Cost and effort How much demand or margin may be withheld? What setup, lead time, and analysis workload are required? The learning is inexpensive and operationally realistic

Business impact should include more than spend. A smaller channel can rank highly when the business is about to scale it, when performance estimates differ enough to reverse the allocation decision, or when the test will inform several comparable markets. Conversely, a large channel with recent direct evidence may not need another test.

Treat confidence as an evidence gap: weaker confidence means a stronger reason to test. Record the calibration source, date of the last test, number of relevant tests, uncertainty interval, and whether results agree. A Marketing Mix Model calibrated with a direct experiment for the same channel and market usually needs less new evidence than one relying on an industry benchmark or a weak prior.

Which evidence gaps should be tested first?

Prioritize the gaps where a plausible alternative result would change a real decision. This is the difference between uncertainty and decision-relevant uncertainty. A wide range is not urgent if every credible value leads to the same budget choice.

Use an evidence ladder for each row:

  1. Direct experiment for the same channel and market
  2. Closely related channel experiment in the same market
  3. Same channel in a structurally comparable market
  4. Relevant industry benchmark
  5. Weak or uninformative prior

Rows near the bottom deserve attention when the business impact is material. They do not all deserve immediate tests. Evidence from another market can be a reasonable interim prior, especially for similar digital mechanics, while print, promotions, and local media often need direct market-level evidence.

Conflicting evidence also raises priority. If several valid tests disagree after the team checks scope, KPI, campaign mapping, and execution quality, another test can be more valuable than the first test in an unmeasured but immaterial cell. The new design should target the reason for disagreement rather than simply add another point estimate.

When should a test candidate be deferred?

Defer a candidate when the test cannot produce a decision-quality answer, when the evidence already covers the question, or when the commercial cost is out of proportion to the learning.

Run a power or minimum detectable effect check before the roadmap is approved. A small channel may have a large evidence gap but still be unable to move total sales enough for a geo test to detect. In that case, the right response may be to pool comparable markets, use a platform conversion-lift design, rely on a wider prior, or wait until spend reaches a testable level.

Analyze existing tests before commissioning more. I have seen examples where a well-covered platform activity moved down the list, while a channel with surprising and inconsistent results moved up. In another case, several historical tests first had to be mapped and reviewed. More testing would have added cost before the team had extracted the value of the evidence it already owned.

Timing matters as well. A test during a peak trading period may carry high lost-margin risk and may answer a seasonal question rather than provide a useful annual baseline. Long-lead geo and print tests belong in quarterly planning. Faster conversion-lift studies can run through a monthly or continuous program, provided each one answers a defined gap rather than filling the calendar.

How should the roadmap change over time?

Recalculate the roadmap whenever a credible result arrives, a channel changes materially, or a major planning cycle begins. A completed test should improve the evidence score for the cells it actually informs and push them down the queue. This keeps testing focused on marginal learning rather than permanent channel entitlements.

A practical cadence has two levels. Use a quarterly steering session to select geo-lift, print, and other tests with long lead times or cross-team dependencies. Let channel teams schedule faster platform studies monthly within the agreed priorities. Keep both in one master country and channel matrix so local execution does not fragment the evidence base.

The roadmap should also track collisions. Two experiments that alter the same audience, geography, or outcome can contaminate each other or make interpretation harder. Avoid overlap where practical. When it cannot be avoided, record it in advance and decide whether the designs remain usable before either test starts.

How this looks in practice

Consider a hypothetical retailer reviewing five candidates for the next quarter. The figures and decisions below are illustrative and do not reproduce any customer data.

Candidate Impact and evidence Feasibility and cost Roadmap decision
Market A, paid-search brand
$4.0M annual spend
No direct test and a large gap between attributed and expected incremental return Clean holdout is feasible; moderate demand risk Test first
Market B, display and CTV
$2.5M annual spend
Benchmark prior only; result would affect the next brand budget Geo test is feasible but needs eight weeks of setup Design now for next quarter
Market A, paid social
$6.0M annual spend
Twelve recent, relevant lift cells agree within their uncertainty Easy to test, but another result has low marginal value Monitor, do not repeat yet
Market C, short-form video
$0.4M annual spend
No local evidence, but limited current budget impact Expected lift is below the geo design's detectable effect Defer or pool markets
Market D, print
$3.2M annual spend
A previous test predates new targeting and format choices High operational effort and a six-month lead time Reserve the next viable window

The largest spend line is not automatically first. Paid social has the highest spend, but another study would add little. Paid-search brand comes first because the decision is material, evidence is weak, and the test is viable. Display and CTV may be equally important, but the long setup time means design work starts now for a later launch.

Experiment library used to prioritize incrementality tests across markets and channels
An experiment library supplies the evidence inputs for the roadmap: market, channel, test type, dates, result, and uncertainty. Demo data

The master matrix should show why each row ranks where it does. A useful roadmap is auditable enough for finance, channel owners, and analytics teams to challenge the inputs without turning the score into false precision.

Should the largest channel always be tested first?

No. Spend raises the potential business impact, but priority also depends on what is already known and whether a different answer would change the decision. A slightly smaller channel with weak or conflicting evidence can have greater learning value than a large channel supported by several recent tests.

Can evidence from another market lower test priority?

Yes, when the channel mechanics, campaign objective, KPI, and market conditions are comparable. Treat the evidence as a wider prior rather than a local fact, and keep direct local testing higher on the evidence ladder. See Can we use incrementality evidence from one country to calibrate an MMM in another?.

How do existing test results affect the roadmap?

Map every valid result to the exact features it informs, then update confidence, recency, and coverage for those rows. Several weak or mismatched tests should not create false confidence. See How should we weight multiple incrementality tests when calibrating an MMM?.

Should geo-lift and conversion-lift have separate roadmaps?

They can have different planning cadences, but they should share one evidence matrix. Conversion lift often answers granular platform questions quickly, while geo-lift can cover broader business outcomes and channels without user-level holdouts. See How should geo-lift and conversion-lift studies be combined when calibrating MMM?.

How Sellforte helps

Sellforte brings geo-lift, conversion-lift, and other experiment results into one library and connects them to the MMM features they inform. Teams can see where material spend rests on strong direct evidence and where weak, stale, or borrowed evidence makes the next test more valuable. Book a demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.