How should we weight multiple incrementality tests when calibrating an MMM?
A senior marketing analytics leader at a large ecommerce company had hundreds of conversion-lift studies spanning various campaign objectives and needed a defensible way to use them in feature-level MMM calibration.
The short answer
Weight multiple incrementality tests with a systematic framework covering relevance, recency, statistical confidence, and spend. First map every test to the MMM feature it informs. Exclude unreliable results, split or allocate mixed-objective tests where defensible, and give the strongest matching evidence the greatest influence on each feature's calibration.
This article is part of Asked by Marketers, a series answering real questions from marketing leaders.
Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →
Why do marketers ask this?
Running one incrementality test creates a result. Running dozens creates a calibration challenge.
The tests may cover different markets, dates, campaign objectives, spend levels, and conversion events. Some test cells map neatly to one MMM feature. Others mix sales, traffic, engagement, and awareness campaigns. Statistical confidence also varies.
The difficulty becomes obvious once a company has built a substantial experiment program. More tests should improve the MMM, but only if each test is mapped to a relevant MMM feature and its signal quality is evaluated against the same framework for relevance, recency, statistical confidence, and spend.
The goal is a repeatable process for bringing the best current evidence into the calibration of each MMM feature. The framework below is one practical way to do that.
Why should you not average the test iROAS?
When several tests are available, the easiest options are to average every result, average only the latest three, or use only the most recent test.
These shortcuts make the calibration workflow easy to explain, but they ignore material differences in relevance, recency, statistical confidence, and spend.
Suppose one test is low quality, one is medium quality, and one is high quality. A simple average gives them equal influence. The calibration should instead draw most heavily on the highest-quality test.
How this looks in practice
The weighting workflow starts whenever a new incrementality test ends. It has five steps.
1. Analyze the test. Use the method appropriate for the test type. The example below shows the output of a synthetic-control geo test, including iROAS, confidence interval, spend, and other diagnostic results.
2. Add the test to the experiment library. The library stores every test so the results can inform MMM calibration together. The example below combines geo tests, conversion-lift tests, and A/B tests in one view, with filtering options available on the left.
3. Connect the test to an MMM feature. Map the new result to the one or more specific MMM features it should inform. If a platform-wide test covers several channels or campaign types, apply a consistent mapping rule to assign the relevant part of the result to each feature.
4. Recalculate the test weights for the affected MMM features. Evaluate the connected tests for each feature against the same framework and calculate new calibration weights. The next section defines the four factors.
5. Calculate the new MMM feature priors and integrate them into the MMM. Use the weights to form new priors and pass them into the model. The calibrated results become available in the next refresh cycle, which may be daily, weekly, or monthly depending on the marketer's setup.
What factors should a systematic framework for incrementality test weighting include?
We have found that the following four-factor framework works for many of our customers:
1. Relevance asks how closely the test matches the MMM feature and decision. A test of the same market, objective, KPI, and campaign type deserves more influence than a broad platform test that only partly overlaps. When a mixed test has been allocated using an imperfect proxy, its relevance score should reflect that limitation.
2. Recency asks whether the underlying media system is still comparable. An older test may remain useful when targeting, creative format, auction dynamics, conversion tracking, and customer behavior are stable. It should lose influence when those conditions have changed. There is no universal expiration date, so the decay rule should be explicit and reviewed by channel and market. If recency is scored numerically, recalculate it whenever a new test is added or the calibration is refreshed.
3. Statistical confidence reflects how precise the experiment's estimate is. Lower-confidence tests should get less weight than tests with higher confidence. Teams typically also set a threshold for excluding tests with very low confidence.
4. Spend determines how much observed activity sits behind the result. All else equal, tests with more spend get higher weight.
Give each factor a numerical weighting score, then combine them into one score for each test.
Related questions
Should low-confidence incrementality tests always be excluded?
No. Exclude a test when execution or data quality makes it invalid for the calibration question. When the design is valid but the estimate is imprecise, down-weight it or carry its wider uncertainty into the model. In either case, retain the result in the experiment library.
How recent must an incrementality test be for MMM calibration?
There is no fixed shelf life. A test remains relevant while the market, targeting, creative format, conversion tracking, and campaign objective remain comparable. Use an explicit recency rule, and test whether a reasonable alternative changes the decision.
Can a mixed-objective lift study still calibrate MMM?
Yes, if its campaigns can be mapped and the result can be allocated with a defensible, consistent key. If that allocation is weak, use the result at a broader level or reduce its influence. Do not apply the full result independently to every objective.
Why can the calibrated MMM still differ from a lift study?
The experiment measures one defined setup, while MMM applies the evidence across a wider period and model feature. The difference is defensible only when the mapping, KPI, prior, time window, and uncertainty remain traceable. See Why can a calibrated MMM show a different incremental ROAS than a conversion-lift study?.
How Sellforte helps
Sellforte connects conversion-lift, geo-lift, and other experimental results to the MMM features they inform. Teams can inspect the mapping, allocation key, confidence, and feature-level calibration instead of relying on one opaque platform average. Book a demo.
Authors

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.
