Why can a calibrated MMM show a different incremental ROAS than a conversion-lift study?

8 min read
Aug 11, 2026

A senior marketing analytics leader at a large ecommerce company saw a conversion-lift study report one incremental ROAS while the calibrated MMM showed a much larger number and needed to know whether the difference was legitimate or a calibration problem.

The short answer

A calibrated MMM can show a different incremental ROAS because a conversion-lift study measures one defined set of campaigns, users, dates, conversions, and test cells, while MMM applies that evidence across a broader model. The gap is acceptable only when the mapping, KPI, time window, prior, and uncertainty remain visible and defensible.

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why do marketers ask this?

Calibration sounds as if it should make two numbers match. A team runs a conversion-lift study, treats the result as causal evidence, and feeds it into MMM. When the MMM dashboard later shows a different incremental ROAS, the natural reaction is to ask what went wrong.

That reaction becomes stronger when the gap is large. The experiment may have been the main justification for changing the model, so a result that moves well beyond the study can feel circular: if the test is the anchor, why is the calibrated output somewhere else?

The answer depends on how the evidence moved from the test into the model. A conversion-lift result is not a universal channel constant. It belongs to specific test cells, campaigns, objectives, dates, conversion events, attribution settings, spend levels, and a stated level of uncertainty. MMM calibration must preserve that scope before the model can use the evidence elsewhere.

The size of the difference alone cannot tell you whether the calibration worked. The chain from test cell to model feature can.

Why should the two incremental ROAS figures differ at all?

They differ because the study and the model do not cover exactly the same thing.

A conversion-lift study estimates the incremental outcome caused by the campaigns included in the test. Its result reflects the audience split, experiment dates, conversion definition, post-test measurement window, spend, and campaign setup used at that moment. It is the strongest causal evidence for that defined test scope, subject to its confidence interval and execution quality.

A calibrated Marketing Mix Model covers a wider system. It may estimate the channel across more weeks, several campaign groups, different spend levels, and a broader sales KPI. The model also accounts for promotions, seasonality, baseline demand, other media, and the relationship between spend and returns. Calibration tells the model what the experiment taught us, but the model still has to reconcile that evidence with the rest of the data.

The experiment and MMM can therefore differ for legitimate reasons:

  • The MMM period includes weeks when the campaign mix or customer demand was different.
  • The experiment measured platform-observed conversions while MMM models total ecommerce or omnichannel sales.
  • The test covered selected campaigns or objectives, but the MMM feature covers a broader group.
  • The MMM estimates an average response across several spend levels, while the test observed one spend level.
  • The model incorporates lagged effects that fall outside the test's primary reporting window.
  • The experiment has a confidence interval, so its reported point estimate is not the only value consistent with the evidence.

Calibration should make the two sources coherent. It should not force a broad model output to copy one experimental point estimate in every period.

How does a conversion-lift result become an MMM calibration?

Start by matching the experiment to the same campaigns, dates, market, KPI, and spend in the attribution or performance dataset. That matched comparison is what connects a narrow experiment to a model feature.

One practical approach is to calculate an incrementality factor:

incrementality factor = experiment iROAS / attributed ROAS for the matched test scope

Suppose a conversion-lift study reports an iROAS of 4.0 and the matched attribution data reports a ROAS of 8.0 for the same campaigns and dates. The factor is 0.5. If the attributed signal for the broader model feature has a prior ROAS of 10.0, applying the factor gives a calibrated prior of 5.0. The model then updates that prior using the full MMM data and may produce a posterior iROAS above or below 5.0.

That arithmetic explains how the calibrated MMM can exceed the test result without ignoring it. The 4.0 was used to estimate that the matched attribution ROAS of 8.0 overstated incremental performance, producing a factor of 0.5. That factor was then applied to a different, broader attribution level before the MMM fit.

This approach is valid only if the numerator and denominator describe the same test scope. If the experiment combines traffic, engagement, and awareness campaigns but the attribution denominator mainly observes click-heavy traffic campaigns, the ratio is not a clean correction factor. Applying it to every objective can magnify the mismatch.

When is the difference acceptable?

The difference is acceptable when the team can trace it from the experiment to the final model result and each step still reflects the decision the marketer is making.

Before accepting the model result, ask four concrete questions:

  1. The scope matches. The experiment is mapped to the campaigns, market, objective, dates, KPI, and test cell it actually measured.
  2. The comparison basis is consistent. Experiment and attribution ROAS use compatible spend, currency, conversion value, returns treatment, and measurement windows.
  3. The model respects uncertainty. A noisy test influences the prior less than a precise one. Low-confidence results are excluded or down-weighted rather than treated as exact truth.
  4. The extension is explicit. If the factor is applied beyond the tested period or pooled across similar campaigns, the model records that assumption and shows where direct evidence ends.

An MMM result does not have to equal the lift-study point estimate to pass these checks. It does need to stay plausible relative to the study's uncertainty, the model data, and results from comparable markets or campaigns.

What should you check when the gap is large?

Treat a large gap as a calibration audit, not as a debate over which dashboard wins.

First, rebuild the matched test cell. Confirm that the campaign and ad-set IDs in the experiment are the same ones used in the attribution comparison. Check whether campaign objectives were pooled, whether any cells were omitted, and whether the model feature contains activity the test never covered.

Second, recalculate the factor from raw inputs. Compare spend, currency, revenue, conversion event, attribution window, and post-test window. Ratios can avoid some unit problems, but they do not repair different KPIs or mismatched campaign groups.

Third, inspect the prior before and after calibration. A reasonable factor applied to an already inflated attributed prior can still produce an implausible calibrated prior. Show the factor, the original prior, the calibrated prior, and the model posterior separately so the multiplication is visible.

Fourth, compare the result with other experiments and markets. A single outlier may be real, but it should not silently become the rule for an entire platform. Weight evidence by relevance, recency, spend, and statistical confidence, and keep contradictory tests visible.

Finally, check whether the model is being used outside its current measurement horizon. A short-term conversion-lift study cannot by itself validate a claim about brand effects several months later. Those effects require their own evidence and should not be smuggled into the calibration factor.

How this looks in practice

Consider a hypothetical ecommerce team with a conversion-lift study covering one market and a defined group of paid-social campaigns. The study reports iROAS of 4.0. Matched attribution for those same campaigns and dates reports 8.0, so the incrementality factor is 0.5.

Illustrative calibration audit showing how a conversion-lift iROAS becomes an MMM prior and posterior
Illustrative demo showing the full calibration chain from a scoped conversion-lift result to a broader MMM estimate. All figures are hypothetical.

The current attribution prior for the broader MMM feature is 10.0. Multiplying it by 0.5 gives a calibrated prior of 5.0. After the model considers the full sales history, seasonality, promotions, other media, and the response curve, its posterior iROAS is 5.3 with an illustrative interval of 4.5 to 5.5.

The MMM result of 5.3 does not reproduce the test result of 4.0, but the difference is explainable. The lift study first corrects the matched attributed ROAS downward from 8.0 through the factor of 0.5. The model then applies that factor to a broader prior and updates it with data from a wider period and feature. Each step and the uncertainty remain visible.

Now suppose the team discovers that the test combined traffic and awareness campaigns, while the 8.0 attribution denominator captured mostly traffic conversions. The factor of 0.5 may then be misleading for awareness. The correct response is to split the objectives where evidence supports it, pool sparse cells conservatively, or keep a wider uncertainty range. It is not to accept 5.3 simply because the model is calibrated.

Related questions

Does calibrating MMM mean it must reproduce the lift-study result?

No. The model should respect the experiment within the scope and uncertainty the study measured, but it may estimate a different result for a broader period, campaign group, KPI, or spend level. An exact match outside the test scope can create false precision.

Which result should marketers use for budgeting?

Use the experiment as the causal authority for the tested setup and period. Use the calibrated MMM for broader cross-channel and cross-period planning when the evidence has been mapped correctly and the model shows its uncertainty. For the full division of labor, see When should an omnichannel retailer use MMM, MTA, or incrementality testing?.

Can one conversion-lift study calibrate every campaign objective?

Usually not. Traffic, sales, engagement, video-view, and awareness campaigns can have different attribution visibility and incremental response. Calibrate at the most decision-useful level supported by enough spend and experimental evidence, then pool sparse cells conservatively instead of applying one precise-looking platform factor.

What if the experiment itself is uncertain?

Keep the confidence interval and test quality in the calibration. A low-confidence or poorly executed study should have less influence, and in some cases should be excluded. The model should not turn a noisy point estimate into a confident channel-wide rule.

How Sellforte helps

Sellforte connects conversion-lift, geo-lift, and other experimental evidence to the exact MMM features they inform. Teams can inspect the calibration basis, compare experiment and attribution results, keep uncertainty visible, and use the calibrated model for broader budget decisions. Book a demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.