How should geo-lift and conversion-lift studies be combined when calibrating MMM?

7 min read
Aug 21, 2026

Context: A senior marketing analytics leader at a large ecommerce company was adding geo-lift results to an experiment library already rich in conversion-lift studies and needed to decide whether the two methods should be treated as equal evidence in MMM calibration.

The short answer

Combine geo-lift and conversion-lift studies only after translating them to the same MMM feature, KPI, period, and uncertainty scale. When both cover the same decision, weight distributions by relevance and precision, and investigate disagreement instead of averaging point estimates.

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series →

Why do marketers ask this?

The question usually appears after an experiment program matures. A team may have dozens or hundreds of conversion-lift cells for a large digital platform, then run its first geo-lift study on the same channel. Both aim to estimate causal lift, but they may report different incremental ROAS figures.

The obvious responses are also the risky ones: average the results, trust the study with the narrower confidence interval, or decide that one method always outranks the other. None of those rules checks whether the studies measured the same campaigns, outcome, sales channel, spend level, or time horizon.

The two methods often answer related but different questions. A conversion-lift study can isolate a platform's selected campaigns at user level and provide repeated, granular evidence. A geo-lift study can measure total business outcomes across ecommerce and stores, cover media that cannot support user-level holdouts, and reveal effects outside a platform's conversion view. MMM calibration needs both kinds of evidence, with their boundaries intact.

What does each type of lift study contribute?

Conversion-lift studies are often most useful for campaign-level depth, while geo-lift studies add coverage at the business level. The practical differences matter more than the method labels.

Calibration question Conversion-lift evidence Geo-lift evidence
What is held out? Eligible users within an ad platform Regions, postal areas, or markets
What can it isolate? Selected test cells on the platform, such as campaigns, objectives, audiences, or  One channel, a channel group, or a broader marketing change that can be varied geographically
Which KPI is observed? The conversion or revenue definition available to the platform The business KPI supplied for the analysis, often the same sales or profit measure used by MMM
Main calibration strength Granular and repeatable platform evidence Business-level coverage, including offline outcomes and media without user-level holdouts
Main limitation Platform eligibility, tracking, conversion definitions, and post-view windows may not match the MMM The result may be less granular, harder to power, and exposed to geographic spillover or simultaneous regional changes

Geo-lift does not automatically deserve more weight because it is broader. A poorly powered geo test can be less useful than a precise conversion-lift study that matches the MMM feature closely. Nor should a narrow platform result automatically win because it is randomized at user level. The better evidence is the evidence that best matches the calibration decision and retains credible uncertainty.

How should the studies be made comparable before calibration?

Put every study into one evidence table, but do not blend the estimates yet. First record the study type, market, platform, campaign objective, participating campaigns, dates, spend change, KPI, sales channels, attribution or observation window, point estimate, uncertainty interval, and any overlap with other tests.

Experiment library combining geo-lift and conversion-lift evidence for MMM calibrationAn experiment library keeps study type, market, channel, dates, result, and confidence visible before the evidence enters MMM calibration. Demo data

Then map each result to the MMM feature it can actually inform. A conversion-lift study for a set of Meta sales campaigns should not silently calibrate all paid social. A geo test that turns off Meta and TikTok together should not be presented as direct evidence for either platform on its own.

KPI alignment is the next gate. If the MMM models net omnichannel sales but the conversion-lift study reports platform-observed ecommerce revenue, the two iROAS figures are not directly poolable. The team needs an explicit bridge between the definitions or must use the results at different levels of the calibration.

One useful bridge for platform evidence is an incrementality factor:

incrementality factor = conversion-lift iROAS / attributed ROAS for the same campaigns and dates

That factor can correct the matching attributed signal used to form a granular MMM prior. A geo-lift result measured on the MMM's business KPI can then serve as a broader constraint or validation point. This keeps the platform study close to the campaigns it observed and the geo study close to the total business effect it measured.

When should the two results be pooled, and when should they stay separate?

Pool the studies only when they address materially the same market, campaigns or model feature, KPI, sales channels, spend range, and effect window. Even then, combine probability distributions rather than averaging point estimates. Give more influence to the result that is more relevant and precise, and reduce the weight of stale, weakly matched, or poorly executed evidence.

Keep the studies on separate calibration layers when their scopes differ. In a hierarchical setup, conversion-lift studies can inform granular campaign or objective priors, while a geo-lift study constrains the combined channel or platform total. The granular priors should add up to a total that remains plausible relative to the geo result and its uncertainty.

Check independence before combining anything. If a geo-lift study and a conversion-lift study used the same campaigns during the same period, they are two readings of one marketing intervention. Counting both as independent evidence would make the prior look more certain than the underlying evidence allows. Record the overlap and model the dependence, or use one study as validation for the other rather than adding both weights in full.

What should you do when geo-lift and conversion-lift disagree?

Treat disagreement as an audit trigger. Do not settle it with a method-wide rule or a simple average.

Start with the outcome. Check gross versus net revenue, ecommerce versus omnichannel sales, conversion event, returns, new versus returning customers, and the post-treatment window. A geo-lift study using the same KPI as the MMM may have higher relevance, but only for the scope it actually tested.

Next, inspect the treatment. Confirm campaign IDs, objectives, audience eligibility, geographic exclusions, spillover, concurrent campaigns, promotions, spend levels, and whether either control group was contaminated. Compare the complete uncertainty distributions, not only the two central estimates.

If the disagreement survives those checks, keep it visible. Use a wider prior, retain separate feature-level estimates, and run sensitivity analyses with each study. The next experiment should be designed to resolve the specific uncertainty. A matched geo-lift can test whether a conversion-lift program has systematic bias, but one paired study should not automatically validate or invalidate every historical platform test.

How this looks in practice

Consider a hypothetical retailer with two new results for paid social. The figures in this written example are deliberately rounded placeholders and do not reproduce any study result from the source transcripts. A conversion-lift study covered Meta sales campaigns and reported a platform-observed ecommerce iROAS of 5, with an illustrative 90% interval of 4 to 6. Matched platform attribution for the same campaigns and dates reported a ROAS of 10, giving an incrementality factor of 0.5.

A geo-lift study covered Meta and TikTok together and measured net omnichannel sales. It reported an iROAS of 6, with an illustrative 90% interval of 5 to 7.

Evidence What it calibrates What it does not prove
Conversion lift
iROAS 5, attributed ROAS 10, factor 0.5
The prior for matching Meta sales campaigns, after applying the factor to a compatible attributed signal Total paid-social return on net omnichannel sales
Geo lift
iROAS 6 on net omnichannel sales
A broader constraint on the combined Meta and TikTok response for the tested market and spend range The separate return of Meta, TikTok, or individual campaign objectives
Calibrated MMM Granular feature estimates whose combined paid-social result remains plausible relative to the geo-lift distribution A claim that the two experiment point estimates measured the same quantity

The team should not average the two headline iROAS figures into a single prior. The figures use different revenue definitions and channel scopes. The conversion-lift result informs the Meta sales component. The geo-lift result checks whether the combined paid-social estimate is credible at the broader business level.

If the calibrated Meta and TikTok components imply a paid-social iROAS inside the geo-lift interval, the two evidence layers are coherent. If the combined model result sits far outside that range, the team has a concrete calibration problem to investigate.

Related questions

Should geo-lift always receive more weight than conversion lift?

No. Weight follows relevance, precision, recency, and execution quality, not the study label. Geo-lift often matches the MMM's business KPI more closely, while conversion lift may match a specific platform feature more closely and provide a more precise estimate.

Can one geo-lift study validate a library of conversion-lift studies?

It can test for systematic differences when the geo and conversion studies cover comparable campaigns, dates, outcomes, and spend. One matched result is useful evidence, but a broader validation claim needs repeated comparisons across settings. Preserve outliers instead of declaring the entire library valid or invalid.

What if geo-lift and conversion lift ran at the same time?

Treat the results as correlated when they observed the same intervention. Do not give both full independent weight. Use one as the calibration input and the other as validation, or represent their dependence explicitly when forming the prior.

How should several studies for the same feature be weighted?

Map each test to the feature first, then weight it by relevance, recency, statistical confidence, and spend. Exclude invalid studies and retain wider uncertainty for valid but imprecise ones. See How should we weight multiple incrementality tests when calibrating an MMM?.

How Sellforte helps

Sellforte brings geo-lift, conversion-lift, and other experimental evidence into one library and connects each result to the MMM features it can inform. Teams can inspect scope, KPI, overlap, uncertainty, and prior-to-posterior movement before using the calibrated model for budget decisions. Book a demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.