What is the best incrementality test?

5 min read
Published Sep 9, 2026
Updated

Context: A senior marketing leader at a large retail company evaluating measurement options needed to understand which experiment types would cover the company's different marketing channels.

The short answer

There is no single best universal test type. Start with Meta Conversion Lift for Meta, geo-lift for Google Search, Shopping, and Performance Max, and customer holdout A/B tests for email, push notifications, and direct mail. A marketing portfolio needs several test types, with platform eligibility checked before planning.

Here's a summary video from Sellforte CEO, Juha Nuutinen:

This article is part of Asked by Marketers, a series answering real questions from marketing leaders.

Sellforte's team holds more than 1,450 meetings each year with marketing leaders in Ecommerce and Retail about Marketing Mix Modeling and incrementality testing. Each week, we anonymize at least one question from those conversations and answer it in depth, based on what marketers are actually struggling with, not what keyword tools suggest. About the series

Why marketers ask this

Choosing one testing method would make life easier. A marketing team could use the same setup across its budget and compare results in a familiar format. But the team has very different control over a Meta ad, a Google Search campaign, and an email sent to its own customers. That changes how it can create a credible group that does not receive the marketing.

Incrementality testing estimates the additional sales or other outcomes caused by a marketing activity. Conversion lift, geo-lift, and customer holdouts approach that question through different experimental units. The useful starting point is the channel and the decision you need the result to support.

Which incrementality test should you use for Meta?

For Meta, our starting recommendation is Meta Conversion Lift. Its combination of ease, speed, granular reporting, and a randomized experimental design makes it particularly useful for repeated testing.

Meta manages the split between a test group eligible to receive the campaigns being measured and a control group held back from those campaigns. The study compares conversion outcomes across the groups to estimate the additional conversions caused by the advertising. This is a different measurement from the conversions Meta attributes to an ad after a click or view.

The practical advantage is repeatability. A team can plan successive Conversion Lift studies around its campaign questions without organizing a regional blackout each time. More granular questions still need enough data, so splitting a study into smaller audiences or campaign groups does not guarantee a useful answer.

Meta controls audience assignment, which limits what advertisers can inspect independently. A well-matched geo-lift study can provide a separate check. Our article on whether you can trust Meta Conversion Lift tests explains that distinction in more detail.

Which incrementality test should you use for Google?

Geo-lift is a practical starting point for testing Google Search, Shopping, and Performance Max. You change the selected advertising in treatment regions and use untreated regions to estimate what would have happened without that change.

Google's user-based Conversion Lift availability needs a closer check. Its documentation lists Video, Discovery, and Demand Gen campaigns, and now directs advertisers to their account representative about user-based studies for Search, Shopping, Performance Max, App, and Display. The Performance Max route excludes Store Goals. Access therefore varies by campaign type and account. Check the options available to your account before choosing the design. Google's comparison of lift types explains the current access routes.

Geo-lift also lets you work with your own regional sales data. That is useful when the question concerns a business outcome that a platform's conversion tracking does not fully capture.

How should you test email, push notifications, and print?

Use a customer holdout A/B test when you can directly control who receives the marketing. Randomly assign eligible customers to receive the activity or to a control group that does not receive it, then compare their outcomes over the agreed measurement period.

For email, this means comparing a send with no send. Testing two subject lines answers which email performs better; it does not establish how much additional revenue sending the email creates. The same distinction applies to push notifications.

For print, this approach fits addressable activity such as catalogs, postcards, and direct mail. You know which customers would receive the mailing and can withhold it from the control. Print distributed without a customer list may need a geographic design instead.

The holdout must also match the scope of the question. If you want to measure a recurring print program over a season, a persistent no-print group can capture the cumulative effect across sends. If you want to measure one mailing, define the eligible audience and comparison for that mailing. Customers held out of one catalog but sent another are not a no-print control for the whole program.

What else should determine your choice of test?

The test must measure the outcome you will act on, with enough precision to inform that decision. Channel fit gets you to a viable method; it does not guarantee that the study will answer your question.

Agree on the conversion or revenue definition before launch. A platform-tracked purchase and an order in your backend may use different definitions or cover different customers. If you use geo-lift, check that the regional data actually covers the outcome you want to measure. Missing geography can limit what the result represents.

Next, assess whether the expected effect, available audience, and test duration can produce useful evidence. Holdout size should follow that assessment and the commercial cost of withholding marketing. There is no percentage that works for every channel and market.

Review the uncertainty around the result as well as the headline lift. A test can leave the exact effect uncertain while still helping you assess whether a channel is likely to meet your profitability requirement. Our guide to judging whether an incrementality test is credible covers the checks to make before using the result.

How this looks in practice

Consider a hypothetical ecommerce team using Meta, Google, email, push notifications, and direct mail. Assume it can withhold owned-channel messages from selected customers and has usable regional sales data for geo-lift. Its initial testing plan could look like this:

Marketing activity Starting test What to settle before launch
Meta campaigns Meta Conversion Lift Campaign scope, conversion definition, eligibility, and power
Google Search, Shopping, and PMax Geo-lift Regional data and treatment design; check user-based lift access too
Google Demand Gen or Video User-based Conversion Lift where available Account eligibility and the outcome the study can measure
Email Customer holdout A/B test Send versus no-send assignment across eligible customers
Push notifications Customer holdout A/B test Which notifications the control will be excluded from
Recurring direct mail Customer holdout A/B test One mailing or the whole program, plus the measurement period

Write a test type next to every channel taking a meaningful share of budget. A blank identifies a gap in your testing plan. Then record whether each channel already has relevant completed evidence: assigning a future test does not mean the channel's incrementality has been measured.

Build the schedule around the operational differences. Meta studies may fit a monthly planning cycle, while regional spend changes and print holdouts can require earlier coordination. For each planned test, write down the budget decision its result will inform and who will make that decision.

How Sellforte helps

Sellforte brings Conversion Lift, Geo Lift, and A/B test results into one Experiments Hub, where teams can review incremental ROAS and uncertainty across their experiments. See how Experiments Hub works or explore the Sellforte demo.

Authors

Lauri Potka

Lauri Potka is the Chief Operating Officer at Sellforte and has over 15 years of experience in Marketing Mix Modeling, marketing measurement, and media spend optimization. Before joining Sellforte, he worked as a management consultant at the Boston Consulting Group, advising some of the world's largest advertisers on data-driven marketing optimization. Follow Lauri on LinkedIn, where he is one of the leading voices in MMM and marketing measurement.