Incrementality Testing
What Is Incrementality Testing?
A plain-language guide to incrementality testing: what it measures, how holdout and geo tests work, and how it grounds your MMM in real causal evidence.

TL;DR
- Incrementality testing measures the overall causal impact of a media channel: the revenue that would not have happened without the channel.
- It uses a test group (exposed to media) and a control group (withheld from media) to isolate true contribution from baseline activity.
- The core formula: (Test Conversion Rate – Control Conversion Rate) / Test Conversion Rate = Incrementality %.
- Unlike platform attribution or MTA, incrementality is unaffected by signal loss, walled-garden bias, or cookie deprecation, making it the foundation of modern, durable marketing measurement.
- Leading brands use incrementality results to calibrate Marketing Mix Models (MMM), validate channel performance, and reallocate budget with confidence.
What Is Incrementality Testing?
Without incrementality testing, everything we know about marketing channels is correlation or anecdote. Incrementality testing is a controlled, randomized experiment that measures the true incremental lift of a channel: the revenue that would not have happened without it.
Why Should You Do Incrementality Testing?
For e-commerce and retail brands, uncalibrated observational models systematically over-spend on channels that look good in-platform but aren’t incremental: branded search, retargeting, promotions that mostly reward customers already about to buy.
As we’ve discussed before, ad platforms systematically overstate the reported ROAS.
As a documented example, a published analysis of Dropbox’s incrementality testing program found that click-based ad attribution meaningfully overstated the actual impact of its paid channels. Acting on the incrementality results let Dropbox redirect budget away from low-incrementality spend toward what actually moved the business.
Nielsen’s own research points the same direction: marketers keep shifting budget toward performance channels at the expense of brand building, which Nielsen warns creates a “vicious circle” of spending more to convert an already-shrinking pool of interested customers.
Where E-commerce Brands Should Test First
Branded search, Retargeting, Promotions and Affiliate programs are top candidates for running incrementality testing. More specifically:
- Branded search. Someone searching your brand name is often already planning to buy; a holdout usually shows only a small fraction of that traffic is genuinely incremental.
- Retargeting and cart-abandonment ads. These chase shoppers who already browsed or added to cart, many of whom would have returned to check out on their own.
- Promotions and discount codes. A well-timed promo can pull forward demand that would have happened anyway at full price a few weeks later, rather than creating a new sale.
- Affiliate, marketplace, and last-click channels. These sit closest to the point of purchase and tend to claim credit for a customer’s entire journey, not just their own contribution.
This cluster is also the fastest place to start a measurement program, since it self-funds: budget freed from low-incrementality spend is what feeds the upper-funnel prospecting.
E-commerce brands with high order volume and short purchase cycles are also well suited to holdout and geo tests, which need a steady stream of conversions to reach statistical confidence quickly.
How Does Incrementality Testing Work?
Incrementality testing measures the overall causal impact of a channel by exposing two comparable groups of people to two different media treatments. The test group is exposed to the channel as usual. The channel is turned off for the other group, the control group (also called the holdout).
An example of incrementality testing
Say a brand runs a four-week holdout test on retargeting. The test group, still exposed to retargeting, converts at 3.2%. The control group, with retargeting withheld, converts at 2.6%.
Incrementality % = (3.2% – 2.6%) / 3.2% = 18.75%
Roughly 19% of the test group’s conversions are incremental, caused by retargeting. The other 81% would have converted anyway.
If retargeting’s in-platform ROAS looked like 8x, the actual incremental ROAS is closer to 1.5x, a very different number to plan a budget around.
What is the difference between A/B testing and Incrementality Testing?
Incrementality testing is a form of randomized controlled experiment, similar to A/B testing. But while A/B testing is for testing features of a site or aspects of a campaign, incrementality testing is about the bigger strategic decisions around budgets and the overall marketing and media mix.
What Are the Different Types of Incrementality Testing?
There are two major types of incrementality testing: holdout experiments and scale experiments.
Holdout Experiments. A holdout experiment measures the causal impact of a marketing channel by completely switching the channel off for the control group and comparing business outcomes to the test group, which continues receiving it as usual.
Scale Experiments. When switching off a channel is too expensive or risky, a scale experiment provides the solution. Instead of switching off a channel, a scale test increases media dosage (2x–4x) for the test group. Unlike holdout experiments, which measure the overall impact of a channel, scale experiments measure the marginal ROAS (mROAS) of the channel. mROAS determines the incremental return of the next dollar (or unit of spent).
What Are the Different Ways of Implementing Incrementality Testing?
There are two main ways of implementing incrementality testing: User-level RCTs / Ghost Ads and Geo-Testing. User-level RCTs / ghost ads, is where individual users are withheld from ad exposure, and geo-testing / synthetic controls, is where geographic markets are held out and modeled against a synthetic baseline.
User-Level Holdout Groups
User-level Randomized Controlled Trials (RCTs) is where individual customers are randomly assigned to an exposed test group or an unexposed holdout group. This test provides a causal impact of the treatment.
What is holdout contamination?
In user-level randomized controlled experiments, cookie expiration, cross-device switching, and mobile privacy restrictions let holdout users get exposed through a secondary device, a problem known as holdout contamination.
How does Ghost Ad work?
A Ghost Ad is a user-level holdout test that runs inside a platform like Meta or Google. When an ad auction occurs for a user assigned to the holdout, the platform suppresses the ad, serves the next-highest bidder instead, and records a “ghost impression” to track who would have been exposed. Marketers then compare the behavior of the test group with the holdout to determine true incrementality.
Ghost ads are precise within a single platform, but can’t account for cross-platform interaction effects, media cannibalization, or offline conversions without external integration.
How Does GeoLift for Incrementality Testing Work?
By partitioning geographic regions into non-overlapping areas known as “geos,” GeoLift studies measure the true causal, by exposing one geo to the media treatment while withholding another geo from the media exposure. This provides a way to measure of incremental impact of a media channel.
Geo-testing is particularly useful for offline channel measurement, as we’ve discussed before.
In the United States, geos are commonly designated as Designated Marketing Areas (DMAs), of which there are 210.
Switzerland and the EU don’t have a DMA equivalent, but the same logic works using cantons, postal-code regions, or metro areas like Zurich, Geneva, and Basel.
Because geo-testing measures aggregate, non-PII outcomes rather than tracking individuals, it fits naturally with Switzerland’s FADP and the EU’s GDPR: a rigorous read on channel performance without touching customer-level data. When designing a geo test, account for seasonality and each region’s baseline market penetration.
Interrupted Time Series (ITS)
Interrupted Time Series (ITS) is a causal inference method used when a standard control group isn’t possible. A clear change happens at a specific point in time, launching a national campaign, turning off a paid channel, and a model predicts what would have happened based on pre-intervention trends and seasonality. ITS is best for macro-level interventions, like nationwide campaigns, that can’t be split by geography or user ID.
Define Incremental ROAS (iROAS)
Incremental Return on Ad Spend (iROAS) is a metric that measures the net-new, causally driven revenue generated per dollar of marketing spend on a given channel. Incrementality testing is the gold standard for experimentally measuring the iROAS.
Unlike standard platform-reported ROAS, which divides total conversions touching an ad by media spend, iROAS subtracts the counterfactual baseline, the volume of sales that would have occurred naturally through organic search and brand equity.
How Marketers Use Incrementality Testing
Incrementality testing is performed alongside Marketing Mix Modeling (MMM) to ensure MMM insights are grounded in real-world causal evidence, not just correlations. This is called MMM calibration.
iROAS serves as an empirical ground truth: estimates from an incrementality experiment are injected into the Bayesian MMM to calibrate it. At ELIYA, we use iROAS as informative Bayesian priors to calibrate top-down MMM response curves, reviewed by our team rather than left to run unchecked, giving teams the most grounded, evidence-based basis for budget allocation and planning.
Attribution vs. Incrementality
Incrementality testing is a randomized controlled experiment that measures a channel’s true causal impact. Multi-Touch Attribution (MTA), by contrast, measures the contribution of each touchpoint to a sale, but doesn’t report causal lift.
Incrementality testing also doesn’t rely on cookie tracking, so it isn’t affected by signal loss the way MTA is: it proves a touchpoint adds new value, not just that it’s correlated with sales.
Can Incrementality testing prevents Lower-Funnel Death Spiral
Yes, Incrementality testing shows which channel genuinely creates new demand and which channels intercept it, so the budget allocation can be balanced by allocating across the entire marketing funnel.
What is the Lower-Funnel Death Spiral
Lower-Funnel Death Spiral occurs when lower-funnel channels absorb most of the media budget, increasingly undermining brand building and distorting baseline demand, leading to a long-term over reliance on performance channels.
Uncalibrated attribution doesn’t just misreport channel performance, left unchecked, it pulls a brand’s budget in the wrong direction over several stages: starting from observational attribution bias, to capital reallocation to lower-funnel channels and finally platform auction hyper-inflation.
Observational methods are dangerous even when accurate-looking: correlation-based models consistently overestimate purchase-driven conversion lift.
How Long, How Much, and How Disruptive Is a Test?
Reasonably, these are the first questions most CMOs ask before agreeing to change how a channel runs. How long: several weeks to a couple of months, depending on conversion volume and baseline variation. How much: usually just the reallocated media spend during the test window, not a large budget on top of it; the bigger cost is operational, holding a group back long enough to see a clean answer. How disruptive: a holdout run on a portion of traffic or geos is designed to be low-risk, and a well-designed test caps how much of the audience or budget is withheld.
FAQ
What is the difference between Holdout and Scale Experiments?
A holdout experiment switches a channel off for part of the audience to measure its overall lift; a scale experiment increases spend instead, to measure marginal ROI. Use a holdout to learn whether a channel matters at all, a scale test once you already know it works and want to know how far to push it.
Does incrementality testing require cookies or user-level tracking?
No. Geo-testing and Interrupted Time Series measure outcomes at the market or time-period level, not the individual, so they hold up as cookies and device IDs disappear.
How does incrementality testing relate to MMM?
They’re complementary. MMM models the whole mix from historical data; incrementality testing provides real experimental ground truth for individual channels, and feeding those results in as calibration priors keeps the model grounded in evidence, not just correlation.
Is incrementality testing only useful for large brands with big budgets?
No. Brands with lower spend or conversion volume often get cleaner results from geo-tests or scale tests, which need less volume to reach statistical confidence than a full user-level holdout does.
Where This Leaves Your Budget
Incrementality testing won’t tell you everything about your marketing mix on its own, but it will tell you which channels are creating growth and which are just taking credit for it. At ELIYA, we run incrementality testing and MMM calibration together, using agentic AI to move faster and at a fraction of the cost of a traditional agency, with our team reviewing every result before it shapes a budget decision.
Ready to find out where your budget is actually working? Talk to ELIYA about running an incrementality test for your brand. In a follow-up guide, we’ll cover how to decide between incrementality testing, MMM, or both.

