After years auditing paid-media programs, the question I get more often than any other isn’t “which attribution model should we use?” It’s a quieter, more frustrated version of the same problem: “We doubled the Meta budget and sales barely moved. Did the ads even do anything?”
That question has a name — and an answer. Incrementality testing is the discipline of measuring what would have happened without a given piece of marketing. Not what the platform reported. Not what attribution assigned. What actually changed in the world because you ran that campaign. It is, by most measures, the most honest thing you can do with a marketing budget.
This guide covers how incrementality tests work, when geo holdouts make more sense than conversion lift studies, the math behind reading results correctly, and where the approach falls short. If you’ve been wondering why your MMM says one thing and your platform dashboards say another, incrementality testing is the bridge.
What Incrementality Actually Means
Platform attribution assigns credit to touchpoints. Incrementality measures cause. The difference sounds philosophical, but it shapes every budget decision you make.
Consider branded search. A typical account might report ROAS of 12–20x on brand keywords. If you pause brand search entirely, some percentage of those clicks will still arrive via organic — the user would have searched, found your organic listing, and converted anyway. The ad spend didn’t create the conversion. It just claimed the credit. The incremental ROAS on branded search can be anywhere from 0.3 to 2.0, depending on the brand’s organic strength and competitive landscape.
Or consider a retargeting campaign hitting users who already added to cart. Last-click ROAS looks stellar. But these users had high purchase intent before you showed them the ad — a large share would have returned and converted regardless. The ad spend may have accelerated the purchase by a day, but it didn’t create it.
The formal definition from the measurement literature: a channel or campaign is incremental to the degree that removing it would reduce the outcome. Incrementality testing creates the conditions to measure that counterfactual directly, rather than inferring it from attribution models that were never designed to answer causal questions.
Two Main Test Designs
There are two dominant approaches. Which one fits depends on your traffic volume, geography, and what you’re trying to measure.
Conversion Lift Studies (User-Level)
Conversion lift studies work by randomly splitting users into a test group (sees the ads) and a holdout group (sees PSA ads or nothing) within the same campaign. The platform records conversions in both groups and computes the lift. Meta’s Conversion Lift tool and TikTok’s conversion-lift A/B test feature both operate on this principle; Google runs lift primarily as geo-based measurement rather than a user-split tool.
The advantage: high statistical power at lower traffic volumes than geo tests, because you’re splitting users within a market rather than across markets. A 5% holdout on a campaign reaching 500,000 users gives you 25,000 people in the control group — plenty for most conversion events.
The limitation: you’re entirely inside the platform’s ecosystem. Meta runs the randomization, Meta measures the outcome, and Meta reports the results. The test is methodologically sound — holdout randomization is genuinely better than attribution — but you’re trusting a platform with a financial incentive to show high lift to be the neutral judge of its own performance. That’s not a fatal problem, but it’s worth naming. Cross-platform conversion lift tests (where you can independently verify conversions in GA4 or your own data) are materially more credible.
Geo Holdout Tests
Geo tests work differently. You identify matched pairs or groups of geographic areas — cities, metro areas, postal codes, or entire regions — and run the campaign in some geos while going dark or reducing spend in others. You then compare conversion rates between the test and control geos.
The advantage: platform-independent. You measure the outcome in your own analytics, your own CRM, your own revenue data. No platform can adjust the numbers. You can also test the effect of going entirely dark (true incrementality) or of specific creative strategies, channels, or spend levels.
The limitation: geographic confounds. Markets aren’t identical. A test geo might have a different baseline conversion rate, a local competitor promotion during the test window, unusual weather, or a PR event that affects sentiment. The more different your test and control geos are to begin with, the noisier the results. This is why geo selection — and pre-test equivalence checking — matters so much.
| Design | Split unit | Who runs it | Typical minimum | Best for |
|---|---|---|---|---|
| Conversion lift (platform) | User | Platform (Meta / Google) | ~20K users/week in campaign | Single-channel channel health check |
| Geo holdout | Geography | You (independent) | ~8–10 matched geo pairs | Cross-channel and budget-level tests |
| Ghost ads holdout | User (independent) | Third-party (Measured, Northbeam) | ~50K impressions/week | Independent cross-channel measurement |
Ghost ads tests sit in between: a third-party tool (Measured.com, Northbeam, Triple Whale’s lift product) randomizes users independently of the platform, serves “ghost ads” to the holdout group that register as impressions but aren’t real ads, and measures conversion delta in first-party data. The methodological independence is better than platform lift; the cost is higher than a DIY geo test.

Running a Geo Holdout: The Mechanics
A basic geo test doesn’t require a vendor. Here’s how to run one without the enterprise price tag.
Define the question precisely. “Does our Meta prospecting drive incremental revenue?” is testable. “Is our marketing working?” is not. Pick one channel or one spend level and one conversion event.
Select and match geos. Pull 6–12 months of weekly conversions by geography from your analytics platform. Group geos by conversion volume, baseline conversion rate, and demographic composition. You want pairs that look similar before the test starts — not identical (impossible) but statistically comparable. Tools like Google’s open-source Meridian GeoX, Meta’s open-source GeoLift, or even a basic correlation analysis in a spreadsheet can identify good pairs. Aim for at least 4–5 test geos and 4–5 matched control geos.
Pre-test calibration period. Before the test starts, run a synthetic test on historical data. Pretend you had randomly paused campaigns in your chosen control geos 8 weeks ago. Does the model predict actual performance? If it’s off by more than 10–15%, your geo matching needs work.
Run the test. The minimum recommended duration is 4 weeks; 6–8 weeks is better for purchase-cycle products with longer consideration windows. During the test: don’t change any other campaigns, don’t run promos that only cover some geos, don’t change prices. Any of these introduces confounds that contaminate the result.
Measure the lift. Compare test and control geos on the metric you defined in step 1. The basic calculation:
Incremental conversions = (Test geo conversion rate − Control geo conversion rate) × Test geo traffic Incremental ROAS = (Incremental conversions × Average order value) ÷ Spend in test geos
If the conversion rate in your test geos doesn’t move relative to your control geos, you have low or zero incrementality in that channel. That’s valuable information, even if it’s uncomfortable.
Account for uncertainty. Geo test results are noisy. A 15% lift in a single 4-week test with 6 geo pairs isn’t a hard number — it’s an estimate with a confidence interval. Report the range, not the point estimate. And replicate before making permanent budget cuts. One inconclusive geo test is not evidence; it’s a first signal.
Reading Incremental ROAS Correctly
The key output of any incrementality test is incremental ROAS (iROAS), sometimes called true ROAS. It’s the number that deserves to sit in a budget allocation decision — not platform-reported ROAS.
How to interpret what you find:
- iROAS well above your platform ROAS: Rare for prospecting. More common in upper-funnel brand campaigns where attribution systematically under-credits the channel because it rarely closes the last click. If your display iROAS exceeds what GA4 shows, the channel is doing more than your reporting reveals.
- iROAS roughly equal to platform ROAS: Your attribution is relatively accurate for this channel, or the test is underpowered and the signal is too noisy to detect a gap. Run a longer or larger test before concluding accuracy.
- iROAS well below platform ROAS: The normal finding for retargeting, branded search, and channels that over-credit intent-rich users who would have converted anyway. Typical finding: retargeting iROAS lands well below reported ROAS — often around 0.4–1.2 against a reported 3–6 — though the exact figures vary widely by brand and audience. The channel is doing something — accelerating conversions, perhaps — but not what the dashboard claims.
- iROAS below 1.0: The channel is spending more per incremental conversion than a conversion is worth. You’re paying for purchases you would have gotten for free, or you’re stealing conversions from a future time window and calling it current-period revenue. These campaigns should be restructured before scaling.
A practical frame I use with clients: run a three-tier test. First test prospecting (most likely to show positive incrementality, usually under-credited by last-click). Then test retargeting (most likely to show low incrementality, usually over-credited). Then test brand search (depends entirely on organic strength). The results in those three tests will tell you more about your actual marketing efficiency than a year of attribution modeling.

Incrementality vs. Attribution: When to Use Which
Attribution and incrementality answer different questions. Using the right tool for the right question prevents most of the “our numbers don’t make sense” conversations.
Attribution (last-click, data-driven, position-based) answers: Which touchpoints were present in the paths that led to conversion? It’s a descriptive answer about correlation. It tells you where to look, not what caused what. Attribution is useful for tactical in-flight decisions — which ad set is seeing more engagement, which creative variant is getting clicks — because it’s near-real-time and granular.
Incrementality answers: What would have happened without this campaign? It’s a causal answer. It’s slower (tests take weeks), less granular (geo-level or user-cohort-level, not individual ad), and expensive to run at scale. But for the question of where to put next quarter’s budget, it’s the right tool.
Marketing mix modeling sits in the same causal family as incrementality testing. The key difference: MMM is retrospective (it models past spend to estimate contribution), while incrementality tests are prospective experiments (they manipulate spend to observe effect). MMM and incrementality testing are complementary: run the holdout tests to calibrate the MMM, and use the MMM to extrapolate beyond the test conditions.
If you only have capacity for one measurement investment, start with incrementality testing. It produces a number (iROAS) that directly informs budget decisions and is harder to misinterpret than a regression coefficient.
Common Mistakes That Invalidate Tests
Incrementality tests are not self-executing. The most common failure modes:
- Testing a period with unusual external events. Running your holdout over a major sale event, a product launch, a PR moment, or a seasonal outlier poisons the comparison. Your test and control geos probably don’t respond to those events identically. Run in a clean, average-ish period if possible.
- Too-short test duration. Two weeks is rarely enough for anything with a purchase cycle longer than a few days. For B2B, subscription, or high-consideration products, run for at least 6 weeks. Catch the full consideration-to-conversion loop.
- Insufficient holdout size. A 2% holdout on a campaign that generates 50 conversions per week means 1 conversion in the holdout group. Statistical noise will swamp any real signal. Rule of thumb: target at least 30–50 conversions per week in the holdout arm, minimum.
- Changing the campaign during the test. Restarting a paused ad set, changing creative, shifting audiences, or adjusting bids during the test period contaminates the result. Treat the test window as a controlled environment: no changes until it’s done.
- Geographic spill. If your test and control geos are adjacent (e.g., neighboring cities), brand advertising in the test geo will spill into the control geo through social sharing, word of mouth, and national media. This compresses the measured lift and makes effective campaigns look ineffective. Use geographically distinct markets — not cities in the same metro area.
- Cherry-picking the lookback window. If you run a 6-week test and only report the 2 weeks where lift looked best, that’s not a test result, it’s a selection artifact. Pre-register the test window and report it in full.
- Ignoring statistical significance. Incrementality results need confidence intervals. A reported 12% lift with a 95% confidence interval of [−8%, +32%] is not evidence of lift — the zero is inside the range. Use a proper test for the difference in proportions or conversion rates, and don’t declare victory on a wide interval.

Where Incremental ROAS Belongs in Your Bid Strategy
The practical outcome of incrementality testing is a set of per-channel iROAS estimates that you use to set bid targets. This is the bridge between measurement and media buying that most teams are missing.
The framework I use:
- Run holdout tests on each major channel once per year, minimum. Incrementality shifts as creative ages, as competitors adjust, and as audience saturation builds. An iROAS estimate from 18 months ago may not reflect today’s channel dynamics.
- Set bid targets using iROAS, not platform ROAS. If your prospecting campaign shows an iROAS of 2.1 and you need a 2.0 to break even, you’re bidding correctly. If you’re bidding to a platform-reported ROAS of 4.5 that reflects an iROAS of 1.3, you’re systematically underbidding.
- Apply a correction factor to platform ROAS targets. For channels with consistent measurement: correction factor = platform ROAS ÷ iROAS. If platform ROAS is typically 3.8 and your iROAS test showed 2.1, the factor is 1.8. Apply it when setting bid targets to translate iROAS goals into the language the platform algorithm understands.
- Segment retargeting audiences by incrementality profile. Not all retargeting is equally non-incremental. High-intent cart abandoners often show better iROAS than broad retargeting pools that include people who bounced from the homepage after two seconds. Testing sub-audiences within retargeting is a high-leverage use of lift methodology.
This connects directly to the ROAS vs. MER vs. CAC framework: MER anchors you to business reality, iROAS tells you which channels are earning their share of that MER, and CAC tracks whether you’re acquiring new customers or just re-converting existing ones.
When Incrementality Testing Doesn’t Apply
It would be clean if incrementality testing solved every measurement problem. It doesn’t. Honest accounting of where it falls short:
- Small conversion volumes. If you’re generating fewer than 100 conversions per week across your entire program, most holdout designs won’t have enough statistical power to detect realistic lift. MMM with good priors may be more useful than a test you can’t power adequately.
- Highly localized businesses. If your customer base is entirely in one city or one region, you can’t construct a credible geographic control that isn’t also exposed to your marketing (brand awareness, PR, word of mouth).
- Testing long-cycle B2B sales. If your average sales cycle is 6 months, a 4-week geo test can’t catch conversion outcomes — you’d need to run the test for 6+ months plus an observation window, which is logistically complex and expensive.
- One-channel businesses. If you spend entirely on Google Search, you can run a geo holdout, but the “what happens without us?” scenario is partially answered by organic — and disentangling paid-search incrementality from SEO health requires additional instrumentation (tracking branded vs. non-branded organic separately during the test window).
- Measurement is not a substitute for strategy. Incrementality testing tells you whether a channel is working, not why, and not what to do instead. A low-iROAS retargeting campaign might need creative work, audience refinement, or a recency filter — not just a budget cut.
FAQ
What’s the difference between a geo holdout and a conversion lift study?
How many geos do I need for a valid geo test?
What’s a typical incrementality result for retargeting?
How does incrementality testing relate to marketing mix modeling?
Do I need a vendor tool to run incrementality tests?
My geo holdout showed near-zero lift. Should I cut the channel?
Can I run incrementality tests on organic channels like SEO or email?
Attribution models are navigation. Incrementality testing is calibration. You need the first to run campaigns. You need the second to know whether the campaigns are actually working.
The teams that get this right run holdout tests on their top two or three spend channels at least once a year, use iROAS (not platform ROAS) to set bid targets, and treat every retargeting and branded-search campaign with healthy skepticism until the numbers hold up under a causal lens. The teams that skip it end up in the meeting I described at the top: budget doubled, results flat, nobody sure what happened.
Start with a geo holdout on your largest paid prospecting channel. Keep the test clean, run it long enough to capture your full purchase cycle, and measure the outcome in your own data. The result — however uncomfortable — is worth more than six months of platform-reported ROAS.
For the measurement layer that sits above individual channel tests, the marketing mix modeling guide covers how to combine holdout results into a portfolio-level view of spend efficiency. And if your attribution setup isn’t solid enough to detect a clean lift signal in the first place, the tracking plan guide covers the data hygiene work that has to come first.
