The Definition
A geo holdout test measures what your ad spend is likely to have caused by running the spend in some geographic regions and deliberately withholding it in others, then comparing business results between the two groups. When the regions are well matched and the main planned difference is the spend, the difference in outcomes becomes evidence of what the spend likely contributed.
It is one of the closest paid media methods to a controlled experiment, and it answers the question that attribution cannot: not "which sales touched our ads" but "which sales likely happened because of them". That question, and the cheaper rungs on the ladder below this one, are covered in our guide to whether your paid media spend is truly incremental. This piece covers the geo rung specifically: how it works, when it is valid, and when it is not worth running.
How a Geo Holdout Works
The mechanics are simple to describe and easy to get wrong.
- Split comparable regions into two groups. Test regions receive the spend; holdout regions do not. The groups should be matched on the things that drive your results: historical sales levels, seasonality patterns, market maturity. You are trying to build two versions of the same market.
- Run the spend difference for a fixed window. Keep the material conditions as stable as possible in both groups: same site, same pricing, same promotions, same other channels. One planned variable moves.
- Compare business outcomes, not platform metrics. Revenue, orders, or new customers per region, from your own records. The gap between test and holdout regions, adjusted for their historical baseline difference, is the measured lift.
An invented example for shape: a retailer picks six comparable city regions, runs the new channel in three, holds out three, and spends for eight weeks. Test regions grow revenue 6% over their baseline; holdout regions grow 4% over theirs. The 2-point gap, on the test regions' revenue base, is evidence of what the spend likely contributed, and it may tell a very different story from the platform's attributed ROAS, for the reasons explained in platform ROAS vs incremental ROAS.
The same logic has a sibling: the audience holdout, where a random slice of the target audience is withheld instead of a region. Geo versions can be practical when platform-side experiment tooling is not available, and blended regional revenue is measurable from your own systems.
When a Geo Holdout Is Valid
The test earns trust when a few conditions hold, and weak geo tests often break one of them:
- The regions are genuinely comparable. Matched on baseline levels and trends, not just picked for convenience. A test region with a stadium event or a local competitor closure mid-test can quietly break the comparison.
- The signal can clear the noise. Regional revenue moves week to week on its own. The expected lift needs to be large enough, and the window long enough, to read beyond that normal variance. As a design principle: if the spend is a small fraction of regional revenue, the lift it could plausibly produce may be smaller than the noise, and the test is unlikely to answer the question in a useful window.
- The groups stay clean. Geo targeting leaks: people travel, VPNs blur locations, and delivery systems spill impressions across borders. Some leakage is normal; wholesale contamination, such as retargeting pools or brand campaigns still running in holdout regions, can break the test.
- One change at a time. A price change, site migration, or promotion landing mid-test in either group makes the comparison harder to read. Boring test hygiene is what makes the result useful.
- The pass mark was written first. Like every test in the channel testing framework, the lift threshold, window, and decision date are agreed before launch, so the result decides rather than the meeting.
When the Sample Is Too Small
Geo holdouts are one of the stronger forms of evidence on the ladder, and one of the most demanding. Honest reasons not to run one:
- Not enough regions. With only one or two meaningful markets, there may be nothing comparable to hold out. Single-market businesses usually need a different measurement design.
- Not enough conversions. If a region produces a handful of sales a week, the difference the test needs to detect drowns in randomness. The test can look scientific while still giving a noisy answer.
- Not enough patience. Readable geo tests usually need weeks, not days. If the decision deadline arrives before the signal can, run a cheaper rung instead.
None of this means smaller advertisers cannot measure incrementality. It means they should start lower on the ladder: reconciliation, structural separation, and pause tests, from the incrementality guide, produce useful answers at a fraction of the volume requirement. A pause test is, in effect, a holdout in time rather than in geography, and for many accounts it is the right-sized version of the same idea.
How to Design One
A practical sequence, assuming the volume is there:
- Write the hypothesis and pass mark first, in the bracket format from the channel testing pillar: expected lift, window, budget, decision date.
- Pick matched regions using at least a year of baseline data where seasonality allows; check the groups tracked each other historically.
- Clean the edges: pause overlapping campaigns, retargeting, and brand spend in holdout regions; note anything you cannot control.
- Run the full window without touching the settings. Mid-test optimisation can contaminate the result, even when the intent is sensible.
- Read blended regional outcomes against the pre-agreed pass mark, and record the result either way; a clean negative is a spend decision funded by evidence.
Designing the regions, windows, and pass marks so the result stands up to a finance review is part of how we run new channel testing engagements: the geo test is the validation step between modelling a channel's potential and scaling it.