What Is a Holdout Test? Definition, Uses, and Business Value
Holdout Test
Definition
A test where a group is intentionally not shown ads so performance can be compared against an exposed group.
Overview
Holdout Test A test where a group is intentionally not shown ads so performance can be compared against an exposed group. This simple definition masks a range of practical choices marketers make: sample selection, randomization method, measurement window, and which KPIs to compare (sales, conversions, brand metrics, etc.).
Holdout testing isolates the causal effect of advertising by comparing outcomes between an exposed population and a similar group that did not see the ad. Unlike observational comparisons that can be biased by differing audiences or seasonal effects, correctly executed holdouts provide a defensible estimate of incremental impact — how much incremental sales, sign-ups, or other outcomes the ads produced.
Why Marketers Use Holdout Tests
Marketers use holdout tests to answer the central business question: did the campaign cause additional customer behavior beyond what would have happened anyway? Common business reasons include budget allocation, creative validation, channel mix decisions, and proof of performance for stakeholders or clients. Brands rely on holdouts when they need a causal estimate for ROI calculation or when attribution models are known to over- or under-count effects.
What The Test Typically Covers
- Population Selection: Defining the universe from which holdout and exposed groups are drawn (geography, user segments, cookie pools).
- Randomization: The method used to assign units to holdout or exposure (random user IDs, ZIP-code level, device IDs).
- Measurement Window: The period over which outcomes are measured (immediate conversions, multi-week lifecycle effects).
- KPIs: Primary and secondary metrics such as purchases, revenue, lead rate, or lift in brand awareness.
- Attribution & Leakage Controls: Rules to prevent holdout contamination (e.g., excluding household members of exposed users).
How It Differs From Other Tests
Holdouts are a form of randomized controlled test but differ from conventional A/B tests in scope and objective. A/B tests typically compare two or more active treatments (creative A vs. creative B) among users who are all exposed to treatment variants. A holdout explicitly includes an untreated control — a group that receives no ad — enabling a direct estimation of incremental effect versus no advertising at all.
How Sample Size And Statistical Power Vary
Detecting small incremental effects requires large sample sizes. Power calculations depend on baseline conversion rate, expected lift, variance, and desired confidence level. For example, detecting a 1–2% sales lift on a low baseline conversion may need hundreds of thousands of users or sizable geographic clusters. Marketers often balance statistical requirements against business constraints (e.g., allowable holdout size, campaign reach).
Practical Example
A retailer running a six-week online display campaign might set a 10% holdout at the user-id level. The exposed group sees display ads; the holdout sees none. After the campaign, the team compares incremental revenue per user across groups, controlling for seasonality and any site-wide promotions. If exposed users show $2.50 incremental revenue per user versus holdouts, the campaign ROI can be computed from that causal lift.
Common Pitfalls
- Contamination: When holdout users are exposed through other channels or household members, biasing results.
- Non-random Assignment: Assigning holdouts by convenience (e.g., last digit of user ID) that correlates with behavior.
- Insufficient Power: Running a test that’s too small to detect realistic lift.
- Short Measurement Windows: Missing longer-term effects like repeat purchases or brand consideration.
When A Holdout Is The Right Choice
Use a holdout when you need a clean estimate of incremental impact versus doing nothing — for new channels, cross-channel budget decisions, or when attribution models are unreliable. Holdouts are also appropriate for validating uplift from brand campaigns or expensive media where proof of causality is required.
Operational Tips For Implementation
- Randomize At The Right Level: Choose user, household, or geographic units to minimize contamination depending on how ads are delivered.
- Document Blocking Rules: Record exclusions, deduplication rules, and how repeat users are handled over multiple campaigns.
- Plan The Measurement Window: Predefine primary and secondary windows (e.g., 7-day, 28-day, and 90-day) and include holdout leakage checks.
- Control For External Factors: Account for promotions, price changes, or market events during the test window.
In short, the Holdout Test gives marketers a straightforward, defensible estimate of ad-driven uplift when designed and executed with attention to randomization, sample size, and contamination controls. Properly implemented, holdouts inform budget decisions and demonstrate causal impact to internal stakeholders and clients.
Sources And Additional Reading (4)
- Controlled Experiments on the Web: Survey and Practical Guide
Kohavi, Ron, Longbotham, Roger, Sommerfield, Daniel, and Henne, Randal. “Controlled Experiments on the Web: Survey and Practical Guide.” Microsoft Research, 2009, https://www.microsoft.com/en-us/research/publication/controlled-experiments-on-the-web-survey-and-practical-guide/.
- About experiments
“About experiments.” Google Ads Help, https://support.google.com/google-ads/answer/2404180.
- A Refresher on Randomized Controlled Trials
“A Refresher on Randomized Controlled Trials.” Harvard Business Review, May 2016, https://hbr.org/2016/05/a-refresher-on-randomized-controlled-trials.
- Media Rating Council
“Media Rating Council.” Media Rating Council, https://mediaratingcouncil.org/.
More from this term
Looking for a 3PL?
Compare warehouses on Racklify and find the right logistics partner for your business.