Racklipedia
Racklify
​
Marketing

Common Biases That Break Incrementality Tests And How To Fix Them

Updated October 1, 2026
Published October 1, 2026
William Carlin

Incrementality

Definition

Measurement of whether ads caused additional sales that would not have happened without the advertising.

Overview

Incrementality is the extent to which sales, customers, or other outcomes were caused by a marketing or promotional activity rather than occurring anyway. Even well-intentioned tests can produce biased incrementality estimates if common failure modes—selection bias, contamination, instrument changes, or seasonality—are not addressed.


This entry catalogs the usual biases that invalidate incrementality studies, explains how each bias works in practice, and offers concrete fixes operations teams can implement before, during, and after tests. The goal is to help marketers, analysts, and operations leaders design resilient tests and trust their causal estimates.


Selection Bias


Selection bias happens when treatment and control groups differ on unobserved factors that affect outcomes (e.g., higher-value customers more likely to be exposed). In observational methods, selection bias is the primary concern. In experiments, selection bias arises from improper randomization or post-randomization reallocation.


  • Fix: Use true randomization and deterministic assignment (hashed user ID). For observational work, apply matching, control for covariates, and run placebo tests.


Contamination And Spillover


Contamination occurs when control units are exposed to the treatment (cookie churn, cross-device exposure, household members). Spillover is when treated users influence controls (word-of-mouth, local store effects). Both bias estimates toward zero or create unpredictable distortions.


  • Fix: Use higher-level randomization (household, household+device, geo) or design the experiment to limit cross-exposure. Measure and correct for contamination where possible (instrumental variables or adjusted estimators).


Instrumentation And Measurement Changes


Changes in tagging, analytics platforms, or conversion definitions during a test break comparability. For example, a new analytics SDK deployed mid-test can increase observed conversions in one arm and invalidate results.


  • Fix: Freeze instrumentation for the test window or run a short A/A validation after any change. Track release windows and tag deployments in test documentation.


Seasonality And External Shocks


Seasonal demand swings, promotions from competitors, or macro events can confound tests if they disproportionately affect treatment or control groups. Geo tests are especially sensitive if regions experience different externalities.


  • Fix: Use pre-test windows to check parallel trends, lengthen the test to cover multiple cycles, and include time-fixed effects in models. If a shock occurs, pause the experiment or treat results as invalid.


Regression To The Mean


Picking treatment units because they recently performed poorly or well leads to regression to the mean: subsequent changes may reflect natural reversion rather than treatment effect.


  • Fix: Avoid choosing test groups on the basis of short-term outliers; use longer pre-periods and stratified randomization.


Underpowered Tests And Multiple Comparisons


Small sample sizes fail to detect meaningful lifts; running many tests or segments without correction inflates the false-positive rate.


  • Fix: Perform power calculations before launching, adjust for multiple comparisons (Bonferroni, Benjamini–Hochberg) when testing many segments, and favor fewer, higher-quality tests over many noisy ones.


Cookie Churn, Identity Decay, And Cross-Device Issues


When identities are short-lived (cleared cookies, ad ID resets), treatment assignment and exposure tracking can flip during the test, eroding measured lift.


  • Fix: Use persistent identifiers where privacy-compliant (logged-in user ID, CRM match), adopt household-level assignment for CTV/OTT, and model expected churn in sample-size estimates.


Practical Example: Fixing Geo Contamination


A brand runs a geo TV campaign and assigns neighboring DMAs as treatment and control. During analysis they find controls were served ads via streaming channels that cross DMA boundaries, contaminating the control. The fix was to re-run the analysis excluding streaming impressions by geography, reassign slightly larger non-adjacent geos as controls, and report results with contamination-adjusted confidence intervals. The team also shifted future tests to use market clusters with limited cross-media leakage.


Prevention Checklist


  • Label: Pre-register the test plan including KPI, assignment unit, sample-size calc, and analysis script.
  • Label: Run an A/A test or compare pre-period trends to validate balance.
  • Label: Lock analytics instrumentation and log any platform changes during the test window.
  • Label: Monitor exposure and contamination metrics in real time and pause if major leaks occur.


In short, the Incrementality estimate is only as credible as the experiment’s defenses against bias. Anticipate selection, contamination, measurement drift, and external shocks, design tests to limit them, and apply corrective analyses when issues are unavoidable.


Sources And Additional Reading (3)

More from this term
Looking for a 3PL?

Compare warehouses on Racklify and find the right logistics partner for your business.