The Incrementality Test You Ran Measured Lift Against the Wrong Baseline Population

Incrementality experiments that use platform-constructed holdout groups inherit the same identity graph biases as your active audience, so the lift you measure describes graph behavior, not true causal response.

A brass surveying optical level on a wooden tripod sits in the left foreground with its eyepiece lens visible, pointing across a flat arid plain toward a red and white graduated surveying staff rod standing in the middle distance at right, with a white circular crosshair reticle superimposed in the sky between them under an overcast gray sky.

Incrementality testing has become the preferred alternative to last-click attribution for media buyers who want a cleaner answer to a direct question: did this campaign actually cause conversions, or would those buyers have converted anyway? The logic is appealing because it borrows from controlled experiment design. You split your audience into an exposed group and a holdout group, run the campaign against the exposed group, and compare outcomes. The gap between the two groups is your lift, and lift is your proof of causal effect.

The problem is not with the logic. The problem is with where the holdout group comes from.

In most platform-native incrementality setups, the holdout group is constructed by the same infrastructure that built your active audience. The identity graph that resolved your targeting file is also the graph that assigns users to the control cell. That means both groups are drawn from the same resolved population, which is not your full intended audience. It is the subset of your intended audience that the graph could find.

This matters because identity graphs resolve different people at different rates. Users with dense cross-device signal, stable email addresses, and frequent logged-in browsing are resolved with high confidence. Users who are less active online, who browse across many devices without logging in, or whose email addresses appear in fewer data partnerships are resolved with lower confidence or not resolved at all. These lower-confidence users are systematically underrepresented in both your active audience and your holdout group.

When you measure lift between two groups that share this structural bias, you are measuring causal response within the portion of your audience that was easiest to find. That is a meaningful number, but it is not the number your media plan assumed you were measuring. The buyers excluded by graph coverage gaps never appear in your experiment at all, and their conversion behavior goes uncounted on both sides.

What a biased holdout group actually produces

Consider a hypothetical scenario to make this concrete. Suppose your campaign targets a B2B audience of mid-market procurement managers. The identity graph resolves roughly 60 percent of your CRM file into addressable profiles. Your incrementality test splits those resolved profiles into an 80 percent exposed group and a 20 percent holdout group.

The 40 percent of your CRM file that the graph could not resolve is absent from both groups. Those buyers may convert during your campaign window through other touchpoints, through direct outreach, or through organic search. Their conversions are real, but they are invisible to your experiment. Your lift calculation reflects only the resolved 60 percent, and any conclusion you draw about campaign effectiveness is scoped to that subpopulation.

If the unresolved 40 percent converts at a meaningfully different rate than the resolved 60 percent, your lift figure overstates or understates the campaign's true causal contribution depending on which direction that difference runs. You cannot know which way the error points without auditing who was excluded.

Why this bias is rarely surfaced in standard reporting

Platform incrementality reports do not typically show you the size or composition of the population that was excluded from both cells. They show you exposed group size, holdout group size, conversion counts for each, and a lift calculation. The denominator of your intended audience, and the gap between that denominator and your experiment population, is not a standard output.

This is not a disclosure gap driven by bad intent. It reflects the fact that most platform measurement infrastructure is built to report on what it can observe, not on what it cannot. The excluded population is, by definition, outside the system's observability. Buyers who accept the lift figure as a complete answer are not being unreasonable. They are working with the information available, which is the issue.

Practical steps for auditing your holdout group composition

Before treating any incrementality result as a reliable input to budget decisions, it is worth asking a few structural questions about how the holdout was built.

First, ask your platform partner or measurement vendor what share of your input audience file was resolved into the experiment population. If you onboarded 50,000 CRM records and only 28,000 appeared in the combined active and holdout cells, the 22,000 gap represents buyers whose behavior your experiment never captured. Understanding the size of that gap is the starting point for knowing how much to trust the lift figure.

Second, ask whether the holdout group was constructed before or after graph-based expansion. Some platforms apply lookalike or probabilistic extension before assigning holdout cells, which changes the population being tested from your first-party audience to a modeled approximation of it. Lift measured against an expanded lookalike holdout answers a different question than lift measured against a matched first-party holdout.

Third, if your organization has access to a data clean room environment, consider whether you can match your full intended audience file against the experiment population and characterize the excluded group by available firmographic or behavioral attributes. The goal is not to add the excluded group to your experiment retroactively, which is not possible. The goal is to understand whether the excluded group is systematically different in ways that should limit how broadly you apply your lift conclusions.

How to scope your conclusions appropriately

An incrementality result from a graph-constrained experiment is still useful. Lift measured within the resolved population tells you something real about how that population responds to paid exposure. If the resolved population is reasonably representative of your broader audience, the result may transfer well. If the resolved population is structurally distinct, the lift figure should be scoped accordingly.

A practical framing for reporting these results internally is to label them as lift within addressable reach rather than lift across intended audience. That distinction signals to budget owners and planning teams that the measurement reflects a subset of the campaign's target, not the full population the campaign was designed to reach. It also creates a natural prompt for follow-on questions about how to extend measurement coverage over time as identity resolution improves.

Incrementality testing is a more honest measurement approach than most alternatives available to media buyers today. But the quality of the experiment is bounded by the quality of the population it can observe, and that population is shaped by graph infrastructure before a single impression runs. Auditing the holdout group's composition is not a technical exercise reserved for data science teams. It is a basic planning discipline that belongs in any media measurement workflow.

Keep up with B2B Solution Journal

Practical guidance and new coverage. You can withdraw your permission at any time.

Read our privacy and data-use policy.