The Incrementality Test You Ran on One Campaign Does Not Generalize to Your Next Audience or Channel
Incrementality results are specific to the audience, channel, and timing of the test that produced them, so reusing findings across different campaigns overstates what you actually know.

Incrementality testing has become one of the more credible tools available to media buyers who want to move beyond last-touch attribution. When a test is designed and executed carefully, it can tell you something genuinely useful: whether the specific audience you reached, through a specific channel, during a specific period, bought more because of your advertising than they would have without it. That is a narrow, precise finding. The problem arises when teams treat that finding as a durable property of their media strategy rather than a snapshot of one moment in one context.
What an Incrementality Result Actually Measures
Every incrementality test, whether a ghost bidding design, a geographic holdout, or a matched panel test, is answering a question about a particular combination of variables. Those variables include which audience was exposed, which inventory carried the ads, what creative was in market, what the competitive environment looked like at the time, and what else the brand was doing simultaneously. Change any one of those variables and you have a different question. The lift estimate from a connected TV test run against a retargeting audience in Q4 does not describe what you would observe running a prospecting campaign on display in Q2.
This matters practically because incrementality findings are often summarized into a single number, a lift percentage or an incremental cost per outcome, and that number gets attached to a channel in planning documents and budget conversations. Once it is attached, it tends to stay there. Teams cite it the following quarter, and sometimes the quarter after that, as evidence of channel efficiency, without examining whether the conditions that produced it still hold.
Three Ways Findings Stop Traveling
Audience composition changes the baseline. Incrementality measures the difference in behavior between an exposed group and a holdout. That difference depends heavily on the baseline purchase rate of the holdout. If your retargeting audience consisted largely of recent site visitors with high intent, a small advertising stimulus may produce measurable lift because the baseline group was already primed to convert. When you run a similar test against a cold prospecting audience with no prior brand contact, the baseline behavior is different, the time-to-conversion is longer, and a test window designed for the first audience may simply not be long enough to observe the effect in the second.
Channel mechanics change exposure quality. The same creative delivered through a high-attention placement and a low-attention placement will produce different lift. Inventory quality, format, and context all affect whether the exposure produces any persuasive effect worth measuring. A lift result from a premium publisher direct deal does not predict what you would observe from open exchange inventory targeting a similar audience definition, even if the targeting parameters look identical in your plan.
Market conditions shift the counterfactual. The holdout group in your test represents what would have happened without your advertising. That counterfactual is not a fixed number. It moves with seasonality, competitor activity, economic conditions, and changes in your own organic marketing. A holdout measured in a low-competition period will show a different baseline than one measured when a competitor is running a heavy brand campaign. Lift is always relative to the counterfactual, and the counterfactual changes.
What to Do Instead of Generalizing
The practical implication is not that incrementality testing is unreliable. It is that test design and finding interpretation both need to be scoped correctly from the start.
When you commission or interpret a test, document the specific audience definition, the inventory mix, the test window, the creative, and any known external factors that were active during the test period. Treat the result as describing those conditions, not as describing your channel broadly. When those conditions change meaningfully, the previous result should be flagged as context-specific rather than carried forward as a benchmark.
For channel-level budget decisions, consider running smaller, faster tests across different audience tiers rather than relying on one large test to speak for all of them. A hypothetical example: if a retargeting audience and a lookalike prospecting audience are both in market, a paired test design that holds out a portion of each independently will give you two incrementality estimates, each relevant to its own audience, rather than a blended number that accurately describes neither.
Test windows also deserve deliberate design rather than convention. A 14-day holdout may be appropriate for a category with short purchase cycles. For categories where consideration takes weeks or months, a 14-day window will undercount incrementality not because the advertising failed but because the outcome had not yet occurred at measurement time. Before interpreting a result as evidence of low lift, check whether the window was long enough to observe conversion in that category.
The Planning Implication
Incorporating incrementality into planning works best when findings are tagged with their conditions and stored that way. Some teams maintain a simple internal log of what was tested, when, against which audience, and what the result showed. Over time, that log starts to show patterns: which audiences tend to show strong incremental response, which channels produce consistent results across multiple tests, and where findings have not held up when conditions changed. That accumulated, conditioned evidence is more useful for planning than a single number treated as permanent truth.
It also changes how you present findings to stakeholders. Rather than stating that a channel delivers a certain lift percentage, you can describe what conditions produced that result and note which of those conditions will or will not be present in the next campaign. That framing is more defensible and, over time, builds more accurate expectations about what measurement can and cannot confirm.
Incremental lift is a real thing worth measuring. The findings just belong to the test that produced them, not to the channel or strategy indefinitely. Treating them that way keeps your planning honest and makes each subsequent test more informative than the last.