The Lookalike Model You Seeded With Your Best Customers Was Trained on Your Graph's Best Guesses
Lookalike models expand from a seed audience that was already shaped by identity graph resolution, so the population you scale toward reflects graph structure as much as customer similarity.
Lookalike modeling is one of the most widely used tools in digital media buying, and its appeal is intuitive. You hand a platform your highest-value customers, and the algorithm finds more people who resemble them. The logic sounds like it starts with your data and ends with more of your market. The problem is that it does not quite start with your data. It starts with what an identity graph decided your data looks like.
Understanding where that distinction matters, and how to account for it, is one of the more practical skills a media buyer can develop when evaluating lookalike-based spend.
What Actually Goes Into the Seed Audience
When you upload a CRM file to a DSP or activation platform as a lookalike seed, that file is not used in its raw form. It is resolved through an identity graph first. The graph maps your hashed emails, postal addresses, or phone numbers to device IDs, cookie spaces, or platform identifiers so the modeling system has something it can read.
That resolution step is not neutral. Every identity graph makes coverage and confidence tradeoffs. Some identifiers in your file resolve cleanly to a single profile. Others resolve to a probabilistic cluster. Some do not resolve at all and are dropped. The seed audience the model actually trains on is the subset of your CRM that survived resolution and was mapped in ways the graph was confident enough to include.
If your best customers are disproportionately represented among the identities that resolve cleanly, your seed is a reasonable proxy for your intended population. But if the graph's strongest coverage skews toward certain demographics, device behaviors, or data-rich geographies, then your seed is already a shaped version of your customer list before any modeling has begun.
How the Model Learns From a Shaped Seed
Lookalike algorithms learn what your seed audience looks like by reading the signals associated with resolved identities. Those signals, behavioral categories, interest clusters, purchase history indicators, are also products of graph resolution. The model is finding patterns in data that the graph already organized.
This creates a compounding effect. If the graph overrepresents desktop-heavy users in your seed because mobile resolution rates are lower for your file, the model learns that desktop behavior is a meaningful signal of your customer profile. It then expands toward people who also show that desktop behavior pattern, not because those people are genuinely more like your customers, but because the graph made desktop users more visible in the seed.
The model is optimizing accurately for what it was given. The issue is that what it was given was already filtered by infrastructure, not purely by customer quality.
Why Scale Decisions Made on Lookalike Audiences Inherit This Problem
Once a lookalike audience is built, buyers typically evaluate it by looking at predicted audience size at various similarity thresholds. A tighter threshold returns fewer people who score as very similar to the seed. A looser threshold returns a larger population with lower average similarity.
Buyers who scale out to broader lookalike tiers on the assumption that they are still reaching people with genuine affinity to their brand may actually be reaching people who share graph-visible traits with a graph-filtered version of their customers. The similarity the model measured was real within the graph's frame of reference. Whether it corresponds to actual customer affinity in the market is a separate question that the model cannot answer on its own.
This does not make lookalike modeling useless. It means the evidence you use to evaluate lookalike performance needs to account for the graph layer. Observing that a lookalike audience converted at a higher rate than a broad audience is not sufficient evidence that the model found genuinely similar customers. It may be that the graph's best-resolved identities are also the most reachable and therefore the most likely to show any measurable conversion event, independent of actual affinity.
Practical Ways to Pressure-Test Lookalike Quality
A few approaches can help buyers get more signal about whether a lookalike audience reflects real customer similarity or primarily graph structure.
First, consider comparing lookalike performance across two different activation platforms that use different identity graphs. If the lookalike audience built on Platform A and the one built on Platform B produce structurally similar customer overlap when matched back against your CRM post-campaign, that is a reasonable signal that the model is finding something real. If the two audiences have very little overlap with each other, that suggests graph structure is driving a significant share of who got selected.
Second, look at what percentage of your lookalike audience, after delivery, resolves back to records already in your CRM or suppression list. A high recapture rate can indicate the model is staying close to the graph's densest clusters rather than genuinely expanding outward.
Third, treat lookalike audience performance as a hypothesis to test rather than a result to report. Running an incrementality test against a random holdout drawn from outside the lookalike pool, rather than a platform-constructed holdout, gives you better evidence about whether the model is finding new demand or finding the same well-resolved population through a different path. This approach has its own design challenges, but it surfaces graph-driven results in ways that platform reporting alone will not.
Fourth, when evaluating vendors, ask specifically how the seed resolution process works before modeling begins. How are unresolved identifiers handled? What minimum confidence threshold is used? Are probabilistic matches included in the seed or excluded? Vendors that can answer these questions specifically are more likely to give you a seed that reflects your actual customers rather than the most graph-friendly version of them.
The Framing Shift That Helps
Lookalike modeling is a useful and legitimate technique. The framing adjustment that benefits buyers most is thinking of a lookalike audience not as an expansion of your customers but as an expansion of your resolved customers within a specific graph context.
That framing does not require abandoning the tactic. It does require treating the graph's coverage choices as a variable in your planning, not a neutral backdrop. When you evaluate a new lookalike strategy, knowing which graph resolved your seed, how comprehensively it covered your file, and how that coverage compares across potential activation paths gives you much better footing for deciding how much of your budget the model has actually earned.
The model is doing what it was designed to do. Making sure it was given an accurate picture to work from is the part that falls to the buyer.