The Audience Your Clean Room Query Returns Is Shaped by Which Party's Identity Keys Anchor the Join

Clean room query results depend on whose identity keys drive the join logic, and that anchor choice silently determines which records survive and which disappear.

Two open hinged brass casting molds sit side by side on a dark stone surface, the left mold holding a narrow cross-shaped silver cast and the right mold holding a wider mushroom-key-shaped silver cast, with a matching finished cross-shaped piece and a smaller rougher key-shaped piece resting below each respective mold.

When two parties bring data into a clean room and run a join, the output looks like a neutral overlap count. It feels objective because neither party handed the other their raw records. But the number that comes back is not a fixed fact about the overlap between those two populations. It is a fact about the overlap as seen through whichever party's identity keys were used to anchor the join. Change the anchor, and the count changes.

Understanding why this happens, and what it means for decisions built on clean room output, is practical work for any media buyer who uses a clean room to plan, measure, or negotiate.

How a clean room join actually works

A clean room join matches records from two datasets by finding shared keys. In practice, those keys are almost never raw email addresses or phone numbers sitting in both tables simultaneously. They are hashed identifiers, pseudonymous IDs, or resolved graph nodes that each party brings in.

When a publisher and an advertiser run an overlap query, someone's keys have to serve as the left side of the join. The records that survive are the records that found a match on those left-side keys. Records that the left-side graph cannot represent simply do not appear in the output, not because they don't exist in the right-side data, but because the anchor graph has no key for them.

This means the join is not measuring overlap between two populations. It is measuring overlap between two populations as both of them are visible through one party's identity resolution layer.

A hypothetical example of how the anchor changes the result

Imagine an advertiser brings a CRM file of 200,000 customers into a clean room. A publisher brings their logged-in audience. The advertiser's identity vendor resolves those 200,000 CRM records to 140,000 graph nodes. The publisher's own identity layer covers a different slice of the same real-world population.

If the join anchors on the advertiser's graph nodes, the query finds matches only within those 140,000 resolved records. Customers whose CRM records did not resolve to a graph node are invisible to the query regardless of whether the publisher actually reaches them.

If the join instead anchors on the publisher's identity keys, a different population anchors the search. Some records that were invisible under the advertiser's graph now surface. Some records that appeared before may not have a match on the publisher's side. The reported overlap number shifts, sometimes substantially, without either party's underlying data changing at all.

Neither result is wrong in a simple sense. They are just answers to subtly different questions.

Why this matters for measurement and negotiation

Clean room outputs are increasingly used for three high-stakes decisions: verifying audience overlap before a buy, measuring reach after a campaign, and negotiating cost-per-reached-customer arrangements with publishers. All three decisions become unreliable if the anchor assumption is invisible or inconsistent.

For overlap verification before a buy, an advertiser anchoring on their own graph will see a conservative overlap estimate if their graph has weaker coverage than the publisher's. They may underestimate how much of their CRM the publisher actually reaches, and underspend on that publisher as a result.

For post-campaign reach measurement, if the clean room query anchors on the publisher's delivery log keys, reach counts will reflect how many of the advertiser's customers appeared in the publisher's identity layer, not necessarily how many the advertiser's own graph would recognize as customers. Those two counts can differ meaningfully depending on how the two graphs were built.

For cost-per-reached-customer negotiations, the agreed definition of a reached customer needs to specify which identity layer governs the count. Without that specification, each party can run the same query and read a number that genuinely reflects their own graph's answer to the question, and those numbers can diverge enough to create billing disputes that are genuinely hard to resolve because neither number is fabricated.

What to ask before you run or commission a clean room query

Before treating a clean room output as a reliable planning input, it is worth asking four specific questions about how the query was constructed.

First, whose identity keys anchor the join, and what is the coverage rate of that graph across the relevant population? A graph with strong coverage among younger urban consumers and weaker coverage among rural buyers will systematically undercount overlap in the latter group regardless of what data the other party brings.

Second, is the join symmetric or directional? Some clean room implementations support bidirectional joins that can surface records visible to either party's graph. Others are strictly directional. Knowing which you are running changes how you interpret the output.

Third, what happens to records that one party's graph cannot represent? Do they drop silently, or does the query flag unmatched volume so you can see how much of each dataset is invisible to the other? Flagged unmatched volume is significantly more useful for planning purposes than a clean overlap count alone.

Fourth, is the same anchor logic being used consistently across measurement periods? If an early campaign flight used the publisher's keys as anchor and a later flight used the advertiser's keys, comparing reach numbers across those two flights will confound the anchor difference with any real change in campaign performance.

A practical suggestion for structuring clean room agreements

When negotiating a clean room measurement arrangement with a publisher or data partner, consider specifying the anchor party in the contract language alongside the query definition. This does not have to be technically complex. It can be as simple as agreeing that the baseline overlap count uses the advertiser's resolved graph as anchor, and that any publisher-anchored query is labeled separately.

This kind of documentation also makes it easier to revisit results if a graph is updated mid-flight, which happens more often than either party tends to anticipate. Identity graphs are updated continuously, and a graph update can change what the anchor covers even if neither party's underlying customer data changed.

The goal is not to make clean room measurement more complicated. It is to make the assumptions explicit enough that the output can be interpreted consistently by both parties across planning, delivery, and reporting cycles. A number without a documented anchor is harder to act on than a number with one, because you cannot know whether a change in the output reflects a real change in audience overlap or a shift in what the anchor graph can see.

Keep up with B2B Solution Journal

Practical guidance and new coverage. You can withdraw your permission at any time.

Read our privacy and data-use policy.