Clean rooms have earned a serious position in enterprise media buying, and for legitimate reasons. They allow first-party data collaboration without raw file exposure, satisfy legal and data governance requirements, and produce outputs that feel authoritative precisely because they are the product of a controlled, permissioned environment. The security architecture works as advertised. The problem is that buyers have migrated their trust in that architecture onto the data quality of the outputs—and those are two different things entirely.

What the Clean Room Actually Guarantees

A clean room enforces rules about what leaves the environment and who can see it. It prevents raw record-level exposure. It enforces minimum thresholds so individual users cannot be reverse-engineered from aggregate queries. It provides an audit trail. These are meaningful protections, and they address real regulatory and competitive concerns.

What a clean room does not do is validate the identity resolution that made the match possible in the first place. Every match result a clean room produces—every overlap count, every audience intersection, every reach projection—is downstream of an identity graph that resolved disparate identifiers into unified person-level records before either party's data entered the room. The clean room's output reflects the graph's resolution decisions, the graph's coverage gaps, and the graph's error rate. It then presents that result with the same visual authority as the governance controls surrounding it.

Buyers read the output. They see a number. They act on it as if the clean room certified the number's accuracy, when it only certified the number's safe transit.

The Resolution Layer Is Invisible to Both Parties

In a standard clean room deployment between an advertiser and a publisher or retail media network, both parties bring first-party data files. Those files are resolved against an identity graph—typically operated by a third-party identity infrastructure provider—before matching occurs. The graph decides which email address, which device ID, and which hashed phone number belong to the same person. The clean room then performs set operations on those resolved populations.

Neither party has direct visibility into how those resolution decisions were made. Neither party knows the graph's match confidence score for any given record, which signals were used to bridge identifiers, or how recently those signals were validated. The resolution layer sits beneath the clean room interface, and the clean room's output inherits its accuracy without surfacing its assumptions.

This means two things practically. First, the overlap your clean room query returns is an overlap of the graph's resolved identities, not necessarily an overlap of real people. Duplicate nodes, stale bridges, and probabilistic links that the graph treats as confident can inflate overlap figures. Second, gaps in graph coverage—populations that are under-represented in the graph's training data, less active across tracked digital surfaces, or frequently changing devices—will produce systematically understated overlap counts. Both errors are invisible in the query output.

How Budget Decisions Get Distorted

The downstream consequences are concrete. When a buyer uses clean room overlap data to negotiate inventory commitments with a publisher, the negotiation is anchored to a number that reflects graph fidelity, not audience truth. If the graph over-resolves—confidently linking records that belong to different people—the overlap appears larger than it is, and the buyer overcommits to inventory that will not reach the intended population. If the graph under-resolves—failing to link records that belong to the same person—the overlap appears smaller, and the buyer undervalues a partnership or pursues supplemental buys unnecessarily.

In either direction, the distortion is structural and invisible. It does not appear in post-campaign reporting. It does not surface in the clean room's audit log. It exists in the gap between the graph's internal state and reality—a gap that no standard SLA or privacy certification addresses.

The same dynamic affects suppression logic. When buyers use clean room outputs to build exclusion lists—suppressing existing customers from a prospecting campaign, for example—the suppression operates on the graph's resolved identities. Records the graph failed to resolve correctly will not be suppressed. Customers whose identifiers changed between graph ingestion and campaign delivery will not be recognized. The clean room enforced the suppression rule correctly; the graph simply did not give it the right people to suppress.

Separating the Security Audit from the Data Audit

The practical correction is procedural rather than technical, and it starts with recognizing that clean room governance and identity graph quality require separate evaluation frameworks.

On the governance side, the existing tooling is mature. Differential privacy mechanisms, query minimum thresholds, and access controls are well-understood and auditable. Legal and procurement teams have developed frameworks for evaluating them. This infrastructure is doing its job.

On the graph quality side, buyers need to develop a parallel discipline. Before relying on clean room outputs for budget allocation or audience strategy, the relevant questions are about the graph underneath, not the room around it: What is the graph's resolution methodology for the specific identifier types in both parties' files? How recent is the signal used to maintain graph linkages? What is the documented false-positive rate for the match confidence threshold being applied? How does coverage vary across the demographic and behavioral segments that matter to this specific campaign objective?

These questions do not require technical access to the graph's internals. They require that buyers make graph quality a term of evaluation alongside governance compliance—and that identity infrastructure providers be asked to produce documentation that addresses data accuracy separately from data security.

Why the Conflation Persists

The market has structural incentives to conflate these two properties. Clean room vendors have invested heavily in compliance architecture, and compliance is a legible, certifiable, commercially valuable attribute. Identity graph quality is harder to quantify, varies by use case, and degrades continuously in ways that resist point-in-time auditing. Separating the two would require buyers to ask questions that neither the clean room vendor nor the identity infrastructure provider has a strong incentive to answer in full.

The result is a market where the privacy-safe label has migrated from the governance layer to the output layer—where buyers treat clean room results as validated data rather than as the product of a resolution process that has its own accuracy profile and its own failure modes.

Data that is safely transferred and data that is accurately resolved are not the same thing. Clean rooms guarantee the former. The work of establishing the latter belongs to buyers who are willing to evaluate the graph as rigorously as they evaluate the room.