- From: Martin Thomson <notifications@github.com>
- Date: Tue, 15 Sep 2026 17:13:59 -0700
- To: w3ctag/design-reviews <design-reviews@noreply.github.com>
- Cc: Subscribed <subscribed@noreply.github.com>
- Message-ID: <w3ctag/design-reviews/issues/1229/5689958678@github.com>
martinthomson left a comment (w3ctag/design-reviews#1229)
Regarding the explainer:
The working group has bad experiences with explainers and the balance of value that they represent. Such documents tend to be throw-away artifacts that decay badly as the main specification evolves, leaving either an unworkable maintenance burden or an untrustworthy artifact that does not represent reality.
It’s hard enough to have a specification that matches tests and implementations. Adding an explainer is not just one more thing to synchronize, but something that is naturally stuck at a point in time.
The things that the Attribution API specification lacks are precisely those things: the analysis that was done to show that what is captured is most appropriate and any alternatives that were considered. These things are the most prone to decay over time and so are best suited to supplementary material.
We could generate documentation of aspects like that, but feel it is not a good use of our resources, even if that makes it harder for the TAG to do a review. The TAG are uniquely affected here, so I’d be willing to discuss any of this with the TAG to help build that understanding; I’m sure others would too.
The other function of an explainer is as introductory material. To some extent, this is so far outside the present zeitgeist that it is understandable that the leap in understanding is a bit hard to make for some. This is an area where the group is still discussing how to make the API more comprehensible in a broader sense. The current thinking seems to be that this is going to involve multiple artifacts rather than the typical “explainer” pattern (for example. [this bit on DAP](https://lowentropy.net/posts/dap-basic/)).
Brian’s [AI-generated explainer](https://gist.github.com/bkardell/495b7d717b6e25867c1015815c0c29c1) is slop. It is not uniformly bad, as a good amount of the content is a straight regurgitation of what is in the specification. Of course, duplication without adding value is not worthwhile. The added content is both interesting and where things go less well. In terms of giving potential readers an appreciation for the spectrum of possibilities that converged on this particular design, it’s not a great reference (some of the text is implausible to the point of being fantasy).
As a personal note, I find that the ability of an LLM to produce more accessible interpretations of complex documents to be one of the more positive aspects of the technology. I have no problem conceding that this is a complex document. If you found Brian’s explainer helpful as a way of gaining an initial understanding, that’s great. That cannot cross the gap to a publishable and well-maintained artifact without significant effort.
---
Just a brief note. Some of this feedback appears to have been generated by an LLM. It is replete with the usual telltales. Even with those telltales, I’d have chosen to bite my tongue if it weren’t for one particular point below.
---
If you will entertain a digression on that same point…
By the way, I do think that LLM-generated reviews are useful. Mark Nottingham has developed a review suite for IETF documents that is quite handy. I don’t like the way he dumps its output on mailing lists though. Many of the “findings” that come from these tools are garbage. In my view, such tools are best suited to giving the authors of specifications some pointers to areas they might like to improve. For a review, such as those the TAG provides, I hope we eventually get past the point where those sorts of issues are raised during review, so that the TAG can concentrate on the higher-level issues.
In this case, the sorts of high-level questions you might ask here are the sorts of things that might be best directed at a working group charter, rather than a specification. That raises interesting questions for the overall process, because the charter is “agreed” much earlier on. The problem being that it is often impossible to gain the understanding necessary to really decide whether you might agree to something until the specification is in front of you. Generally, we frown on people who challenge the charter at this stage. It’s almost like the game is rigged in favor of forward motion. I’ve seen the same bias in many places, which is often accompanied by either indignant or frustrated proponents, who justifiably protest that they followed the rules scrupulously.
In other words, I’m not offended if you want to challenge the underpinnings of this work. I know it’s controversial. I happen to think that it’s worth doing, but many people hold understandable reservations.
---
Regarding the specific points raised:
> The formal DP guarantee is conditional on assumptions the spec states cannot hold. The per-site DP claim is not unconditional, and the spec acknowledges it.
I could choose to interpret this as positive feedback, even if it was under a “please fix me” category. This sort of assumption or limitation is very common in formal analysis of complex systems. That these “gotchas” are so narrow should be reassuring, not a cause for concern.
The work that was done on the privacy analysis shows a very narrow case where the DP guarantees do not hold in the very strictest of senses. The first listed “gap” recognizes that, for a given epoch, generic learnings from some sites (e.g., people shown ads like this buy the green option more often) can be exploited by other sites that have not yet expended their privacy budget to learn new things that build on that information (e.g., let’s focus on the green options and see what shade is selling better).
In a sense, it’s a special form of adaptive composition that the DP analysis cannot allow.
The model we use includes the possibility that a single entity controls multiple sites or that sites can collaborate to maximize advantage. That’s necessary, because that’s how the web works. Including that in the model means that this sort of composition works against the ironclad guarantees that differential privacy seeks to provide.
This is a somewhat theoretical concern, because this is also a good description of how the API is supposed to work, only within a single epoch (the privacy unit) rather than across epochs. We could provide a slightly different design that eliminated same-epoch crosstalk between sites, by delaying the aggregate results by a week. The net effect would be the same: sites could still gain an advantage by sharing information about any given week, using that information in the following week. They would just have to wait longer. At the same time, the utility would crater because of the delays. Increasing the waiting time to seven days makes the API useless for many purposes.
Given that the information obtained is not individually-attributed, such that what sites learn is generic, the conclusion was that addressing this sort of “leak” was not justified. That said, it might slightly motivate parameter choices (like smaller epsilon) that bias toward better privacy.
The second leak mentioned in the spec is leakage through cross-site budgets. That is something that is easy to describe, but not something for which I can articulate any concrete privacy consequence. Hitting a shared budget limit results in obtaining a zero-valued histogram, which is indistinguishable from not hitting that limit, so this is more of a theoretical limitation on the analysis itself than it is something that has privacy consequences. Relative to how theoretical the other leak is — at least in my view — this is not worth further analysis.
> Privacy-determining parameters are implementation-defined with weak floors
Indeed. That was very much deliberate. Browsers wanted the ability to differentiate in the market based on their parameter choices. The working group explicitly ruled agreement on privacy parameters as out of scope.
It is also entirely possible to give users control over these parameters (though we do not expect that to happen).
> Small-cohort / micro-targeting inference - It would be beneficial if the spec can analyse the worst case where an adversary engineers batch composition down to min_batch_size. Either raise the minimum, tie it to ε, or document the residual inference risk explicitly.
The differential privacy guarantees do not depend on the minimum batch size in DAP at all. In other words, differential privacy IS a worst-case analysis (there’s a related discipline that talks about empirical privacy, but we explicitly do not use that).
The value for min_batch_size in DAP exists to support the shuffle DP mode that is used by other applications. We don’t rely on it in any meaningful way. We can’t guarantee that the raw inputs for all the other report submissions aren’t known to the site that receives the final aggregate, no matter what value we set. The value in the Attribution spec is somewhat arbitrary. It could be 1 and the privacy properties would be largely the same.
Not that some aggregation isn’t useful. Intuitively, having your data mixed with others provides some reassurance. After all, other people add their own randomness. It’s just not something we rely on in any analysis.
> Aggregation-service trust and governance are out of scope, but the guarantee depends on them
Yes. And? I mean, this is also true for the Web PKI. Somehow we stagger onward. The requirements are pretty clearly laid out in the specification, which is all that a specification can reasonably do. Having to do extra work as a precondition of deployment is a giant nuisance, of course, but it’s not something that W3C is set up to do.
> Ad fraud / invalid traffic is acknowledged but unmitigated on-device - should be surfaced as an explicit non-goal ("this API does not provide Sybil resistance for measurement integrity") rather than left implicit in a security subsection, since advertisers may otherwise assume the numbers are trustworthy.
That too is an assumption. My experience of advertisers and the advertising at large is that there is abundant and healthy skepticism of any measurement technique. Maybe it’s just my unique exposure to people in the industry, but the level of sophistication I’ve observed when it comes to measurement is quite high.
Still, this is superficially a reasonable ask, except that it is not a non-goal at all. It’s just that this version of the API does not achieve the sort of robustness against IVT that we might eventually aspire to.
It’s probably not hard to imagine that this was a particularly contentious topic in the working group. There is a fairly good appreciation of the complexity of the problem, as well as a general sense of the tools that measurement professionals in the industry handle the problem. Those tools carry their own privacy hazards, as you might imagine, as they include fingerprinting and other techniques that we might consider questionable, all being used extensively.
Still, the working group reached a conclusion regarding this risk and the result is what you see.
This is not the absence of defenses as your comment makes out. Yes, it is in some key ways inferior to the tools available to a measurement provider who is able to use cross-site cookies. Nor is it a complete open door either. Conversions can be authenticated and measurement aggregation delayed until that happens; questionable items can be withheld from aggregates.
Additionally all existing IVT checks and validation can still be put into play before an impression or conversion is triggered. The advertising industry still contains multiple many-multi-million dollar companies who handle validation via client side JavaScript whose technology can still be employed as usual in this process.
What this does not provide is good defense against sites that seek to “snipe” conversion attribution. That’s a pre-existing and fundamentally unsolved problem, but one that this API can make worse. It is most evident when the simplistic last-touch attribution method is used, but the API does not need to be used that way at all. It provides a range of tools that are far more resistant to sniping than it might seem on the surface.
Detailing those is not the business of a specification, but supporting documentation, and we don’t expect that to be really useful until the API is deployed and in use for some time. Right now, the only experience we have in this area is with Google’s Attribution Reporting API, where the finer details of usage are not transferable.
> Enabled by default in third-party contexts, opt-out only: The undetectable-opt-out design is genuinely good as it prevents discrimination against users who decline. But default-on, third-party-exposed, opt-out participation in a cross-site data flow is a configuration that does not follow the Privacy Principles' consent expectations. The "collective privacy" framing argues for a policy outcome (enable for everyone) inside a technical spec; the TAG should decide whether that argument belongs in normative material. We think it is good to justify the * default and third-party availability against Design Principles ("Design for user intent") and ("Help users make good decisions"), which themselves point to the Privacy Principles' consent principles, and separate the normative behaviour from the advocacy (Note the API is already user-activation-gated per spec §4.2.2, so the concern is the consent model and defaults, not the absence of activation.)
The specification might appear to create the impression that it is suggesting that the API be enabled by default. But it does not mandate anything. I suggest that you read it again. What it does is lay out an argument for why enabling it by default could be acceptable or even beneficial. Browser implementations always have discretion in what their defaults are and how to interact with the people who use, or might use, their product.
This is introductory material, which to some extent needs to address the question of why the API exists, so the notion that it is “advocacy” is only true insofar as it advocates for the existence of the API. If that is not acceptable material for a specification, I have some other APIs I could refer you to.
If the reference to “* default” here is with respect to permissions policy, that’s a separable issue. (I’ve read this several times and can’t find any other explanation for its inclusion other than it being LLM output. That is, it looks like LLM output that should not have been forwarded on in the review. That being the case, the TAG can decide if they think this choice is aligned with their view of how the web should work. The other interpretation is less positive: “the TAG should decide” could imply something more than that. If the intent of that statement is that you have some decision-making authority, that’s not a power you have. The W3C membership decides on approving specifications.)
The “* default” in permissions policy was made by the working group based on evidence that deployment of anything that depends on every site acting individually has been slow or unsuccessful in the past. That decision was made in full awareness of some of the problems that decision creates. If concerns about this decision are grounded in the idea that a first- third- party availability distinction has some sort of meaningful impact on privacy in 2026 and beyond, that’s something we should debate.
I also respectfully disagree with your reading of the privacy principles. That document covers a lot of ground, but several sections reference this work indirectly. [Collective governance](https://w3ctag.github.io/privacy-principles/#collective) addresses the question of what it means to determine defaults. (To some degree, this discussion is that principle in action; the decisions that browser implementers make are another aspect.) Similarly, [collective privacy](https://w3ctag.github.io/privacy-principles/#collective-privacy) and [ancillary uses](https://w3ctag.github.io/privacy-principles/#ancillary-uses) are sections that address this style of API. This work is cited in [Example 7](https://w3ctag.github.io/privacy-principles/#example-preventing-profiling), though it doesn’t acknowledge that the mentioned personal data could also be [deidentified](https://w3ctag.github.io/privacy-principles/#deidentified-data) (something this API does).
It’s explicit in the privacy principles that this sort of “ancillary use” thing is contested. So there is no expectation that you, as individuals or a group, need to agree with the conclusion.
> matchValue (and related fields) can carry identifiers
>* suggest cross-referencing §9.5 to the budget/cohort assumptions and state the residual risk when those are configured permissively.
It’s possible that you are not understanding the full implications of this language. This text exists to underscore the importance of the privacy design in the API, nothing more. That is, we are making it clear that the information in the impression store is controlled by adversaries.
https://github.com/w3c/attribution/pull/492 is a short acknowledgement of that. It incorporates the suggestion, which is a good one.
Side-channel mitigation is largely delegated and hard to test - define at least one normatively checkable property?
See https://github.com/w3c/attribution/pull/493 for this. It was previously implied, but now it has the smell of must.
> Wall-clock dependence (§9.8): the one-time budget-renewal risk from forward clock jumps is acknowledged; consider a normative bound on tolerated forward correction within an epoch.
The effect is barely worth mentioning, so I’m surprised you think it worth extra text.
A bound doesn’t help. Accelerated time means accelerated availability of additional privacy budget. So a cap to the one-off jump isn’t addressing any particular concern.
It’s also basically impossible for a browser to do anything. The browser is downstream from the system clock. It won’t necessarily know that the time shift has occurred.
The privacy exposure via the shifting clock is more of a privacy concern. Consider the extreme case where a user’s clock runs seven times faster than ordinary. In this case, their privacy budget renews daily rather than weekly. That’s a real problem because their privacy loss is seven times higher over time. (This demonstrates the futility of the browser attempting to do something: no single detectable event exists where time moves forward in a jump.) The main privacy problem is that a browser attached to a 7x clock is trivially fingerprinted by any site that is visited over any non-trivial time period that happens to call Date.now(). (Not to mention that this sort of weirdness is a tell that should trip any IVT filtering.)
A user that experiences a one-time shift forward, at worst, experiences a temporary increase in privacy loss through this API, sure. That exists regardless of where the time shift occurs, or how long it is. A site that can observe time on either side of the shift is better able to treat the shift as a strong fingerprint than use the one-off boost to privacy budget.
From the perspective of the privacy budget release over time, a one-off correction is not so different to the budget renewal that happens every week. A one-off increase does increase the overall privacy loss rate, but the idea behind setting budgets is that whatever privacy loss occurs over several epochs needs to be within acceptable bounds. (There is considerable debate about the stability of user preferences and activity such that composition over time is a material privacy risk, which is another reason why the specification is silent on the exact value for privacy budgets.)
> Unconfigured browsers (§9.4): the requirement to obtain aggregator config out-of-band to avoid a timing/fingerprinting leak can be interop constraint and should be a MUST, not a SHOULD, given the consequence.
If it weren’t for the fact that browser startup is already trivially fingerprintable, I’d agree. The working group couldn’t agree to make this a hard requirement. No browser was willing to extend the time it takes to launch the browser or compromise the ability to run offline. I doubt that this will change for this API. As the text notes, the alternative is to have the API. This fails the information-leakage test in a technical sense, but it avoids worse leakage.
As a practical matter, the expectation is that this configuration will be relatively static. Configuration can likely be baked into browser builds, with opportunistic updates deployed through browser configuration services in ways that don’t cause the API to become unavailable at any time. So this is more of a theoretical risk, which we note, just as we do all the other theoretical risks. Risks are to be understood then managed if they can't be eliminated. And the cost of eliminating this one was considered too high. I'd be curious whether you agree, armed with this new knowledge.
---
As a final note, I appreciate the time you have taken to look at this work. I know it isn't easy and I can understand why you might prefer this work not continue. It's hard to look at the damage that advertising has done and then be faced with something that will ultimately help that industry. Personally, I'm sick of the sorts of zero-sum games we play. Those only escalate into non-constructive outcomes, like outright conflict. We can be better than that.
--
Reply to this email directly or view it on GitHub:
https://github.com/w3ctag/design-reviews/issues/1229#issuecomment-5689958678
You are receiving this because you are subscribed to this thread.
Message ID: <w3ctag/design-reviews/issues/1229/5689958678@github.com>
Received on Wednesday, 16 September 2026 00:14:04 UTC