DisclosedConsumer transparency infrastructure · US & Canada

The record

Methodology — published limitations

Rendered at build time from docs/METHODOLOGY.md in the repository. The page and the file are the same text; there is no separate public version of it.

Draft content for the eventual /methodology page. Written 2026-09-06, when each finding was fresh, and published before any score is.

These are stated the way SPEC-scoring.md §6 states its confounds: by us, first, unprompted. A limitation someone else discovers in our numbers is a scandal. A limitation we publish alongside them is a methodology. Not one of the sections below is a weakness to be buried; each is a fact a reader needs in order to read our numbers correctly.


Some publishers' affiliate links cannot be resolved, and that is their configuration

Many publishers route outbound product links through a redirect endpoint on their own domain rather than linking to the merchant directly. The link a reader clicks says publisher.com/out/...; where it actually goes is only visible by following the redirect.

We follow those redirects server-side, from our own infrastructure, without cookies, and without ever requesting the destination — we read the forwarding address out of the response header and stop. The reader's browser is never involved and no affiliate cookie is ever set by us. We also respect robots.txt on every request we make, including requests to redirect endpoints. A redirect endpoint is not an exception to that rule.

Wirecutter's robots.txt disallows its own affiliate redirect endpoint. As of 2026-09-06, nytimes.com/robots.txt contains Disallow: /wirecutter/out/ under User-agent: *, and Wirecutter routes its outbound product links through exactly that path.

Re-measured 2026-09-07 over the 345-article corpus, and the earlier figures are corrected rather than quietly updated. This paragraph previously read "Of 1,117 cloaked links in our corpus, 1,029 are Wirecutter links we are not permitted to resolve. The same applies to click.linksynergy.com (23 links) and a small number of others." Every number in it was stale and the ratio was wrong. Measured now:

2,175 links in the corpus are same-host cloaks, and 2,175 of them — every single one — are nytimes.com/wirecutter/out/. Not 1,029 of 1,117 (92%); 2,175 of 2,175 (100%). There is no "and a small number of others": on this definition there are no others at all.

The click.linksynergy.com sentence is removed rather than renumbered, because it was never true under this definition. A linksynergy link is CROSS-host, so it is not a same-host cloak and was never among these links whatever the count. (The corpus holds 148 of them; they are classified on the rakuten signature, which is verified: false, so they are likely and have nothing to do with this paragraph.)

Updated 2026-09-07. Two of the three things we did not know about those links, the publisher had written into them all along. Wirecutter marks each of those links rel="sponsored" in its own markup, and 918 of the 1,066 distinct sponsored-cloak URLs also carry a merchant= parameter naming the merchant — merchant=Amazon, merchant=Walmart, merchant=Best%20Buy. So the three questions come apart, and we now answer them separately instead of collapsing all three into "unknown":

QuestionAnswerOn whose evidence
Is this link paid?YesThe publisher's own rel="sponsored", which is the attribute the HTML specification defines for paid placement and what a publisher sets to satisfy 16 CFR 255.
Which merchant is paying?The one named in the linkThe publisher's own merchant= parameter — its own word, in its own markup, for the party paying it.
Where does the link actually go?Still unknown, and always will beNobody's. We are not permitted to look.

The monetization and the merchant are the publisher's own declaration. The destination remains unobserved, because /wirecutter/out/ is robots.txt-disallowed to us. And the rate is the declared merchant's own published commission schedule, cited and dated on the card like every other rate we publish. We did not estimate it, and we did not derive it from the link.

Two limits on that, stated by us rather than left for a reader to find:

  • Declared is not observed. We know the merchant Wirecutter named. We do not know that a reader clicking that link arrives there, because checking would mean following a redirect we have been told not to fetch. Every card that prices a pick this way says so in those words.
  • A name is not a storefront. merchant=Amazon names a company; amazon.com and amazon.ca publish different commission schedules and the parameter says nothing about which one a reader reaches. We price a bare "Amazon" against the US schedule, and that is an assumption we are making, not something we observed.

Where a declared merchant resolves to no rate card we hold — an opaque publisher id, or a merchant whose schedule we have not fetched — the pick keeps r_i = null and drops out of the correlation exactly as before, and the count of those picks is published rather than quietly absorbed. On cnet.com that is every declaring pick: CNET writes an opaque id (merchant=05kie42h3YvHwjr4G1w80Qq) rather than a name, so its declarations price nothing. Its rates come from a different declaration — the destination URL CNET's own redirector carries in plain sight.

For the links that remain genuinely unresolved: we do not know the merchant, so we do not know the commission rate, so those picks are recorded as unknown and excluded from the Payout–Rank correlation. They are not counted as unmonetized, and no rate is estimated for them. Where a publisher's picks are predominantly of this kind, that publisher may have too few priced picks to be scored at all, and its card will say so.

This is a property of the publisher's configuration, not a limitation of the method. We could resolve those links by ignoring robots.txt; we do not, and we would rather publish a smaller dataset gathered within a publisher's stated wishes than a larger one gathered by disregarding them. We state which publishers this affects and how many of their links it covers, so a reader can weigh our coverage rather than assume it.


M̄ = 0 means no affiliate monetization was detected — not that a publisher earns nothing

Monetization exposure () is the share of a publisher's ranked picks that carry a link we can establish is affiliate. When M̄ = 0 we do not score the publisher at all: with no commission on any pick, there is no conflict of interest for a Payout–Rank correlation to measure, and the card reads "Takes no affiliate revenue."

That statement is narrower than it sounds, and the narrowness is ours to disclose. This method observes one thing: outbound links on public pages, and whether those links carry affiliate attribution. It is blind to every other way a publisher can be paid — display advertising, sponsored placements paid outside an affiliate network, subscriptions, licensing, paid product placement, and commercial relationships disclosed nowhere in the page's links. A publisher that earns entirely from display advertising and takes no affiliate commission is indistinguishable, to this pipeline, from a publisher that earns nothing at all. We measure affiliate conflict, and only affiliate conflict.

Two publishers in the current corpus are examples. headphonehorizon.com (15 of 15 harvested articles) and standingdeskpicks.com (6 of 6) carry no outbound merchant links at all. We verified this rather than assumed it: the pages render fully, and repeated renders — including one waiting twenty additional seconds for product links to appear, and one using an ordinary browser user-agent to check whether the sites treat our crawler differently — returned identical results. Both sites publish ranked "best of" articles with no affiliate links in them.

What we can say about those two publishers is that we detected no affiliate monetization in the articles we sampled, and we will say how many articles that was. What we cannot say is that they are unmonetized, and we will not imply it. Where a publisher's M̄ = 0 result rests on a small sample, the card states the sample size next to the claim.


Part of one publisher's monetization figure rests on a signal we had to execute code to see

Added 2026-09-07 with SPEC-scoring.md v0.6.9 §0.1.0. It is the first signal in this system with this property, and it is disclosed rather than absorbed.

A publisher can mark a link rel="sponsored" — the attribute the HTML specification defines for paid placement, and what a publisher sets to satisfy 16 CFR 255 — while the link points at a merchant's own site with no affiliate mechanism visible in the URL. We do not badge those links and never will: rel="sponsored" is also the correct attribute for paid placement that pays no commission, so badging on it would mark a publisher for complying with the law.

But on some of those links the publisher's own client-side code, running in our browser, writes an affiliate attribution subtag onto the anchor. On nymag.com that is data-affiliate-subtag, written by its own ensureSubtag() routine, which sets the literal string "undefined" when it recognises no affiliate program for the destination and a real identifier when it recognises one. That is the publisher's own software stating whether it believes the link is in a program.

We count such a link as monetized and we still do not badge it. The asymmetry is deliberate:

  • rel="sponsored" is in the bytes the publisher served. Anyone can fetch the page and see it.
  • The subtag exists only because we executed the publisher's JavaScript in our own browser. It is a fact about what their code did in our render. A third party checking it must re-render the page rather than read the stored HTML, and a render is a thing that can differ.

Weaker provenance is admitted to the weaker claim — a corpus statistic we publish with its inputs and its date — and refused the stronger one, which is an assertion made to a reader about one link in front of them. That stronger claim is governed by precision-over-recall, where one false accusation is unrecoverable.

Every card that uses this signal publishes both figures. and counting badged picks only appear side by side, with the count of picks that used the weaker basis, so a reader can subtract it. On nymag.com today: = 0.886, and 0.744 counting badged picks only, from 66 of 419 picks.

The literal "undefined" is counted as NOT monetized, and separately from links carrying no such attribute at all. Those are two different facts — the publisher's code ran and found no program, versus nothing was observed either way — and only the first is a positive observation.

What this cost. Recording these picks as monetized-with-an-unknown-rate removed them from the correlation and lowered the affected articles' rate coverage. nymag.com's qualifying article count fell from 35 to 26 and its rate coverage from 0.795 to 0.676. The amendment bought monetization coverage and spent rate coverage, and both figures are published.


Some merchants' published commission rates cannot be read, and that is their configuration

This is the same rule as the Wirecutter case above, applied to the other party in the transaction. There, a publisher's robots.txt disallowed the redirect endpoint its own outbound links run through. Here, a merchant's robots.txt disallows the page on which it publishes its commission rate. Our obligation does not change with which company set the directive.

B&H Photo Video's robots.txt disallows the path its affiliate page lives on. As of 2026-09-06, bhphotovideo.com/robots.txt returns HTTP 200 and contains Disallow: /find/shared/ under User-agent: *. B&H's affiliate-recruitment page is at /find/shared/affiliates.jsp — B&H's own 404 page footer links there. A crawler that respects robots.txt therefore cannot read it, and we do not have B&H's commission rate.

The consequence, stated exactly: bhphotovideo.com carries no rate in our cards. That is not a statement that B&H pays nothing, and it must not be read as one. It is a statement that the page on which B&H publishes what it pays is one we are not permitted to fetch.

What this costs, measured rather than asserted. In the current corpus, 129 ranked picks link to bhphotovideo.com, across 27 articles and 3 publishers. Of those, 124 picks — across 24 articles and 2 publishers (nymag.com, rtings.com) — are picks we badged as monetized and in-market. The remaining 5, all on cnet.com, were not badged. Both numbers are given with their own denominators because they answer different questions, and welding one count to the other's article and publisher totals would overstate the reach of the monetized figure.

None of those 124 picks is currently left without a rate, because every one of them also links to at least one merchant we can price. What the missing B&H rate changes is not whether those picks have an r_i, but which merchants that r_i was minimised over: SPEC §0.1.2 sets a multi-merchant pick's rate to the minimum across its merchants, and a merchant we cannot price is a merchant absent from that minimum. If B&H pays less than the merchant that currently sets the floor, our recorded r_i for those picks is too high. We state the direction of that error rather than estimating its size.

We followed a redirect into that disallowed page once, and we discarded what it returned. bhphotovideo.com/find/affiliates.jsp is robots-allowed, returns HTTP 200, and 302s to the disallowed /find/shared/affiliates.jsp. Our rate probe checked robots.txt against the URL it started from and not against the URL it ended on, so the redirect was followed and a page we are not permitted to fetch was rendered. The content is not recorded anywhere in this repository and no rate was taken from it: a rate whose provenance is a URL we may not crawl is not a rate we may publish. The tooling has since been changed so that robots.txt is re-evaluated before every request on every redirect hop, in the redirect resolver and in the harvester alike — an allowed entry URL never legitimises a disallowed destination.

A rate for B&H is obtainable only by asking B&H. We have not made a deliberate, documented decision that an allowed entry URL licenses reading the page behind it, and we do not intend to make one quietly.


Some intermediaries choose the affiliate network AFTER the click, so no one rate card applies

Recorded 2026-09-08. SPEC-scoring.md §0.1.4 and §6 confound 8.

A commission rate is a fact about a relationship between a publisher, a network and a merchant. Two intermediaries in this corpus break that chain by design: they do not carry one relationship at all, but decide per click which network to route the reader through.

NucleusLinks, operated by Affinity Global Inc., is the intermediary behind ncls1.com, through which nymag.com routes 1,030 of the product links in this corpus. Its own published product pages, fetched 2026-09-08 with our own user agent under a robots.txt that permits them, describe two mechanisms. Smart Redirects (nucleuslinks.ai/smart-redirects) is "an AI-powered, multi-dimensional decision engine that dynamically redirects each merchant click to the highest EPC-yielding network", and states that it is "Pre-integrated with 70+ networks". Network+ (nucleuslinks.ai/network-plus) "creates an auction-style environment where primary and sub-networks (CPC & CPA) compete to maximize affiliate revenue" and "automatically optimizes click allocation daily, directing traffic to the highest-paying network".

Geniuslink, operated by GeoRiot Networks, Inc. (geniuslink.com), is the intermediary behind target.georiot.com and buy.geni.us. Its published purpose is narrower — it localises an Amazon link to the reader's own country's store — but it has the same consequence for a rate: amazon.com and amazon.ca publish different schedules, and which one a given click reaches is decided after the click.

These are the companies' own descriptions of their own products, and that is a claim by an interested party, not an observation of any link. We say so here in the same words we said it to the rater who screened these hosts. We have never followed a link on ncls1.com — its robots.txt is an explicit User-agent: * / Disallow: / — so we have not seen a routing decision and do not assert that any particular reader's click went anywhere in particular.

What it does to r_i, exactly. A link through such an intermediary gets no rate. There is no single published schedule that governs the click, and every network that might govern it keeps its rates private. Where the intermediary's own URL declares a merchant that publishes a minimum, that floor is used and the pick is flagged dynamic_routing_floor — a merchant's floor is a fact about the merchant, and routing among ways of being paid for the same sale can only raise what is earned above it, never lower it. That figure is therefore a lower bound, in the direction that flatters the publisher.

It is a rule about a LINK, not about a ranked product. A pick carrying one routed link and one ordinary priced link keeps the rate we hold for the second, cited as always. On this corpus 121 of the 203 routed picks are exactly that shape.

Measured cost, and it is smaller than it sounds. All 203 routed picks are on nymag.com. Of those, 121 price from another link and 82 do not. Re-computing that publisher's card with this rule switched off produces identical published numbers — the same 17 qualifying articles, the same median, the same , the same rate coverage — because the merchants those 82 links declare are ones we hold no rate card for either way. The rule changed the stated reason for a null, not the value of one.

The sign of the bias is indeterminate and we do not correct for it. Dropping these picks removes observations whose true rank-payout relationship we do not know; they could sit anywhere. Where an auction operates, the rate a publisher earns is a function of what other bidders offered on that day, which is neither published nor stable — so even a rate we could observe today would not be the rate that applied yesterday.


A geo-targeting redirect shows us a different destination than it shows a US reader

Recorded 2026-09-07. Evidence in docs/RESOLVER-SCOPE.md §3–§4.

Some publishers route outbound product links through a geo-targeting redirector, which sends each reader to their own country's storefront. tomsguide.com uses one — GeoRiot / target.georiot.com — on 505 links in our corpus.

Our crawl egress is in Canada. So when we resolve one of those links we are shown the Canadian destination. Measured on a 10-URL sample: a page whose own markup declared amazon.com/…?tag=ftr-tomsguide-us-20 redirected our request to www.amazon.ca/…?tag=ftr-tomsguide-ca-20. The redirect substituted the storefront and the publisher's Associates tag. A reader in the United States following that same link reaches the US storefront under the US tag.

amazon.com and amazon.ca publish different commission schedules. So:

  • Which rate applies to such a link depends on the reader's country, and we cannot observe the reader. No r_i we publish for a geo-targeted link can be correct for every reader of it.
  • We do not let our own vantage point decide. Where the page declares its own destination — GeoRiot carries it in a GR_URL parameter — that declaration is what we price against, because it does not change with who is looking. Our observed destination is used only where the page declares nothing, which is the plain-shortener case (amzn.to).
  • The confound is not removed by that choice, only kept out of the rate join. It is stated here rather than estimated.

This is the same class of limitation as the merchant rate pages we cannot read from a Canadian egress (above): a fact about where we fetch from, deliberately never reported as a fact about the publisher or the merchant.


What our gate does NOT measure: recall, and what a missed link does to the correlation

Recorded 2026-09-07. Protocol in docs/LABELING.md §7–§8.

Our published gate is a precision gate. It asks whether the links we badge as affiliate really are affiliate, and it is set brutally high on purpose — 0.99 on the top tier, and a single false positive anywhere in the negative control is a hard stop. That asymmetry is deliberate: a false accusation costs a publisher its reputation and cannot be taken back, and a missed link costs us one link.

Precision cannot see a missed link at all. So the same run reports recall — of the links an independent rater labeled affiliate, the share we badge — and the count of missed ones, and both belong beside the gate rather than in a footnote:

Labels the figure rests onRecallAffiliate links missed
independent (human + AI rater, 92 links)0.96743
human alone (15 links)1.00000
AI rater alone (77 links)0.96103

Updated 2026-09-07 after a ruling and a fresh 40-link sample. Recall was 0.9429 on 70 links with 4 misses. Two of those four are now detected — a publisher's own redirect endpoint is recognised by the DOMAIN it sits on rather than by whether its URL path matched a list of shapes we had seen before. The other two, and one newly found, are all the same host and are described below.

A miss can also move between the two tiers we do not badge, so we count them apart and never fold one into the other. Otherwise the missed-link count could improve because a link changed tiers rather than because we found it, which is the opposite of what this table is for.

The evaluation sample is stratified, deliberately over-weighted toward the places a missed link would hide. So these are per-stratum evidence about where the classifier fails, not an estimate of how often it fails across the corpus, and we do not present them as one.

A missed link is not a harmless omission — it enters the correlation as a rate of zero. SPEC §0.1 gives an unmonetized pick r_i = 0, which is the right rule: a publisher who earns nothing on a pick earns nothing, and that is a known rate rather than a missing one. But when we fail to detect a pick's affiliate link, that pick is recorded as unmonetized and enters the Payout–Rank correlation at r_i = 0 as though we knew the publisher earned nothing. We did not know that. We missed it.

The direction of the resulting error is indeterminate, and we will not claim otherwise. A missed link on a top-ranked pick puts a spurious zero where the publisher's highest prominence is, which pushes ρ down and flatters the publisher. A missed link on a bottom-ranked pick puts that zero at low prominence, which pushes ρ up and reads against them. Which one happens depends on where in each list the miss falls, and a recall figure says nothing about that. So we state the mechanism and the count rather than a correction: there is no adjustment we could apply that would not be an invention.

What we can say is exactly what the missed links currently are, because we publish them. Every one is a link we were not permitted to look at. Not one is a link we looked at and misjudged. Two hosts are involved and the two reasons are different and are not merged:

  • ncls1.com — 3 links IN THE LABELED EVAL SAMPLE, 1,030 links in the corpus. Both numbers are true of different things and this sentence now says which. The 3 is the count a rater judged and is the count that appears in the recall figure above; the 1,030 is how many links on nymag.com route through this host in total. An earlier version of this document said only "3 links", which reads as a claim about the corpus and is wrong by a factor of 343. An unrecognised commerce redirector. Its robots.txt returns HTTP 200 carrying User-agent: * / Disallow: / — an explicit blanket refusal of the whole host. We asked for permission, were refused, and stopped. No destination was observed, so nothing about it is badged. All three of the current misses are this host.
  • cc.cnet.com, 2 links on cnet.com. CNET's own outbound redirector. Its robots.txt cannot be read at all: the request answers 403 behind a bot challenge, and our own conduct rule treats an unreadable robots.txt as a refusal. That is our choice, not the publisher's instruction, and we say so rather than reporting it as though CNET had told us no.

That is still exactly true, and these two links are no longer missed. Nothing about the refusal changed and we still have not observed where they go. What changed is that we no longer need to: cc.cnet.com is a subdomain of cnet.com, the site the links were published on, and CNET marks them rel="sponsored" in its own markup. A publisher saying "this hop through my own infrastructure is paid" answers whether the link is monetized without answering where it goes. We now record the first and still decline to guess at the second.

Being unable to look is a better position than looking and getting it wrong. It is still three picks that would enter a correlation at a rate we do not actually know.


Most of our ground-truth labels were written by an AI, and none of them have been checked yet

Recorded 2026-09-07. Protocol in docs/LABELING.md §1.1.

Everything we publish rests on a classifier, and the only way to know whether a classifier is right is to compare it against links that somebody labeled independently. Those labels are our ground truth, and who produced them is part of the finding, not a footnote.

Some are produced by an LLM rater. It is a different process from the classifier, and it is instructed and observed rather than blind — a distinction we make deliberately, because it is the one we can support. What we can state as fact is what we gave it and what it did: its brief contains no classifier code, no network signature table, no resolutions and no tier, so nothing from the classifier reached it through us; it was shown exactly what a human labeler is shown — the url, the anchor text, the rel attribute, the page the link appeared on and that page's title, and whether that page rewrote its links at runtime; it was instructed never to fetch or follow any link, because following an affiliate link registers a click and credits a publisher for traffic that never happened, and we observed that it made no requests. What we cannot state is what a model already knew about these hosts before we asked. A person can be kept blind to a prediction; a model cannot be emptied of what it has read. Calling it "blind" would claim a property we cannot verify, so we describe the procedure instead. Every label it produces carries a one-line rationale citing the evidence in the url or the context that decided it, and every row is stamped as the AI's rather than a person's.

As of 2026-09-07, the gate rests on 135 independent labels: 25 written by a person and 110 written by the LLM rater. A human has audited 0 of those 110. The audit tool is built and has not been run. When it runs, this paragraph will say how many were checked and how often the person disagreed; until then the number is zero and zero is what we publish.

Note which way that number moved. It was 0 of 70. We then added 40 more AI labels, so the same unrun audit now covers a smaller share of the evidence than it did before. Collecting more AI labels makes this figure worse, not better, and we report it in the direction it actually went rather than reporting the raw count of labels as though it were progress.

A second AI has checked twelve of them, and that is a different number with a different name. A different model was shown the same twelve links, given the same instructions, and never shown the first rater's answers. It agreed on twelve out of twelve. We report the two figures side by side and never add them together:

What it meansValue
K_humanagent labels a PERSON has checked0 of 110
K_agent2agent labels a SECOND AI reached the same verdict on, blind12 of 12 agreed

K_agent2 is not a weaker version of K_human. It is a different measurement. It bounds one specific failure — one model's idiosyncratic reading of a link going unnoticed — and it bounds nothing else. Two models can be wrong in the same way, and on a rule they were both handed in the same brief they are especially likely to be. A perfect score there is not evidence that a person would agree, and we will not let it stand in for a person agreeing. No label was changed by it in either direction: a second rater is a measurement, not a correction.

We report precision and recall four ways, never blended into one: over the human labels alone, over the AI's labels alone, over the two combined, and over everything including the mechanical pre-pass that transcribes publishers' own declarations. A single averaged figure would hide the thing a reader most needs, which is how much of the evidence a person actually looked at.

What this does and does not buy. An AI label is independent of the classifier, which is the property that makes it worth measuring against at all — a label derived from the classifier's own rules would only prove the rules match themselves. It is not a person's judgment, and we do not present it as one. On the one slice where both exist — ten links from our negative control, labeled by a person in September and by the AI afterwards without being shown that person's answer — the two agreed on ten out of ten. That is a small and encouraging number, and it is a sample of ten.

It has already been visibly imperfect, in the direction that matters least and we are saying so anyway. On six links the rater called a publisher's own redirect endpoint an affiliate link where our written protocol says to record it as unknown, because the destination cannot be seen without following it. The rater's reading is defensible; it is not the protocol. We left its labels exactly as it gave them, counted the resulting mismatch against the classifier as a miss rather than quietly reclassifying it, and put those six at the top of the audit queue. We have since concluded the rater was right and changed our rule to match, which is a fact about our rule and not a reason to trust the rater more: a rater that reads a rule differently from us is exactly as likely to be wrong the next time.

And an AI rater has now overturned one of our own decisions in the other direction. Five links had been classified as monetized because their URL path looked like a publisher's redirect endpoint; on the ruling above they were reclassified as not monetized, on the ground that they point at a merchant's own website rather than the publisher's. Nobody had ever labeled them. We put all five in front of the rater as a deliberate test of the reclassification, and it called all five not affiliate — which means that, before the change, five links we badged as monetized were wrong, and our own precision gate would have failed on the next measurement. We publish that because a rule change that removes a badge deserves the same scrutiny as one that adds one, and because the alternative was finding out from somebody else.

Every disagreement of this kind moves the number toward a missed detection, never toward a false accusation. A missed detection costs us a link. A false accusation costs a publisher their reputation, and no measurement convenience is worth that.


Our k-anonymity claim did not hold, so we removed the need for it

Added 2026-09-07, resolved 2026-09-08. This is a limitation of our service, not of a publisher's configuration, and it is on this page for the same reason all the others are.

What we measured, and it is unchanged

When a link cannot be classified on your device, the original design sent the first five hexadecimal characters of the SHA-256 of the URL and received every resolution whose hash starts with those characters, filtering locally. That is the HaveIBeenPwned model, and its privacy rests entirely on the bucket being crowded: the server is meant to be unable to tell which of many URLs you meant.

We bucketed the 1,147 resolutions we hold. There are 1,146 distinct five-character prefixes, and 1,145 of them contain exactly one URL.

At this corpus size the crowd is one person. The mechanism was built and behaved correctly — five characters sent, no URL stored, nothing logged, a longer prefix refused outright — but the property those five characters exist to provide had not arrived. k is a function of how many URLs share a prefix, roughly N / 1,048,576 for a table of N rows, so it arrives with scale and with nothing else.

What we did about it, which is not what we said we would do

We said the property "arrives at scale". Waiting for that would have meant running a crowd-of-one lookup on the live path for as long as it took.

Instead the whole table ships to your browser. All 1,147 resolutions travel inside the rules bundle the extension already downloads — hash-keyed, sha256(url) -> destination, about 200 KiB inside a 533 KiB file. A cloaked link is now answered by a hash and a map lookup on your own device.

  • There is no URL in the file. A URL you are looking at can be looked up; the list of URLs cannot be read out of it. That is a stronger position than the lookup endpoint, whose buckets return the URL in plaintext.
  • A row with no destination still carries a reasonrobots_disallowed, blocked, not_a_redirect. "We looked and were refused" and "we never looked" are different findings and neither is reported as the other.
  • The extension makes zero network calls. Not "zero by default". The bundle carries a flag, resolutions_truncated; it goes true only if the table ever exceeds 2 MiB, and only then may a client call the lookup endpoint. It is false, so the browser is granted no host permission with which to make the call, in any build.

What is still true, and what we are not claiming

The lookup endpoint still exists, live and documented, for the future in which the table no longer fits in a bundle. Its own response says, in its first sentence, that no current client calls it, and it publishes the measured crowd size beside the mechanism — because on the day it becomes live again, k will be whatever the table size makes it, and that is a number rather than a promise.

We are not claiming this makes us unable to learn what you browse. We are claiming there is no request in which we could. The bundle download is the same file for every reader and says nothing about any of them.