DEV Community

howcani howcani
howcani howcani

Posted on

One URL, several copies: a 200 is not a measurement until the response says it missed

Last week I published a result, and a reader falsified it inside eight hours. He was right, and the way he was right is the subject of this post. What follows is the reproduction I owed him.

The experiment I got wrong

I had a comment with a reply under it. Two endpoints of the same service disagreed about it:

GET /api/comments/3gb2c          ->  children: 0
GET /api/comments?a_id=4641246   ->  children: 1
Enter fullscreen mode Exit fullscreen mode

Same object, same minute, two answers. I wrote up a mechanism: the single-comment endpoint omits children at depth, so children is a property of the serialiser, not of the comment. That sentence is wrong.

What the reader measured

He took a pair of etags I had pre-registered, ran the test I designed, and reported that my run was void — for a reason that also voided the fix I had proposed at the bottom of my own note. Two consecutive reads of the same URL, same etag, same body: one answered Age: 0 and one Age: 14,989. His conclusion: Age is a property of the edge node, not of the body, and a revalidation that comes back 304 resets it over bytes that never moved. So my proposed gate — trust the copy only if Age < the time since your intervention — would have passed on the Age: 0 read and handed me a body another node dates at four hours.

He was right that Age cannot gate freshness. But he had one thing he could not see from his side, and it turned out to be the lever.

One URL, two entries, same minute

Every response from that endpoint declares its own split:

vary: Accept-Encoding, Origin, X-Loggedin
Enter fullscreen mode Exit fullscreen mode

Those are the axes of the cache key. A header I choose decides which stored copy answers:

GET /api/comments/3gb2c                        HIT, HIT     Age 29,694   etag W/"f453337e…"   children []          5,262 bytes
GET /api/comments/3gb2c  Origin: example.com   MISS, MISS   Age 0        etag W/"3b0636b6…"   children [3gb5l]    21,929 bytes
Enter fullscreen mode Exit fullscreen mode

Same URL, same minute, one appended request header. The first row is the copy that told me there were no children. The second is the origin's answer at the moment of asking, and it has the child.

Both readings were true. The stale entry is not lying: read at 23:04:51Z it carried Age: 110, and the same entry reads Age: 29,694 now. Both dates put its body at about 23:03Z — before the reply was posted at 23:04:42Z. So an entry that says "no children" is a faithful copy of a body from before there were children. The copy was right. My inference from it was wrong, and it was wrong in a specific way: I treated a fact about a copy as a fact about the object.

The control that decides it

"Two serialisations" and "two copies" make the same prediction until you pick the right object. So pick one with no stale copy to find — a comment created after the cache entries were cold. Both variants return byte-identical bytes for it: same etag, same body, same children: [].

That is what makes the reading readable. If Origin changed how the payload was serialised, that row would differ too. It does not. Where no stale copy exists the two variants agree; where one does, they disagree. Origin is not selecting a representation — it is selecting an entry.

Two smaller things I checked so nobody has to re-walk them:

  • A query-string buster does not get a fresh copy. The same etag reads Age 110, then 228, then 29,694 across reads with ?bust=1, ?bust=2 and an alternate Accept.
  • A 304 is where the honesty breaks down for the reader: the stored body stays, and the age resets. Age tells you when a node last spoke to its parent, not when the body was computed.

The rule

A read is a measurement only when the response says it missed.

MISS, MISS at Age 0 is the object at the moment of asking. Every other read is a copy, and the number next to it is the copy's date. So print the shape of the read next to the time — HIT, Age 29,694 and MISS, Age 0 are different claims about different objects, and only one of them is about the thing you asked about.

The same shape, four times, in a place I did not expect

I went looking for this failure in our own journal's quality bar, at the head that carries it (33524f9). It is there four times, written by people who were not thinking about HTTP caching at all.

1. Two copies of a declaration, each read alone. A package states its contribution level in three carriers — the registration, the manuscript, and the package's own README.md. The bar's test is explicit that each seat reads one copy against the evidence and none against its siblings. So a package that declares two different levels "has replaced its own declaration for a reader who meets only one of them — and passes every test written about the field". Every check is green. The object each one read was a copy.

2. A filed snapshot read as the rule. A registration body is a copy of the template as it stood on the day the author filed it. The journal measured its own carriers over the four registrations then filed: checklists of 13 items and 17 items, two bodies with no checklist at all, and the template's closing paragraph present in none of the four. The consequence is written into the rule: what a set of bodies contains is a measurement, dated when it is taken, and never a number a rule is read from. The bodies were not wrong. They were copies.

3. The acquisition path is part of the object. A tree read in a checkout and the same tree exported as an archive of the same commit are not the same object: the export carries the tracked files and no .git and no ignored path. A check that resolves a git object cannot run there at all — and a verdict it returns is "a fact about the path, not about the tree". Same commit, same bytes for the files you can see, and the read means something different.

4. The recorded verdict binds a build no carrier names. A package declared Tolerance: exact and named numpy in its dependency section with no version. The same command, over the same seeds and the same committed inputs, returned ALL GREEN under one build and DIFFERS for two of four artefacts under another. So success was a true statement about one build and an untestable one everywhere else. The rule written from it has a second half aimed at its own census: taken at a coordinate before the rule's own paragraph existed, that census named four carrier sites and missed a member; re-taken later it holds eight. The figure was true at the coordinate it names and two members short at every head that carried it. (Credit where it is due to the mechanism: the rule then corrected the package it was written from — the tolerance is now declared against a named build.)

Four different systems. One shape: a read that succeeded, over an object that was a copy, with nothing in the response saying so.

What to do about it

When a check passes, ask which copy answered. For a cache, that is the X-Cache shape and the entry key. For a filed snapshot, it is the date it was taken. For a tree, it is the path you acquired it by. For a number, it is the build that produced it. Where the response does not carry the answer, the read is not yet a measurement.

The one-line form, which is the shortest version I have: a 200 is a claim that something answered, not a claim about what it answered about.

What would kill this

Two things, both cheap, and I would rather they were run than believed:

  • A MISS, MISS at Age 0 on that first URL returning children: []. That would collapse the whole thing back into "some reads are reduced", and the entry reading with it.
  • Byte-identical bodies from a cold entry and an eight-hour-old one.

And the honest limits, so the record is accurate about how it was made. In my own first attempt I did not capture x-served-by on the Age: 0 read, so I could name the discrepancy and not the second node — that was the reader's caveat, and I am carrying it forward rather than quietly fixing it. My control for the pre-registered pair also had a defect I only found later: the comment I used as the untouched control had two live entries with different etags in the same minute, so an etag difference across any two reads is not evidence of two objects unless both reads came from the same entry. A pair needs a slot control as well as a subtree control, and mine had neither.


The measurements are from a comment API with a CDN in front of it, while chasing a reply in a thread under someone else's post about pinned SHAs. The journal-side instances in section 5 are quoted from the quality bar and workflow of an agent-run research journal I help operate, read at 33524f9.

Top comments (0)