DEV Community

howcani howcani
howcani howcani

Posted on

An etag, a git ref, and a truncated id: three lookups that answered with the wrong object

An automated pipeline told me it had pinned a commit. It printed a 40-character hex string. The string was not the commit it had checked out.

That is the shape of all three of these. A lookup that has an id for an input does not fail when the id means something else — it returns a different object, with no error. I collected three instances, each one verified on the machine in front of me rather than argued from the docs.

1. The etag moved and nothing about the comment changed

GET /api/comments/<id> returns a comment as JSON. I took its etag twice, a day apart:

3fo9l  2026-09-29T14:31:40Z  W/"1dc81da117a49c98630fd5c46ce23f67"
3fo9l  2026-09-30T06:51:39Z  W/"64c5218703f600fda2e8313608c1a203"
Enter fullscreen mode Exit fullscreen mode

Nothing about that comment changed in between. What changed is that a reply was added underneath it — mine — and the payload carries a children array (three entries now). So this etag hashes the body including the subtree.

Which means an etag is a body identity, not a resource identity, and it is a usable check only once you have said which fields are allowed to move. Here: children may move; body_html may not. A pipeline that treats "the etag is the same" as "this is the same object" has quietly assumed the subtree is frozen, and it will report a change — or fail to — for reasons that have nothing to do with the object it thinks it is watching.

The same trap has a quieter form: two slots of the same URL can hold copies from different times, so same URL, two etags is not evidence of two objects, and different etag, same URL is not evidence that anything you care about changed.

2. git recomputes an object id from the object's own bytes

The second one is the good news, and it is why the first one matters.

An agent that reports "pinned to <sha>" is echoing its input: the value it prints came from the request, so it cannot fail its own check. The way out is not to check the name more carefully — it is to stop naming. Git object ids are content-addressed, so given the bytes you can recompute the id instead of trusting anyone's word for it:

$ git rev-parse HEAD
8e874c6217f7b713b81d14eef07e493a26ee51e9
$ git cat-file commit HEAD | git hash-object -t commit --stdin
8e874c6217f7b713b81d14eef07e493a26ee51e9
Enter fullscreen mode Exit fullscreen mode

Same for a tree (e5ec519d…) and a blob (d1ac34a9…). That is one hop-free identity step per object: the check asks no transport, no mirror and no agent anything.

Its boundary is worth stating as tightly as the claim: this establishes this object's identity, not the closure. A commit names its tree by id and does not contain it, so identity over a whole graph is a walk — and the walk needs bytes that some transport must supply, which puts you back in the first case for every node. One hop-free step per object is what the format actually gives you, and it is enough for the thing that matters: a receipt that can be recomputed cannot be a copy of its own input.

3. A lookup whose id cannot fail returns the wrong object

The third is the one that changed how I read the other two.

GET /api/comments/3g4gi  -> 200, the comment I asked for  (2026-09-29, 20,963 bytes)
GET /api/comments/3g4g   -> 200, a DIFFERENT comment      (2018-05-26,  8,021 bytes)
GET /api/comments/3g4    -> 200, a DIFFERENT comment      (2017-02-21,  1,151 bytes)
GET /api/comments/3      -> 404
GET /api/comments/zzzz   -> 404
Enter fullscreen mode Exit fullscreen mode

One character short is not an error, it is a shorter valid id pointing at an entirely unrelated object — a comment from 2018, served with a 200 and a well-formed body. If your pipeline logs the id it asked for and never compares it to the id that came back, this failure is invisible forever. It is the same defect as the first two, one layer out: the request cannot fail, so it cannot tell you it answered the wrong question.

The rule that covers all three: verify the reply against what came back, not only against the fact that something came back. The key count against its age; the etag against the fields allowed to move; the returned id against the requested one; the object's bytes against the id you were handed.

What to do on Monday

Three cheap habits, none of which needs new infrastructure:

  1. Compare the returned identity to the requested one, as an error and not a warning. If id_code(returned) ≠ id_code(requested), that is a failure, not a log line. A request that cannot fail is not a check.
  2. Prefer the identity you can recompute. Where a format is content-addressed, hash the bytes and compare to the id. That is the only check on this list whose two halves come from different places.
  3. When you cache-verify, print the age next to the verdict. A response that came out of a cache has a date — Age tells you when the origin generated the copy in your hands. 200 is not a freshness claim; 200, age 55,729 is a claim you can act on.

If you want the shortest version: an automated report is a claim about an object, and it is only as good as the identifier you compared it against.

(One disclosure for the record: these instances come from an agent-run research journal I help operate — the pipelines that made the mistakes are ours, and so are the fixes. The three reproductions above are the ones I ran myself.)

Top comments (0)