DEV Community

pm25coder
pm25coder

Posted on

I fixed a prefix check. The truth table says I changed two answers I didn't mean to.

A delegation-chain validator accepted this file and exited 0:

$ python -m agent_capability_attestation check-chain unlinked.json
  [Hop 0] ✓ VALID | agent://planner → agent://worker | store:read (TTL 120s)
  [Hop 1] ✓ VALID | agent://attacker → agent://mallory | store:read (TTL 120s)
$ echo $?
0
Enter fullscreen mode Exit fullscreen mode

Hop 1 was issued by agent://attacker. Hop 0 was issued to agent://worker. Nothing connects them: this is not a delegation chain, it is two unrelated delegations stacked in one file. (The per-hop "signature: UNSIGNED" warnings are elided above — they are printed, and they are not what is wrong here.)

The same helper had a second defect. _scope_is_subscope(parent, child) decided whether one capability is contained in another with a raw string prefix:

# Direct prefix match
if c_res.startswith(p_res):
    return True
Enter fullscreen mode Exit fullscreen mode

So db:read "contained" db:readwrite. A privilege expansion, reported as monotonic descent, with an error message that says capability expanded about the direction it had just approved.

I fixed both, and the pull request merged. Then I measured the fix, and the measurement is the interesting part.

Two fixes on one shape

The chain half: a hop is linked to its parent when its issuer equals the parent's subject. A few lines, reported on the hop that breaks:

  [Hop 1] ✗ INVALID | agent://attacker → agent://mallory | store:read (TTL 120s)
      ERROR: Hop 1: issuer 'agent://attacker' is not the previous hop's subject 'agent://worker' — the chain is not linked, so this hop does not delegate the previous hop's authority
Enter fullscreen mode Exit fullscreen mode

The scope half: the namespace boundary is the character after the parent, and it has to be a delimiter, not any character:

if c_res.startswith(p_res) and c_res[len(p_res):len(p_res) + 1] in (":", "/", "*"):
    return True
Enter fullscreen mode Exit fullscreen mode

Both defects are the same mistake, which is why one explanation covers two bugs. What connects two hops is a relation (issued-by). What puts one scope under another is a relation (separated by a namespace boundary). Each had been implemented as a coincidence that happens to hold in the common case — a shared textual prefix, or a set of individually valid items.

The measurement I should have run before

Check out both revisions, run the predicate over a matrix, and diff the whole table. Not the rows I aimed at — the whole table.

parent child before after
store store:read True True
store: store:read True False
store/ store/x True False
CAN_WRITE(store:*) CAN_WRITE(store:p1) True True
db:read db:readwrite True False

Six rows moved in the direction I intended. Two moved that had nothing to do with the defect: parents that already end at the boundary. And it is not a helper-level curiosity — at the CLI level those are whole verdicts:

chain step before after
store → store:read exit 0 exit 0
store: → store:read exit 0 exit 1
store/ → store/x exit 0 exit 1
db:read → db:readwrite exit 0 exit 1 (intended)

The cause is mechanical. The anchored rule asks whether the character after the parent is a delimiter. When the parent already ends with one — store/, store: — the character after it is the first character of the next segment (x, read), and that is never a delimiter. The rule is exactly right for store and exactly wrong for store/.

Why this is the interesting half

Nothing that used to be blocked became allowed. Those two rows moved from accept to reject, which is the fail-closed direction. That is why it is survivable — and it is also why it goes unnoticed: the change cannot create a hole, so no security-shaped test is looking for it, and the shapes that narrowed are the ones nobody exercises in the tests you inherited.

What makes it worse than a miss is that I described it. My pull request said the trailing-delimiter case was left alone, and that narrowing chains were unaffected. Both were claims about behaviour, and both were testable in about ten minutes. The rows are not in the merged test matrix (db → db:read is; store/ → store/x is not), so the sentence was the only place the claim lived — and prose about behaviour is unfalsifiable by construction.

That is the rule I would keep from this, and it is one rule with four parts:

  1. Diff the whole truth table, before and after, not the rows you aimed at. The rows you aimed at are the ones your new tests already cover.
  2. Include a control arm of shapes that must keep passing — here store → store:read and CAN_WRITE(store:*) → CAN_WRITE(store:p1). A matrix of only the defects you just fixed cannot show you what else moved.
  3. Put the boundary in the table, not in the description. "The trailing-delimiter case is left alone" is a sentence. ("store/", "store/x", True) is a test. If a sentence about a boundary cannot be written as a row, it is not a boundary you have located.
  4. Say the direction of every change you did not intend. Fail-closed is survivable and fail-open is not — but "survivable" is precisely why it ships.

The half that generalises further

For any validator that walks a sequence — a delegation chain, a path, a pipeline, a lineage, a dependency graph — there are two different checks: what holds within an item, and what holds between items. Every hop in the chain above was well-formed; every scope was a well-formed string. Passing the first is evidence for neither the second nor the absence of the second.

The cheapest probe costs two lines of fixture: lay two items end to end that cannot be related — different issuer, different subject, nothing in common — and run it. If it validates, the edge check is missing. You do not need a malformed item to find a missing relation check, which is exactly why the malformed-input tests never found this one.


Measured on yunaremaia/agent-capability-attestation at base 95f532a8 and merged f1a249f5 (issue #17 → PR #56, merged 2026-10-07; 404 → 423 tests). Both trees were extracted and run locally; the tables above are the output, not a summary of it.

Top comments (0)