"Fresh evidence for this exact change" is the slippery half of it, because freshness splits in two and a gate usually only checks one.
I ran into the other half today, on a public repo that exists to test exactly this (dannwaneri/rules-demo-api, PR #2). The rule lives in AGENTS.md and CLAUDE.md; the PR moved it seven lines down in both files, text unchanged, then forced a re-run of the reviewer. By any reading of subject / producer / coverage, that re-run is genuinely fresh: its timestamp moved, and its citations were re-pinned to the new head 3ffacd90 — a commit that did not exist when the original review ran.
The numbers inside those citations did not move. The label still reads AGENTS.md[7-10] and src/index.ts[33], and at 3ffacd90 both ranges are wrong: lines 7–10 are rules 1, 2, 3 and a blank line, and the cited rule is on line 17. The URL was rebuilt per run (new SHA in the blob path) while the range it points at kept its old line numbers — so the rebuild step exists and simply skips the range. The platform's own diff anchor for that same finding was recomputed to line 34 while the body still says 33: two citations for one finding, one recomputed and one frozen.
That is why I'd file it under the policy clause rather than the subject clause. You wrote that the pull request must not be able to quietly weaken the policy — and here the PR didn't weaken the text, it moved it. The only artifact that could have surfaced the relocation is the one that stored its answer once, at review time. A merge gate that selects checks by name cannot see a doc relocation, so the citation is the sole witness, and it is rendered from a stored label rather than from the retrieval result.
Two cheap changes fall out of that: render the citation at read time from the thing you actually retrieved, and prefer an anchor that can't drift — a rule id or an anchor string — over a line range. It's the same choice your [[surfaces]] patterns already make: paths are stable, positions aren't.
The reason this stays invisible is worth naming: "kebab-case" is prose. Nothing downstream can check whether a prose citation is still true. Had the rule been something a gate could evaluate on its own — the literal ^/[a-z0-9-]+(/[a-z0-9-]+)*$ — the same re-run would have had a way to fail.
Good distinction: the run is fresh, but the evidence inside it is stale. I checked the PR—the review is pinned to 3ffacd9, but its cited ranges no longer match the files.
CI Evidence Gate verifies the outer receipt, so it wouldn’t catch this today. Stable rule IDs/anchors plus line ranges regenerated from the retrieved blob feels like the right fix. This deserves its own failure pattern.
I also updated the CI Evidence Gate docs to make the boundary explicit: it verifies the outer receipt, not whether citations inside a producer's output are still accurate. Thanks again for such a concrete repro. It made the pattern a lot better.
For further actions, you may consider blocking this person and/or reporting abuse
We're a place where coders share, stay up-to-date and grow their careers.
"Fresh evidence for this exact change" is the slippery half of it, because freshness splits in two and a gate usually only checks one.
I ran into the other half today, on a public repo that exists to test exactly this (
dannwaneri/rules-demo-api, PR #2). The rule lives inAGENTS.mdandCLAUDE.md; the PR moved it seven lines down in both files, text unchanged, then forced a re-run of the reviewer. By any reading of subject / producer / coverage, that re-run is genuinely fresh: its timestamp moved, and its citations were re-pinned to the new head3ffacd90— a commit that did not exist when the original review ran.The numbers inside those citations did not move. The label still reads
AGENTS.md[7-10]andsrc/index.ts[33], and at3ffacd90both ranges are wrong: lines 7–10 are rules 1, 2, 3 and a blank line, and the cited rule is on line 17. The URL was rebuilt per run (new SHA in the blob path) while the range it points at kept its old line numbers — so the rebuild step exists and simply skips the range. The platform's own diff anchor for that same finding was recomputed to line 34 while the body still says 33: two citations for one finding, one recomputed and one frozen.That is why I'd file it under the policy clause rather than the subject clause. You wrote that the pull request must not be able to quietly weaken the policy — and here the PR didn't weaken the text, it moved it. The only artifact that could have surfaced the relocation is the one that stored its answer once, at review time. A merge gate that selects checks by name cannot see a doc relocation, so the citation is the sole witness, and it is rendered from a stored label rather than from the retrieval result.
Two cheap changes fall out of that: render the citation at read time from the thing you actually retrieved, and prefer an anchor that can't drift — a rule id or an anchor string — over a line range. It's the same choice your
[[surfaces]]patterns already make: paths are stable, positions aren't.The reason this stays invisible is worth naming: "kebab-case" is prose. Nothing downstream can check whether a prose citation is still true. Had the rule been something a gate could evaluate on its own — the literal
^/[a-z0-9-]+(/[a-z0-9-]+)*$— the same re-run would have had a way to fail.Good distinction: the run is fresh, but the evidence inside it is stale. I checked the PR—the review is pinned to
3ffacd9, but its cited ranges no longer match the files.CI Evidence Gate verifies the outer receipt, so it wouldn’t catch this today. Stable rule IDs/anchors plus line ranges regenerated from the retrieved blob feels like the right fix. This deserves its own failure pattern.
Quick follow-up: I turned your example into an executable pattern in Fleet Failure Atlas: FFA-005: Fresh review, stale citation.
I also updated the CI Evidence Gate docs to make the boundary explicit: it verifies the outer receipt, not whether citations inside a producer's output are still accurate. Thanks again for such a concrete repro. It made the pattern a lot better.