When I look at a pull request, I do not really care that a job named test is green. I care that the right tests ran, on the right commit, under a p...
For further actions, you may consider blocking this person and/or reporting abuse
The four-question split is the useful bit, because separating subject from producer is what makes a rerun answerable at all.
On the no-op item in your "does not prove" list: I hit that in a shape the gate could never catch, because the skip decision lived inside the test process. Every live suite opened with
const run = process.env.API_KEY ? describe : describe.skip, so with no secret vitest exits 0 and the only trace is "2 skipped" in a log nobody opens. Green meant "the contract holds" and "nothing was checked" identically.Moving the guard out into its own job fixed it, because then the skip arm shows up in the job list, where a name-based required check can see it. Does that sit inside your producer-identity question, or is it a fifth one?
I'd keep it under producer identity, with one important boundary. Once the applicability decision becomes its own named check, the gate can require it. If
describe.skipstays inside Vitest, CI Evidence Gate cannot see it today.The clean bridge is a separate required check over structured test output, for example
expected > 0andskipped == 0. And yes, a missing secret is evidence unavailable, not "not applicable."This also exposed a gap in FFA-001: the write-up mentions green no-op jobs, but the executable fixture only covers stale SHA binding. I'll split the no-op case into its own runnable pattern. Thanks for the sharp example.
Producer identity — and I think the reason it feels like it might be a fifth question is worth naming.
Your original shape had a producing job with an identifiable name (
test) and a green conclusion; what had no identity anywhere was the decision that the suite applied. So the receipt was complete about the producer it names and silent about the actor that scoped it. Same layer as "which workflow and concrete job produced this result", just with two roles in one process — runner and applicability judge — where only one of them reaches the check-run namespace. The article's own "does not prove" list parks this as "an evidence-producing job did meaningful work instead of a no-op", which reads like an evidence-layer limit; I'd argue it is a provenance one, because a no-op is only invisible while the actor that chose it cannot be named.What your fix actually does is make the decision nameable, and the name is the currency here: required checks select by name, so a decision that never gets one can never be required, forbidden, or read as missing. Worth being precise about the residual hole, though — selection sees a job, not what the job did. A job that shells out and swallows a non-zero exit is green under an excellent name. The split upgrades a silent decision into a checkable one; it does not audit the shell.
If you want the whole family covered with one rule instead of one job per guard, evaluate the run's structured report in a separate required check: assert
expected > 0andskipped == 0over the required set from--reporter=json. That catchesdescribe.skip, at.skip()inside a case, an earlyreturninbeforeAll, and a config-level ignore decision — none of which a job split necessarily catches — and it never asks the test process to grade itself.One polarity note, because it is the cheapest version of the fix: an absent secret is evidence unavailable, not evidence not applicable. Written the other way round — skip only when the secret exists but is malformed, fail when it is missing — the trap does not exist and no new job is needed. The job split is what the genuinely inapplicable case needs, where the right output is a recorded
not_applicablewith a reason rather than a green check.Agreed — and the failure that eroded my trust hardest wasn't a fake green, it was a flaky test I marked retry-tolerant. It stayed green so long that every real failure started looking like noise.
The most useful change wasn't a smarter suite: it was pinning the check to the exact commit SHA and having the job declare which paths it actually exercised, so 'green on a matrix that ran nothing' still stands out. Anything that retries now gets a loud label in the job name. What triggered the post for you — a green job that ran nothing, or a name/workflow mismatch?
Closer to the second one, though it wasn't a single production incident. While building a synthetic fixture, I saw how easily a successful receipt could be accepted for the wrong SHA. “The check passed” is not the same as “this is fresh evidence for this exact change.”
Your retry-tolerant case is the same problem from another angle. I like the loud label — I'd also record the retry count so pass-after-retry doesn't get flattened into plain green.
"Fresh evidence for this exact change" is the slippery half of it, because freshness splits in two and a gate usually only checks one.
I ran into the other half today, on a public repo that exists to test exactly this (
dannwaneri/rules-demo-api, PR #2). The rule lives inAGENTS.mdandCLAUDE.md; the PR moved it seven lines down in both files, text unchanged, then forced a re-run of the reviewer. By any reading of subject / producer / coverage, that re-run is genuinely fresh: its timestamp moved, and its citations were re-pinned to the new head3ffacd90— a commit that did not exist when the original review ran.The numbers inside those citations did not move. The label still reads
AGENTS.md[7-10]andsrc/index.ts[33], and at3ffacd90both ranges are wrong: lines 7–10 are rules 1, 2, 3 and a blank line, and the cited rule is on line 17. The URL was rebuilt per run (new SHA in the blob path) while the range it points at kept its old line numbers — so the rebuild step exists and simply skips the range. The platform's own diff anchor for that same finding was recomputed to line 34 while the body still says 33: two citations for one finding, one recomputed and one frozen.That is why I'd file it under the policy clause rather than the subject clause. You wrote that the pull request must not be able to quietly weaken the policy — and here the PR didn't weaken the text, it moved it. The only artifact that could have surfaced the relocation is the one that stored its answer once, at review time. A merge gate that selects checks by name cannot see a doc relocation, so the citation is the sole witness, and it is rendered from a stored label rather than from the retrieval result.
Two cheap changes fall out of that: render the citation at read time from the thing you actually retrieved, and prefer an anchor that can't drift — a rule id or an anchor string — over a line range. It's the same choice your
[[surfaces]]patterns already make: paths are stable, positions aren't.The reason this stays invisible is worth naming: "kebab-case" is prose. Nothing downstream can check whether a prose citation is still true. Had the rule been something a gate could evaluate on its own — the literal
^/[a-z0-9-]+(/[a-z0-9-]+)*$— the same re-run would have had a way to fail.Good distinction: the run is fresh, but the evidence inside it is stale. I checked the PR—the review is pinned to
3ffacd9, but its cited ranges no longer match the files.CI Evidence Gate verifies the outer receipt, so it wouldn’t catch this today. Stable rule IDs/anchors plus line ranges regenerated from the retrieved blob feels like the right fix. This deserves its own failure pattern.
Quick follow-up: I turned your example into an executable pattern in Fleet Failure Atlas: FFA-005: Fresh review, stale citation.
I also updated the CI Evidence Gate docs to make the boundary explicit: it verifies the outer receipt, not whether citations inside a producer's output are still accurate. Thanks again for such a concrete repro. It made the pattern a lot better.
The distinction between “green” and “bound to the right change” is the important part. A passing check is only useful evidence if we can establish exactly what it evaluated—and what it didn’t.