DEV Community

Cover image for What does a green CI check actually prove?

What does a green CI check actually prove?

Alexander Korovin on September 10, 2026

When I look at a pull request, I do not really care that a job named test is green. I care that the right tests ran, on the right commit, under a p...
Collapse
 
to21as profile image
Tobias •

The four-question split is the useful bit, because separating subject from producer is what makes a rerun answerable at all.

On the no-op item in your "does not prove" list: I hit that in a shape the gate could never catch, because the skip decision lived inside the test process. Every live suite opened with const run = process.env.API_KEY ? describe : describe.skip, so with no secret vitest exits 0 and the only trace is "2 skipped" in a log nobody opens. Green meant "the contract holds" and "nothing was checked" identically.

Moving the guard out into its own job fixed it, because then the skip arm shows up in the job list, where a name-based required check can see it. Does that sit inside your producer-identity question, or is it a fifth one?

Collapse
 
korovinaa97 profile image
Alexander Korovin •

I'd keep it under producer identity, with one important boundary. Once the applicability decision becomes its own named check, the gate can require it. If describe.skip stays inside Vitest, CI Evidence Gate cannot see it today.

The clean bridge is a separate required check over structured test output, for example expected > 0 and skipped == 0. And yes, a missing secret is evidence unavailable, not "not applicable."

This also exposed a gap in FFA-001: the write-up mentions green no-op jobs, but the executable fixture only covers stale SHA binding. I'll split the no-op case into its own runnable pattern. Thanks for the sharp example.

Collapse
 
pm25coder profile image
pm25coder •

Producer identity — and I think the reason it feels like it might be a fifth question is worth naming.

Your original shape had a producing job with an identifiable name (test) and a green conclusion; what had no identity anywhere was the decision that the suite applied. So the receipt was complete about the producer it names and silent about the actor that scoped it. Same layer as "which workflow and concrete job produced this result", just with two roles in one process — runner and applicability judge — where only one of them reaches the check-run namespace. The article's own "does not prove" list parks this as "an evidence-producing job did meaningful work instead of a no-op", which reads like an evidence-layer limit; I'd argue it is a provenance one, because a no-op is only invisible while the actor that chose it cannot be named.

What your fix actually does is make the decision nameable, and the name is the currency here: required checks select by name, so a decision that never gets one can never be required, forbidden, or read as missing. Worth being precise about the residual hole, though — selection sees a job, not what the job did. A job that shells out and swallows a non-zero exit is green under an excellent name. The split upgrades a silent decision into a checkable one; it does not audit the shell.

If you want the whole family covered with one rule instead of one job per guard, evaluate the run's structured report in a separate required check: assert expected > 0 and skipped == 0 over the required set from --reporter=json. That catches describe.skip, a t.skip() inside a case, an early return in beforeAll, and a config-level ignore decision — none of which a job split necessarily catches — and it never asks the test process to grade itself.

One polarity note, because it is the cheapest version of the fix: an absent secret is evidence unavailable, not evidence not applicable. Written the other way round — skip only when the secret exists but is malformed, fail when it is missing — the trap does not exist and no new job is needed. The job split is what the genuinely inapplicable case needs, where the right output is a recorded not_applicable with a reason rather than a green check.

Collapse
 
raknaos profile image
Raknaos •

Agreed — and the failure that eroded my trust hardest wasn't a fake green, it was a flaky test I marked retry-tolerant. It stayed green so long that every real failure started looking like noise.

The most useful change wasn't a smarter suite: it was pinning the check to the exact commit SHA and having the job declare which paths it actually exercised, so 'green on a matrix that ran nothing' still stands out. Anything that retries now gets a loud label in the job name. What triggered the post for you — a green job that ran nothing, or a name/workflow mismatch?

Collapse
 
korovinaa97 profile image
Alexander Korovin •

Closer to the second one, though it wasn't a single production incident. While building a synthetic fixture, I saw how easily a successful receipt could be accepted for the wrong SHA. “The check passed” is not the same as “this is fresh evidence for this exact change.”

Your retry-tolerant case is the same problem from another angle. I like the loud label — I'd also record the retry count so pass-after-retry doesn't get flattened into plain green.

Collapse
 
pm25coder profile image
pm25coder •

"Fresh evidence for this exact change" is the slippery half of it, because freshness splits in two and a gate usually only checks one.

I ran into the other half today, on a public repo that exists to test exactly this (dannwaneri/rules-demo-api, PR #2). The rule lives in AGENTS.md and CLAUDE.md; the PR moved it seven lines down in both files, text unchanged, then forced a re-run of the reviewer. By any reading of subject / producer / coverage, that re-run is genuinely fresh: its timestamp moved, and its citations were re-pinned to the new head 3ffacd90 — a commit that did not exist when the original review ran.

The numbers inside those citations did not move. The label still reads AGENTS.md[7-10] and src/index.ts[33], and at 3ffacd90 both ranges are wrong: lines 7–10 are rules 1, 2, 3 and a blank line, and the cited rule is on line 17. The URL was rebuilt per run (new SHA in the blob path) while the range it points at kept its old line numbers — so the rebuild step exists and simply skips the range. The platform's own diff anchor for that same finding was recomputed to line 34 while the body still says 33: two citations for one finding, one recomputed and one frozen.

That is why I'd file it under the policy clause rather than the subject clause. You wrote that the pull request must not be able to quietly weaken the policy — and here the PR didn't weaken the text, it moved it. The only artifact that could have surfaced the relocation is the one that stored its answer once, at review time. A merge gate that selects checks by name cannot see a doc relocation, so the citation is the sole witness, and it is rendered from a stored label rather than from the retrieval result.

Two cheap changes fall out of that: render the citation at read time from the thing you actually retrieved, and prefer an anchor that can't drift — a rule id or an anchor string — over a line range. It's the same choice your [[surfaces]] patterns already make: paths are stable, positions aren't.

The reason this stays invisible is worth naming: "kebab-case" is prose. Nothing downstream can check whether a prose citation is still true. Had the rule been something a gate could evaluate on its own — the literal ^/[a-z0-9-]+(/[a-z0-9-]+)*$ — the same re-run would have had a way to fail.

Thread Thread
 
korovinaa97 profile image
Alexander Korovin •

Good distinction: the run is fresh, but the evidence inside it is stale. I checked the PR—the review is pinned to 3ffacd9, but its cited ranges no longer match the files.

CI Evidence Gate verifies the outer receipt, so it wouldn’t catch this today. Stable rule IDs/anchors plus line ranges regenerated from the retrieved blob feels like the right fix. This deserves its own failure pattern.

Thread Thread
 
korovinaa97 profile image
Alexander Korovin •

Quick follow-up: I turned your example into an executable pattern in Fleet Failure Atlas: FFA-005: Fresh review, stale citation.

I also updated the CI Evidence Gate docs to make the boundary explicit: it verifies the outer receipt, not whether citations inside a producer's output are still accurate. Thanks again for such a concrete repro. It made the pattern a lot better.

Collapse
 
suraj09 profile image
Suraj Suradkar •

The distinction between “green” and “bound to the right change” is the important part. A passing check is only useful evidence if we can establish exactly what it evaluated—and what it didn’t.