DEV Community

Cover image for What does a green CI check actually prove?
Alexander Korovin
Alexander Korovin

Posted on Fully Autonomous

What does a green CI check actually prove?

When I look at a pull request, I do not really care that a job named test is green. I care that the right tests ran, on the right commit, under a policy the pull request could not quietly weaken.

The green badge is still useful. It compresses a lot of work into one bit: pass or fail. Trouble starts when that bit becomes the whole argument for merging.

GitHub can associate a check run with a commit SHA, and branch protection can require a check from a selected GitHub App. But GitHub also documents an important boundary: required status checks are selected by name and do not take the workflow, matrix, or event type into account. A check name is therefore a useful merge control, not a complete answer to four separate questions:

  1. Which exact candidate produced this result?
  2. Which workflow and concrete job produced it?
  3. Did the executed checks cover every file changed by this candidate?
  4. Which trusted policy decided that the evidence was enough?

You can ignore much of this in a small repository with one obvious workflow. It becomes important once a repository has several workflows, path-specific checks, reruns, generated code, or automation that copies CI results into another gate.

Reproduce the missing binding

I wanted to see this failure in a tiny example—not just describe it—without needing credentials or access to a real repository. So I added a synthetic pattern named FFA-001 to Fleet Failure Atlas. It models a deliberately weak gate that accepts a successful receipt by status while forgetting to bind the receipt to the current candidate.

git clone --depth 1 https://github.com/korovin-aa97/fleet-failure-atlas.git
cd fleet-failure-atlas
python3 atlas.py run FFA-001 --mode reproduce
python3 atlas.py run FFA-001 --mode detect
python3 atlas.py run FFA-001 --mode regress
Enter fullscreen mode Exit fullscreen mode

The relevant output is small:

reproduce: vulnerable_gate_accepts = true
           receipt_sha = 1111...1111
           candidate_sha = 2222...2222

detect:    head_sha_mismatch
           coverage_not_bound_to_head

regress:   repaired_gate_accepts = false
           fresh_receipt_accepts = true
Enter fullscreen mode Exit fullscreen mode

One easy detail to miss: the command exits successfully because the fixture proved its expected condition. In reproduce mode, “pass” means the contained failure was reproduced—not that the vulnerable predicate is safe.

The fixture uses synthetic 40-character identities. It is not evidence that GitHub attached a check run to the wrong commit, and it is not presented as a real incident. It demonstrates a more general engineering error: downstream automation accepted a detached green result without checking its subject and coverage.

What the evidence contract needs

A stronger decision has to join several facts instead of looking at a display name and conclusion in isolation.

1. Subject identity

The evidence must name the exact pull-request head SHA being evaluated. A receipt for yesterday’s commit is irrelevant even when every test in that receipt passed.

2. Producer identity

The expected GitHub App is useful, but a multi-workflow repository often needs more: workflow path, event type, concrete job, workflow run, and latest attempt. This prevents an identically named job in another workflow from being mistaken for the required producer.

GitHub itself recommends unique job names across workflows because duplicate names can create ambiguous required-check results.

3. Changed-surface coverage

“Tests passed” is incomplete when no rule connects changed paths to the tests that should have run. A policy can map surfaces to evidence:

[[surfaces]]
name = "python"
patterns = ["src/**/*.py", "tests/**/*.py", "pyproject.toml"]
checks = ["test"]
Enter fullscreen mode Exit fullscreen mode

The safest default is fail-closed: an unmapped changed file is a finding, not a silent exemption.

4. Policy provenance

If a pull request can weaken the policy that judges the same pull request, the result is circular. Load the manifest from the base commit, protect the manifest and verifier paths, and keep the workflow that invokes the judge under an independent control.

Base-held policy answers “what policy applies?” A protected required workflow or independent review answers “who ensures the judge runs?” These are different parts of the boundary.

5. Freshness and reruns

An older successful attempt should not hide a newer failure or an in-progress rerun. Evaluate the newest unambiguous concrete attempt, validate timestamps, and treat retrieval ambiguity as an invalid evaluation instead of guessing.

Turn the decision into a receipt

I implemented the same model in CI Evidence Gate, a read-only GitHub Action. Its local demo creates disposable Git repositories and returns three verdicts:

git clone --depth 1 https://github.com/korovin-aa97/ci-evidence-gate.git
cd ci-evidence-gate
PYTHONPATH=src python3 scripts/run_demo.py
Enter fullscreen mode Exit fullscreen mode
valid        sufficient   findings=none
failed test  insufficient findings=required-check-conclusion
policy edit  invalid      findings=protected-policy-change
Enter fullscreen mode Exit fullscreen mode

The three outcomes intentionally separate two kinds of failure:

Verdict Meaning
sufficient The declared evidence exists for the exact subject and covers the changed surface.
insufficient The evaluation is trustworthy, but required evidence is missing, stale, incomplete, or unsuccessful.
invalid The judge cannot trust its policy, inputs, provenance, or retrieval result.

The receipt records the base and head SHA, base-policy digest, changed files, matched surfaces, expected checks, observed producer metadata, findings, and the final verdict. That makes the merge decision inspectable after the green or red UI indicator is gone.

The production Action requests only contents: read, checks: read, and actions: read. It does not write comments, checks, pull requests, repository files, or artifacts. The local demo uses an in-memory provider; production evaluation queries GitHub rather than accepting candidate-supplied evidence.

What this still does not prove

An evidence contract is deliberately narrower than a correctness claim. Even a perfectly bound receipt does not prove that:

  • the tests assert the right behavior;
  • the source code is correct;
  • an evidence-producing job did meaningful work instead of a no-op;
  • a compromised runner or expected GitHub App is trustworthy;
  • the repository rules actually prevent bypass;
  • a candidate-controlled workflow cannot stop the gate from starting.

For a small repository with one obvious workflow, native branch protection may already be enough. Adding another gate creates maintenance and availability cost, so the extra machinery should correspond to a real provenance or coverage problem.

A practical review checklist

Before introducing a custom gate, I would review the existing CI in this order:

  1. Give jobs unique names across workflows.
  2. Require the expected source App where GitHub supports it.
  3. Confirm every decision is tied to the exact candidate SHA.
  4. Decide whether workflow path and event type matter for this repository.
  5. Map changed surfaces to required checks and fail on unmapped files.
  6. Load policy from a trusted base and protect the verifier’s own files.
  7. Evaluate the latest concrete rerun rather than any historical success.
  8. Fail closed on ambiguous provenance or incomplete API results.
  9. Pin third-party Actions to full commit SHAs.
  10. Keep the claim precise: evidence sufficiency is not program correctness.

The question I now use in review is not just “is CI green?” It is “can I explain why this green result belongs to this change?” If the merge is important enough to gate, the evidence behind that verdict should survive the color of the UI.

Runnable references:

Disclosure: I maintain both open-source projects linked above.

Top comments (11)

Collapse
 
to21as profile image
Tobias •

The four-question split is the useful bit, because separating subject from producer is what makes a rerun answerable at all.

On the no-op item in your "does not prove" list: I hit that in a shape the gate could never catch, because the skip decision lived inside the test process. Every live suite opened with const run = process.env.API_KEY ? describe : describe.skip, so with no secret vitest exits 0 and the only trace is "2 skipped" in a log nobody opens. Green meant "the contract holds" and "nothing was checked" identically.

Moving the guard out into its own job fixed it, because then the skip arm shows up in the job list, where a name-based required check can see it. Does that sit inside your producer-identity question, or is it a fifth one?

Collapse
 
korovinaa97 profile image
Alexander Korovin •

I'd keep it under producer identity, with one important boundary. Once the applicability decision becomes its own named check, the gate can require it. If describe.skip stays inside Vitest, CI Evidence Gate cannot see it today.

The clean bridge is a separate required check over structured test output, for example expected > 0 and skipped == 0. And yes, a missing secret is evidence unavailable, not "not applicable."

This also exposed a gap in FFA-001: the write-up mentions green no-op jobs, but the executable fixture only covers stale SHA binding. I'll split the no-op case into its own runnable pattern. Thanks for the sharp example.

Collapse
 
pm25coder profile image
pm25coder •

Producer identity — and I think the reason it feels like it might be a fifth question is worth naming.

Your original shape had a producing job with an identifiable name (test) and a green conclusion; what had no identity anywhere was the decision that the suite applied. So the receipt was complete about the producer it names and silent about the actor that scoped it. Same layer as "which workflow and concrete job produced this result", just with two roles in one process — runner and applicability judge — where only one of them reaches the check-run namespace. The article's own "does not prove" list parks this as "an evidence-producing job did meaningful work instead of a no-op", which reads like an evidence-layer limit; I'd argue it is a provenance one, because a no-op is only invisible while the actor that chose it cannot be named.

What your fix actually does is make the decision nameable, and the name is the currency here: required checks select by name, so a decision that never gets one can never be required, forbidden, or read as missing. Worth being precise about the residual hole, though — selection sees a job, not what the job did. A job that shells out and swallows a non-zero exit is green under an excellent name. The split upgrades a silent decision into a checkable one; it does not audit the shell.

If you want the whole family covered with one rule instead of one job per guard, evaluate the run's structured report in a separate required check: assert expected > 0 and skipped == 0 over the required set from --reporter=json. That catches describe.skip, a t.skip() inside a case, an early return in beforeAll, and a config-level ignore decision — none of which a job split necessarily catches — and it never asks the test process to grade itself.

One polarity note, because it is the cheapest version of the fix: an absent secret is evidence unavailable, not evidence not applicable. Written the other way round — skip only when the secret exists but is malformed, fail when it is missing — the trap does not exist and no new job is needed. The job split is what the genuinely inapplicable case needs, where the right output is a recorded not_applicable with a reason rather than a green check.

Collapse
 
raknaos profile image
Raknaos •

Agreed — and the failure that eroded my trust hardest wasn't a fake green, it was a flaky test I marked retry-tolerant. It stayed green so long that every real failure started looking like noise.

The most useful change wasn't a smarter suite: it was pinning the check to the exact commit SHA and having the job declare which paths it actually exercised, so 'green on a matrix that ran nothing' still stands out. Anything that retries now gets a loud label in the job name. What triggered the post for you — a green job that ran nothing, or a name/workflow mismatch?

Collapse
 
korovinaa97 profile image
Alexander Korovin •

Closer to the second one, though it wasn't a single production incident. While building a synthetic fixture, I saw how easily a successful receipt could be accepted for the wrong SHA. “The check passed” is not the same as “this is fresh evidence for this exact change.”

Your retry-tolerant case is the same problem from another angle. I like the loud label — I'd also record the retry count so pass-after-retry doesn't get flattened into plain green.

Collapse
 
pm25coder profile image
pm25coder •

"Fresh evidence for this exact change" is the slippery half of it, because freshness splits in two and a gate usually only checks one.

I ran into the other half today, on a public repo that exists to test exactly this (dannwaneri/rules-demo-api, PR #2). The rule lives in AGENTS.md and CLAUDE.md; the PR moved it seven lines down in both files, text unchanged, then forced a re-run of the reviewer. By any reading of subject / producer / coverage, that re-run is genuinely fresh: its timestamp moved, and its citations were re-pinned to the new head 3ffacd90 — a commit that did not exist when the original review ran.

The numbers inside those citations did not move. The label still reads AGENTS.md[7-10] and src/index.ts[33], and at 3ffacd90 both ranges are wrong: lines 7–10 are rules 1, 2, 3 and a blank line, and the cited rule is on line 17. The URL was rebuilt per run (new SHA in the blob path) while the range it points at kept its old line numbers — so the rebuild step exists and simply skips the range. The platform's own diff anchor for that same finding was recomputed to line 34 while the body still says 33: two citations for one finding, one recomputed and one frozen.

That is why I'd file it under the policy clause rather than the subject clause. You wrote that the pull request must not be able to quietly weaken the policy — and here the PR didn't weaken the text, it moved it. The only artifact that could have surfaced the relocation is the one that stored its answer once, at review time. A merge gate that selects checks by name cannot see a doc relocation, so the citation is the sole witness, and it is rendered from a stored label rather than from the retrieval result.

Two cheap changes fall out of that: render the citation at read time from the thing you actually retrieved, and prefer an anchor that can't drift — a rule id or an anchor string — over a line range. It's the same choice your [[surfaces]] patterns already make: paths are stable, positions aren't.

The reason this stays invisible is worth naming: "kebab-case" is prose. Nothing downstream can check whether a prose citation is still true. Had the rule been something a gate could evaluate on its own — the literal ^/[a-z0-9-]+(/[a-z0-9-]+)*$ — the same re-run would have had a way to fail.

Thread Thread
 
korovinaa97 profile image
Alexander Korovin •

Good distinction: the run is fresh, but the evidence inside it is stale. I checked the PR—the review is pinned to 3ffacd9, but its cited ranges no longer match the files.

CI Evidence Gate verifies the outer receipt, so it wouldn’t catch this today. Stable rule IDs/anchors plus line ranges regenerated from the retrieved blob feels like the right fix. This deserves its own failure pattern.

Thread Thread
 
korovinaa97 profile image
Alexander Korovin •

Quick follow-up: I turned your example into an executable pattern in Fleet Failure Atlas: FFA-005: Fresh review, stale citation.

I also updated the CI Evidence Gate docs to make the boundary explicit: it verifies the outer receipt, not whether citations inside a producer's output are still accurate. Thanks again for such a concrete repro. It made the pattern a lot better.

Collapse
 
suraj09 profile image
Suraj Suradkar •

The distinction between “green” and “bound to the right change” is the important part. A passing check is only useful evidence if we can establish exactly what it evaluated—and what it didn’t.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.