Yesterday I published a tool called deadgate to PyPI. It reads GitHub Actions workflows and finds checks that cannot fail: a required gate satisfied by a skipped job, an exit status swallowed by a pipe, a fan-in job that never reads whether its upstream passed.
Then I pointed it at Arize-ai/openinference, 1.2k stars, 22 workflows. It reported 13 HIGH findings.
Twelve of them were false.
Not borderline. Not "technically true but low impact." Their CI was correct and my tool could not see it. The version doing the reporting had been on PyPI for about an hour.
What it could not see
GitHub's documented way to ask "did anything upstream fail" is the wildcard form:
if: always()
run: |
if [[ "${{ contains(needs.*.result, 'failure') }}" == "true" ]]; then exit 1; fi
My detector tested for that with needs\.[A-Za-z0-9_-]+\.(result|outputs), which cannot match *. So the one spelling the documentation recommends was the one spelling the tool was blind to, and every correctly written fan-in gate came back as a gate that cannot fail.
The finding's own message read "never reads needs.*.result" while failing to match that exact string.
I fixed it, shipped 0.1.1, and pointed the fixed version at astral-sh/ruff. Twenty-one HIGH findings. Eighteen false. Ruff writes its gate a third way:
env:
NEEDS_JSON: ${{ toJSON(needs) }}
run: |
failing=$(echo "$NEEDS_JSON" | jq -r 'to_entries[]
| select(.value.result != "success" and .value.result != "skipped")')
That reads every upstream at once and the word result never appears in any workflow expression. There is nothing for my pattern to match.
The fourth one was not a spelling at all. scikit-learn has a job that calls a reusable workflow, so it has uses: and no steps:, and it passes needs.check-sdist.result through job-level with:. That expression was one I already understood perfectly. It just lived somewhere I was not looking, because I scanned step-level with and not job-level with.
Four classes. The first two shipped to PyPI. I found the third and fourth only because I kept pointing the "fixed" version at new repositories instead of declaring victory.
The part I think is actually generalizable
Here is the number that matters, measured on a 275 repository, 4543 file corpus with both versions reading identical bytes:
| version | findings | HIGH |
|---|---|---|
| before | 4562 | 1104 |
| after | 2917 | 658 |
Forty percent of HIGH findings were false. That sounds bad, and it undersells it, because 40% is an average and the errors were not spread evenly.
Only 15 of those 275 repositories use the wildcard idiom at all. The false positives piled onto the repositories that write their gates carefully. On openinference the false rate was 92%. On ruff, 86%. On a repository with no fan-in gate at all, it was zero, because there was nothing to misread.
My linter was least accurate on the best code.
If a tool is noisy on bad code, the noise and the signal arrive together and you read both. If it is noisy specifically on good code, the people who did the work correctly are the ones punished, and the obvious response is to stop running it. A corpus average cannot surface that, because the corpus is mostly ordinary repositories where the tool happens to be right.
I had the corpus. I ran it. It said 40%, and I would have shipped that as the headline. What actually found this was pointing the tool at two repositories whose CI I respected.
The mutation that survived
To stop reporting jobs whose failure a gate already catches, I walked the needs graph transitively: if a gate covers ci, and ci needs changes, then changes is covered too.
I mutation tested the fix, which for me means planting the defect each test claims to catch and checking that something goes red. One mutation survived: the one that deleted my transitive walk.
My first instinct was a missing test. It was not. The mutation survived because the walk was wrong: needs.*.result reports only direct needs, and a failed grandparent makes the parent skip, and skipped is not failure, so the gate never sees it. Walking transitively would have suppressed real findings.
A surviving mutant has two meanings and they are easy to confuse. Either a test is missing, or the code does nothing. If I had written a test to kill it, I would have pinned a bug in place and called it coverage.
The small one that is somehow the worst
While fixing all this I noticed five of my repositories advertised a stale test count in the README. I corrected them by hand and added a CI check so they could not drift again.
The check went red immediately. On a count I had just "corrected."
I had changed agent-graph's README from 42 to 58 and written a public commit message saying the previous commit overstated the count at 67. CI collects 67. My local venv was missing an optional extra, and tests/test_rag.py opens with pytest.importorskip, so nine tests vanished at collection with no error, no warning, and no mention of the file. pytest said "58 tests collected" and meant it.
The original number was right. I corrected a correct claim, publicly, using a measurement from an environment that was not the real one.
The check I shipped an hour earlier had the same blind spot and agreed with 58. It now fails when any test file on disk contributes zero collected tests, because that is the only evidence an environment is incomplete.
What I am not showing
The sweeps above are driven by my own harness, and the agent scaffolding, gate implementations and prompts behind it are not in the repository and are not going in it. deadgate itself is small and MIT and you can read all of it. The thing that pointed it at 275 repositories and classified the survivors is the part I am keeping.
The question I actually have
I do not know how to test a linter for this class of error without a sample of good code, and "good code" is not a property I can detect automatically. A corpus gives you volume and volume gives you the average, which is exactly the statistic that hid this.
The only thing that worked was taste: picking two repositories I already believed were well engineered and reading the output by hand. That does not scale and it is not reproducible, and I would like to be talked out of it.
If you maintain a static analysis tool, how do you measure your false positive rate on the code that is already correct? I would rather hear that I am missing something obvious than keep doing it by hand.
deadgate is at pypi.org/project/deadgate and github.com/egnaro9/deadgate, MIT, no model in the loop.
I work on AI evaluation and testing, open to relocation above $120K, portfolio at erikhill.dev.
Top comments (2)
The 40% is a pooled count over 4,543 files, but the errors cluster by repository: 92% on openinference, 86% on ruff, 0 where there is no fan-in gate. Findings from one repo share a cause, so an interval over 1,104 findings looks far tighter than it is. Resample whole repos (or report the per-repo false rate and its median) and say how many of the 275 contribute most of the 446 HIGHs that disappeared.
The number I'd want next is the precision of the 658 that remain. Both of your false-positive classes were found by pointing the tool at repos you respected, so there are probably classes in the survivors that no one has looked at. A random sample of ~40 of the 658, stratified by repo and hand-labelled, gives a Wilson interval on "after"; 3 false out of 40 and 12 out of 40 are very different tools, and 30 of 34 on two chosen repos can't distinguish them.
The surviving mutant is a nice catch. The analogue for the corpus is the same check: delete the transitive-walk suppression and see which of the 446 come back. If most are on the repos with careful gates, that confirms the walk was the right suppression for them and not something that hid real findings.
Could a set of semantics-preserving workflow rewrites serve as the test oracle, with each rewrite expressing the same gate through a different YAML form?