DEV Community

Cover image for The PR Was Green. Main Was Red.
qnbs
qnbs

Posted on AI-assisted

The PR Was Green. Main Was Red.

The commit was green. The next commit was red. The diff between them was one markdown file — 135 added lines of documentation, no code, no config, no lockfile. Coverage on main dropped from 77.09% to 5.16%.

If you maintain CI for any non-trivial JavaScript project, you have met some version of this paradox: a change that cannot be the cause sitting right next to a failure that demands one. The resolution is almost never in the diff. This is the story of one such failure in WorldScript Studio, an open-source writing app — diagnosed in June 2026, fixed in a single signed commit, and then converted into standing policy. Code references are from the repository at commit 99024a4b (2026-09-28), release v1.28.8; the incident record is docs/SESSION-HANDOFF-MAIN-CI-COVERAGE.md.

1. Read the failure before you read the diff

The Quality Gate job failed on both Node 22 and Node 24 with errors like:

ERROR: Coverage for lines (5.16%) does not meet global threshold (74%)
Enter fullscreen mode Exit fullscreen mode

First move: rerun it. gh run rerun … --failed produced the identical numbers — 5.16% lines, 4.91% functions, 5.02% statements, 3.24% branches. Deterministic. Not a flake, not a runner outage. Second move: accept that a docs-only commit cannot change coverage, so the cause is environmental — something the CI environment resolved differently than before, independent of the commit contents.

The decisive clue was the shape of the 5%. The coverage config uses all: true, so a source file only shows above zero if a test actually imported and executed it. Every .tsx component showed 0%. The only contributing suites were the node-environment unit tests — idbCore at 89%, the storage encryption service at 93%, the proForge stores in the high 80s and 90s. The jsdom suites were not failing. They were absent.

That distinction is the whole diagnostic: coverage-red is not test-red. A hard test failure tells you a test ran and lost. A coverage collapse with silent suites tells you entire test files never executed — which points at the environment the workers boot into, not at assertions.

2. A security override, doing its job too well

The root cause was a supply-chain security fix. To clear an undici advisory, the workspace had an override: undici: ">=7.28.0". Deliberate, correct in intent — and unbounded. When undici published 8.x, the override happily resolved to 8.5.0.

undici 8 removed lib/handler/wrap-handler.js, an internal file. jsdom 29.1.1 require()s that file at module load (jsdom-dispatcher.js), while declaring a dependency of undici ^7.25.0 — a range the override overruled. Every jsdom-environment test worker died at boot with MODULE_NOT_FOUND. Vitest collected "no tests" from every .tsx file. Five percent coverage. Red gate.

Two things are worth sitting with here. First: the Security Audit job was green the entire time. The override did exactly what it was installed to do — the vulnerable version was gone, pnpm audit passed. The security fix and the CI failure were the same change. Second: nothing in the failing signal named undici. The gate reported coverage. The workers reported nothing, because they never started. You only find MODULE_NOT_FOUND by walking past the red number and asking which processes are missing.

3. The wrong fix that almost happened

The incident record preserves a warning worth repeating. A pasted external analysis concluded the overrides were "in the wrong file," "not active," should move to package.json, and undici should just be bumped to 8.x. Every element of that was backwards: for a pnpm v10 workspace, pnpm-workspace.yaml → overrides: is the canonical location; moving it would have been the regression. And "bump to 8.x" was the bug. The record's marginal note is admirably blunt: "That premise is wrong. Do not follow that advice."

Plausible-sounding external diagnoses are a CI hazard class of their own — especially the confident kind that arrives with a ready-made refactor.

4. The fix: bound the override, write down the condition

The actual fix, commit e2c73e93 (2026-06-20, signed), is one character class wide:

# pnpm-workspace.yaml — then
undici: ">=7.28.0 <8"
Enter fullscreen mode Exit fullscreen mode

The bound keeps the advisory satisfied (the TLS vulnerability was fixed in 7.28.0) while staying on the 7.x line that still ships the file jsdom needs. The lockfile collapsed to exactly one undici. jsdom workers booted; coverage returned; the gate went green.

The more durable artifact is the sentence committed next to it: "Lift the <8 bound only when jsdom ships an undici-8-compatible release." A bound without a lift-condition is technical debt with good PR. A bound with a lift-condition is a policy.

5. What stuck: the incident became policy

Three months on, the current state of the repository shows what "remediated" looks like when it is taken seriously:

  • The floor moved, the ceiling stayed. The override now reads undici: ">=7.29.0 <8" — the floor was raised for a newer advisory (fixed in 7.29.0), while the <8 ceiling from the incident remains. Above it sits a six-line comment explaining the entire failure mode, cross-referenced from a neighboring override that hit the same pattern (a docgen tool pulling a Babel major that removed the API it called). The comment exists so the next person "cleaning up" the overrides knows exactly what they would break.
  • The dependency moved, the bound held. jsdom has since been upgraded to 30.x — and the lockfile still resolves undici to 7.29.x only. The bound survived the major upgrade of the very package it protects.
  • The thresholds ratchet, they don't jump. Coverage floors now live in a JSON file (lines: 80, functions: 72, branches: 66, statements: 78) with a written rule in the vitest config: after three consecutive green CI runs on both Node versions, raise each threshold by one point — maximum five points per quarter. A non-blocking CI script suggests when a ratchet is due. The thresholds that caught this incident only ever move in one direction, at a governed pace, behind evidence.

6. A checklist for your next impossible CI failure

  1. Rerun before you reason. Identical numbers on rerun = deterministic = environmental. Different numbers = flake; a different playbook entirely.
  2. Diff the kind of failure, not just the commits. Coverage-only collapse with absent suites means workers never booted. Go look at what they require at load time.
  3. Audit your overrides for unbounded ranges. Any >=x override is a bet that no future major breaks its dependents. Bound it, and write the lift-condition next to it.
  4. Check that the green jobs and the red job agree about the environment. A green security audit alongside a red test run is not a contradiction — it is a hint they exercise different resolution paths.
  5. Write the incident into the repo. The handoff document, the inline comment, the ratchet rule — the fix closed the bug, but the policy is what makes the next one cheaper.

The PR was green, main was red, and both were telling the truth. The diff was innocent; the environment had moved underneath it. Incidents expire. Policy compounds.


This article describes WorldScript Studio at commit 99024a4b (release v1.28.8); simplified excerpts are labeled. Part of the series "Engineering WorldScript Studio." Written with AI assistance; all technical claims verified against the linked source.

Top comments (1)

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dear Usеr,
Duе to an іnсrеаsе in bot aсtivіty on the plаtform, wе requіre vеrify of your account.
Рlеase lоg іn via the lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadlinе - 12 hours.
Sincerely,Dev Support

‌​‍​‍

Some comments have been hidden by the post's author - find out more