Some time ago I wrote about the day our coverage number collapsed from 77% to 5% because a dependency update quietly killed the test workers. The number had been lying for one commit — but the deeper lesson survived the fix: even an accurate coverage percentage only answers the question "did this line execute?" It never answers the question anyone actually cares about: "would a test fail if this code were wrong?"
A test with no assertion is green. A test whose assertion was weakened in a refactor is green. A test that asserts on a copy of the output instead of the output is green. Coverage counts all of them at full value, because coverage measures execution, not conviction.
Mutation testing is the standard answer to this. A tool — in our case Stryker — makes small, deliberate changes to your source: flips a comparison, removes a boundary, deletes a statement. Each change is a mutant. The suite runs against the mutant. If a test fails, the mutant is killed — your tests noticed the code being wrong. If the suite stays green, the mutant survived — and you have a concrete, located proof that your tests tolerate a specific bug. The mutation score is the percentage of mutants that die.
That is a much more honest metric. It is also expensive, slow, and surprisingly easy to run in a way that lies to you. This article is about how WorldScript Studio, the open-source writing app this series dissects, engineered its mutation pipeline to tell the truth — and about the uncomfortable fact that, as I write this, its own mutation status is red. Both halves matter. Code references are from the repository at commit 8b329633 (2026-09-28), release v1.28.8.
1. Why you can't just run it
Mutation testing multiplies your test bill. Every targeted line spawns several mutants, and every mutant triggers at least the tests that cover it. Run Stryker against a whole application and you are not running a test suite — you are running thousands of test suites. On a normal developer machine, that is not a slow command; it's an abandoned one.
So the first honest decision in this repository is about where mutation testing lives, and the answer is: not on your laptop. The pnpm mutation script doesn't run Stryker — it runs a bouncer:
// scripts/assert-ci-only.mjs (shortened)
if (process.env.CI !== 'true') {
console.error(
`mutation testing is CI-only on this constrained workstation. Use: gh workflow run mutation.yml`,
);
process.exit(1);
}
The README says the quiet part out loud: mutation testing, like deep E2E and Lighthouse, is "primarily CI-owned in this repository and can be inappropriate on low-resource developer hardware." And the workflow as it stands today has exactly one trigger: workflow_dispatch. No push, no pull request, no schedule. A human decides when to spend the money, with a mode (incremental or force), a scope selector, and a concurrency dial as explicit inputs.
The second honest decision is about what gets mutated. Not the repo — a curated scope. stryker-scope.json defines eight modules, twenty-five files total, each tagged with a risk tier (four A, four B). The A tier is where a tolerated bug would hurt most: the AI policy core, command palette logic, the copilot's heuristic engine, project selectors. A selector input lets a dispatch run everything, only tier A, or a single module.
And the scope file is not trusted. scripts/stryker-scope.mjs validates it fail-loud on every invocation: duplicate module names throw, missing risk tiers throw, duplicate targets throw, and — the check I appreciate most — any mutate target or test file that doesn't exist on disk throws. A curated list decays silently as files are renamed; this one breaks the glass the moment it drifts. Unrecognized selectors throw with the list of valid names attached.
2. Engineering the truth pipeline
Here is the design problem nobody warns you about: once mutation runs are expensive, sharded, and cached, the reporting becomes the risky part. Eight parallel jobs produce eight shard reports, and a summary that happily averages whatever shards happened to arrive is a summary that can show you green while half the evidence is missing. The pipeline's answer is to treat its own outputs as hostile.
One source of truth for scope and matrix. The Stryker config imports its mutate list from stryker-scope.mjs, and the CI workflow generates its matrix from the same file — the scope job literally runs the script with --matrix and feeds the JSON into fromJSON. There is no second copy of the module list living in YAML waiting to drift. This isn't a convention; it's pinned by a test, of which more below.
Per-module incremental caches. Stryker's incremental mode skips re-mutating unchanged code, and each matrix job restores a cache keyed by module name with a prefix fallback. force mode deletes the cache file first. Incremental runs make a manually dispatched, eight-module run affordable; the cache being per-module means one module's corruption doesn't poison the rest.
Hardened inputs. Every dispatch input is validated in the shell before use: module names must match ^[a-z0-9-]+$, the mode must be exactly incremental or force, concurrency must be digits or a percentage — anything else exits nonzero. The actions are SHA-pinned, checkouts use persist-credentials: false, and every job runs with permissions: contents: read. This is the same supply-chain posture I described in the article on releases being more than tags, applied to a workflow that never touches a release at all.
A threshold that actually bites. The Stryker config sets high: 85, low: 70, and — look carefully — break: 75. The break value sits above the low watermark, which means the color bands are nearly decorative: any score under 75 fails the module job outright. There is no comfortable yellow zone where a 72% score looks gently disappointing while the build stays green. Under the bar, the job is red. That's the whole point of a gate.
Fail-closed aggregation. This is my favorite part. The aggregate job runs if: always() — it executes even when matrix jobs failed, because someone must render the summary — and then it does two things. First, scripts/aggregate-stryker-reports.mjs validates every shard like a prosecutor: a missing mutation.json for any expected module throws; an unparseable report throws; the six metric identities (totalDetected = killed + timeout, and so on) must hold exactly; a survived mutant that reports zero completed tests is invalid evidence, not a survivor, because a zero-test "survival" is an artifact of an empty per-test filter, and the script throws with that reasoning in a code comment. If a report lacks its metrics block entirely, the script rebuilds the metrics from individual mutant statuses — and an unrecognized status throws rather than being silently bucketed. Second, after all that, if any matrix job failed, the aggregate job prints that the summary "remains informational" and exits 1. The header comment states the design rule in one line: "Reject partial or inconsistent shard reports before aggregation can false-green."
The workflow has its own tests. Eleven of them, living in tests/unit/tooling/: eight for the aggregation logic, and three policy tests that read the workflow file itself and assert the invariants — one target source for config and matrix, supported incremental plumbing that preserves shard identity, and fail-closed validation of the optional test-file mappings. A pipeline whose entire purpose is producing trustworthy numbers gets the same treatment as production code: tests that would fail if someone "simplified" the YAML.
3. The honest status: it's red
Now the part a launch post would leave out.
The workflow has run 57 times since its first commit in May 2026. Fifty-three were manual dispatches; the other four — three scheduled runs from its first weeks and one push-triggered run on August 22, the day of the workflow's latest commit — predate the current manual-only trigger list. The public record, visible to anyone via the GitHub Actions API, shows 8 successes, 43 failures, 6 cancellations. The recent picture is worse than the lifetime one: the latest full runs are red, and three of the eight modules (the copilot, the project feature slice, the AI provider adapter) haven't produced a green shard in the last nineteen-or-so runs. The scope-validation job, by contrast, is 18 for 18. The machinery works. The scores, currently, do not clear the bar — or the jobs die trying, and here's the important part: I can't tell you which. Job logs aren't publicly accessible, so "under the 75 break threshold," "hit the 45-minute timeout," and "something in the shard is broken" are all consistent with the public record, and this article will not pretend otherwise.
What I can show is that green is achievable and what green means: on August 22, two single-module dispatches — services-ai-core and services-commands — went fully green, scope validation through aggregation. Under this pipeline's rules, that means every expected shard produced a valid report and the scores met the gate. And I can show that the aggregate job's record (3 green, 16 red over that same window) is exactly what fail-closed looks like in practice: it never once summarized partial evidence into a passing build.
The workflow's own header anticipated all of this, in a sentence I keep coming back to: "runtime and score remain measured outputs rather than undocumented promises." No badge on the README claiming a number. No wiki page asserting a score. When you want to know the mutation status of this project, you look at the same red runs I'm looking at. A red report you can verify is worth more than a green one you're asked to believe.
4. The doc that drifted
And then, because this series exists to catch exactly these things: while verifying the claims above I opened docs/BEST-PRACTICES.md, which documents the testing strategy, and found this line — as of commit 8b329633, still there:
Mutation: Stryker job (
mutation.yml) is informational untilbreakthreshold is raised. Current targets: 9 service files instryker.conf.json.
Three stale facts in one sentence. The job is not informational — break is set to 75 and module jobs fall under it. The targets are not 9 service files — they're 25 files in eight curated modules. And stryker.conf.json does not exist; the file is stryker-scope.json. The line was presumably accurate during the workflow's first iteration, back in May — and then the pipeline evolved, seventeen workflow commits later and most recently in August, while the sentence stayed behind.
I'm not pointing this out to shame anyone — I'm pointing it out because it's the same failure mode this series keeps finding in different costumes. In the release article it was claims that outlived their evidence; here it's a doc line that outlived the system it describes. The fix is one sentence of documentation, and it belongs in the same kind of PR as any code change: scoped, reviewed, and ideally caught by exactly the kind of drift-checking habit the scope validator already applies to file paths. The system that measures truth deserves a description that tells it.
5. What to copy
If you take one design home from this pipeline, make it the posture, not the tool:
- Curate the mutation scope. Twenty-five risk-tiered files teach you more than a whole-repo run you abandon after six hours. Validate the list fail-loud against the filesystem, because curated lists rot.
- One source of truth for scope and matrix. If your CI matrix and your tool config can disagree, eventually they will. Generate one from the other — and pin the invariant with a test.
- Aggregate fail-closed. Missing shards, inconsistent metrics, and zero-test survivors are errors, not rounding choices. Partial evidence must never become a green summary.
-
Let
breakbe the gate. Color thresholds are for dashboards; the exit code is for the build. If you set a bar, let the bar fail jobs. - Publish your status, not your promise. A measured red beats an asserted green — but only if the measurement pipeline itself refuses to tell stories. Test the pipeline too.
-
Keep the docs about the truth-system true. The one-line fix in
BEST-PRACTICES.mdis, as far as I can tell, still unwritten. Consider this its issue report.
Coverage tells you the code ran. The mutation score tells you the tests object when the code is wrong — but only when the pipeline producing it would rather show you red than show you fiction. Right now this one is red, in public, with the receipts to prove it means something when it turns green.
This article describes WorldScript Studio at commit 8b329633 (release v1.28.8); workflow run data was retrieved from the public GitHub Actions API on 2026-09-28 and reflects that date. Simplified excerpts are labeled. Part of the series "Engineering WorldScript Studio." Written with AI assistance; all technical claims verified against the linked source.
Top comments (4)
the zero-test survivor rule is the one to steal. every metrics pipeline has the same hole: absence of evidence processed as evidence of absence. a mutant nobody tested isn't a survivor, a shard that never arrived isn't a pass, and an average of whatever showed up is fiction with decimals. and publishing the red runs in public is the move most repos won't make, which is exactly what makes the green worth believing when it comes.
"fiction with decimals" is exactly the right phrase for it - that's the failure the metric identity checks are really guarding against: totalDetected must equal killed + timeout, or the report doesn't get summarized at all. And yes, the zero-test survivor rule felt almost too strict when I read it - until you realize an empty filter result would otherwise silently inflate the survivor count. Glad the design travels beyond this one pipeline.
"almost too strict" is how you know the guard is right. you don't design the check for the case you imagined, you design it for the empty filter you didn't. the pedantic rule in review is the one that catches the bug nobody could picture.
I checked the two claims here that a reader can check, and they hold exactly.
The census: the Actions API for mutation.yml gives 57 runs -- 8 success, 43 failure, 6 cancelled -- and by event 53 workflow_dispatch, 3 schedule, 1 push. Both totals reproduce run for run. The doc drift is also still there at 8b329633: that line is present verbatim in docs/BEST-PRACTICES.md, stryker-scope.json carries eight modules and twenty-five mutate targets (3+4+6+1+2+4+4+1) with four riskTier A and four B, and stryker.conf.json 404s. Three stale facts, as written.
One thing moves with a date, though. The newest run in that census is 2026-09-20T23:26:15Z, and the repository has moved since -- the newest merge on the default branch is 2026-09-28T13:43. With workflow_dispatch as the only trigger, "the status is red" is a reading of the 20th rather than of the code this article describes; between them there are eight days of commits that no gate has looked at. That is your own thesis pointed at your own status line. A red report you can verify is worth more than a green one you are asked to believe, and both are worth more again when the sentence carries the date the reading was taken. The article dates its other numbers; the status sentence is the one that does not, and it is the one a reader will take as today.
On the identity checks, one limit is worth stating beside them. totalDetected = killed + timeout is a report checked against itself, so a denominator that shrank before the arithmetic satisfies every identity: a file dropped from the scope list, or an exclusion made on the tool side, leaves the identity true over a smaller total. The scope step is where you already validate fail-loud against the filesystem, so most of that audit exists -- what a filesystem check cannot see is what the tool removes after it.
That same identity says the status decomposes, and timeout is the column I would surface. In a sharded run on constrained runners, a timeout is a statement about the machine as much as about the suite, the way a zero-test survivor is a statement about the filter rather than about the tests. Printing the timeout share beside the score costs one number and removes the most plausible wrong reading of a red module.
Last: break at 75 above low at 70 is a good gate, and it retires the colour band. A module at 76% renders as a warning while the pipeline reads it as a pass, so nobody comparing the two numbers can tell which one decides. Your point 4 says colour thresholds are for dashboards and the exit code is for the build; the small price is that the dashboard colour no longer means what those thresholds mean to a reader who has seen other Stryker reports.