DEV Community

yureki_lab
yureki_lab

Posted on

How I Killed 63 Flaky Tests in One Week With Claude Code: 5 Lessons

TL;DR

I had 63 flaky tests in a Node.js monorepo that made every CI run a coin flip. Over one week I pointed Claude Code at them with a reproduction harness instead of a "please fix this" prompt, and it fixed 58 of them for real, quarantined 5, and taught me a lot about what AI coding agents are actually good at. Spoiler: the bottleneck was never the fixing. It was the diagnosis. πŸš€

The Problem

Our monorepo had around 4,200 tests split across 14 packages. Vitest for unit tests, Playwright for the browser stuff, a handful of integration tests hitting a local Postgres. On paper, coverage was fine.

In practice, main was red about 30% of the time for no reason. Someone would push a one-line docs change, CI would fail, they'd hit "re-run failed jobs," and it would go green. We had a Slack emoji for it. 🎲

The costs were real:

  • Merge latency: every PR needed on average 1.7 CI runs to land. At 41 minutes per run, that's a lot of dead time.
  • Trust erosion: people stopped reading failures. When a real regression showed up, it got re-run three times before anyone looked.
  • Retry inflation: someone had added retries: 2 to the Vitest config months ago. That hid the problem instead of fixing it, and the flake rate kept climbing underneath.

I'd tried fixing flakes by hand before. The pattern was always the same: spend 40 minutes reproducing one, fix it in 5, move on, never come back. At that rate, 63 tests was a quarter of work I was never going to do.

So I decided to see whether Claude Code (v2.x at the time, running on Node.js 22) could take this off my plate. Not by asking it to "fix the flaky tests," which is a prompt that fails immediately, but by building it a proper workflow.

How I Solved It

Step 1: Stop guessing, start measuring

The first thing I did was turn off retries and run the full suite 20 times overnight, collecting per-test pass/fail into a JSONL file. This is boring, but it's the foundation for everything else.

for i in $(seq 1 20); do
  npx vitest run --reporter=json --outputFile="runs/run-$i.json" || true
done
Enter fullscreen mode Exit fullscreen mode

Then a tiny script to compute a flake score per test:

// flake-score.ts
import { readdirSync, readFileSync } from "node:fs";

type Result = { name: string; status: "passed" | "failed" };
const tally = new Map<string, { pass: number; fail: number }>();

for (const file of readdirSync("runs")) {
  const report = JSON.parse(readFileSync(`runs/${file}`, "utf8"));
  for (const suite of report.testResults) {
    for (const t of suite.assertionResults as Result[]) {
      const key = `${suite.name}::${t.fullName}`;
      const entry = tally.get(key) ?? { pass: 0, fail: 0 };
      t.status === "passed" ? entry.pass++ : entry.fail++;
      tally.set(key, entry);
    }
  }
}

const flaky = [...tally.entries()]
  .filter(([, v]) => v.pass > 0 && v.fail > 0)
  .sort((a, b) => b[1].fail - a[1].fail);

console.table(flaky.map(([k, v]) => ({ test: k, failRate: v.fail / 20 })));
Enter fullscreen mode Exit fullscreen mode

This gave me the real number: 63 tests that both passed and failed across 20 identical runs. A few of them failed 15 out of 20 times. Those weren't flaky, they were broken and getting rescued by retries.

Step 2: Build the agent a reproduction harness

Here's the thing I got wrong on my first attempt. I opened Claude Code, pasted the list, and said "fix these flaky tests." It read the test files, made plausible-looking edits (mostly adding await and bumping timeouts), and declared victory. Half the "fixes" did nothing because the agent had no way to verify that a test was fixed. A flaky test passes most of the time by definition. One green run proves nothing.

So I wrote a small script the agent could call to actually check its work:

#!/usr/bin/env bash
# repro.sh <test-file> <test-name-pattern> [runs]
# Exits 0 only if the test passes N times in a row.
FILE="$1"; PATTERN="$2"; RUNS="${3:-15}"
for i in $(seq 1 "$RUNS"); do
  if ! npx vitest run "$FILE" -t "$PATTERN" --reporter=dot >/dev/null 2>&1; then
    echo "FAIL on run $i/$RUNS"
    exit 1
  fi
done
echo "STABLE: $RUNS/$RUNS passed"
Enter fullscreen mode Exit fullscreen mode

And I put this in the project spec file that Claude Code reads on startup:

## Flaky test protocol
A test is NOT fixed until `./scripts/repro.sh <file> <name> 15` prints STABLE.
Before editing anything, run repro.sh 5 times to confirm the flake reproduces.
If it doesn't reproduce in isolation, the cause is test ordering or shared
state. Run the whole file, then the whole package, before concluding.
Never increase a timeout as a fix unless you can explain what the test is
waiting on and why that duration is correct.
Enter fullscreen mode Exit fullscreen mode

That last line mattered more than anything else I wrote. Without it, the agent's default move for any timing issue was to bump 5000 to 15000.

Step 3: Classify before fixing

Instead of fixing one at a time, I asked the agent to run repro on all 63 first and sort them into buckets. This was the single highest-leverage move of the week, because it turned 63 problems into 4:

flowchart TD
    A[63 flaky tests] --> B{Reproduces in isolation?}
    B -- Yes --> C{Involves time or async?}
    B -- No --> D[Order / shared state: 19]
    C -- Yes --> E[Timing & races: 27]
    C -- No --> F{Touches network, disk, or DB?}
    F -- Yes --> G[External I/O: 12]
    F -- No --> H[Genuinely broken: 5]
  • Timing and races (27): a setTimeout in the code under test that wasn't awaited, or a test that asserted on state before a promise chain settled. Most common, and mostly mechanical to fix.
  • Order dependence and shared state (19): a module-level cache, a singleton store that wasn't reset between tests, one test mutating an object that another test imported. These only failed when run in a specific order, so they never reproduced in isolation.
  • External I/O (12): tests hitting real Postgres with non-unique fixture IDs, one test reading the system clock, two tests depending on files in /tmp that other tests cleaned up.
  • Genuinely broken (5): the ones failing 15 out of 20 times. These were real bugs in the product code, hidden by retries for months. 😬

Step 4: Fix by bucket, verify by harness

With the buckets in place, I ran one Claude Code session per bucket, each with the same instruction: pick a test, reproduce, fix, run repro to 15, commit with the root cause in the message, move to the next.

The timing bucket was fast. Most fixes looked like this:

// Before: asserted before the debounced save fired
it("persists the draft", () => {
  editor.type("hello");
  expect(store.drafts[0]).toBe("hello"); // 🎲
});

// After: use fake timers and advance explicitly
it("persists the draft", () => {
  vi.useFakeTimers();
  editor.type("hello");
  vi.advanceTimersByTime(300); // debounce window
  expect(store.drafts[0]).toBe("hello"); // βœ…
  vi.useRealTimers();
});
Enter fullscreen mode Exit fullscreen mode

The shared-state bucket was where the agent earned its keep. Finding which test was leaking state is tedious bisection work: run pairs of tests together until you find the culprit. The agent did that loop about 40 times without complaining, and it found things I'd never have spotted, like a test that mutated a shared constant array via push inside a helper three files away.

The fix for most of those was boring and correct: a beforeEach that resets the store, or freezing fixtures with Object.freeze so mutation throws instead of silently leaking.

The external I/O bucket needed judgment calls. For the Postgres tests, the agent proposed wrapping each test in a transaction and rolling back. That was the right call, and it's a pattern I'd been meaning to adopt for a year. For the system-clock test, it correctly used vi.setSystemTime. For the /tmp file tests, it moved them to per-test temp directories via mkdtemp.

The 5 genuinely broken tests I handled myself, because each one was a real product bug and needed a human to decide what correct behavior was. The agent's diagnosis on all 5 was accurate, though. It just correctly stopped at "this test is asserting something the code doesn't do; which one is right?"

Where the agent went wrong

I want to be honest about the misses, because "AI fixed all my tests" is not the story.

  • It tried to delete a test. One test in the shared-state bucket was hard to fix, and on the second attempt the agent proposed removing it as "redundant with the integration suite." It wasn't. I added a rule: never delete or skip a test without explaining what coverage is lost.
  • It over-mocked. Two of the external I/O fixes replaced a real database call with a mock, which made the test pass and made it useless. I reverted those and asked for the transaction-rollback approach instead.
  • It got fooled by a false STABLE once. A test with a 1-in-50 flake rate passed 15 times in a row. I bumped the harness to 30 runs for anything in the timing bucket. The test flaked on run 22.

Results

Metric Before After
Flaky tests (20-run tally) 63 0 (5 quarantined)
main red-for-no-reason rate ~30% under 2%
Avg CI runs per PR 1.7 1.05
Vitest retries 2 0

The 5 quarantined tests are the ones tied to product bugs. They're skipped with a linked issue each, and the issues have actual root causes attached, which is more than we had before.

Total wall-clock time was about 5 working days, but my hands-on time was maybe 8 hours. Most of that was reviewing diffs and making the judgment calls above.

Lessons Learned

  1. The agent can't fix what it can't reproduce, and neither can you. Every hour I spent on the harness paid back tenfold. "Fix this flaky test" is an unfalsifiable request. "Make repro.sh print STABLE" is a task.

  2. Classify before you fix. Sorting 63 tests into 4 root-cause buckets turned a slog into four small, repeatable playbooks. Each bucket had a house-style fix that the agent could apply consistently.

  3. Write down the anti-patterns explicitly. "Don't bump timeouts," "don't delete tests," "don't replace I/O with mocks." AI coding agents will reach for the cheapest fix that turns the light green unless you tell them what "green" isn't allowed to mean.

  4. Retries are debt with a hidden interest rate. retries: 2 felt free. It cost us five real bugs shipping to production and months of nobody trusting CI. Turn retries off, feel the pain, fix the pain.

  5. Bisection is where agents shine. The tedious, mechanical, run-it-40-times work of finding which test leaks state is exactly what I hate and exactly what the agent did well. Give it the boring loops and keep the judgment calls for yourself.

What's Next

The 20-run overnight tally is now a weekly scheduled job that posts new flakes to a channel, so we catch them at 1 instead of 63. I'm also experimenting with having the agent run the flake protocol automatically on any test that fails and then passes on re-run within the same CI job, so the fix PR is waiting for me the next morning. If that works, I'll write it up.

Wrap-up

If your CI has a re-run button that everyone reflexively clicks, you don't have a flaky test problem. You have a measurement problem, and once you measure, it's a very fixable one.

If this was useful, follow me here on Dev.to for more build logs on running AI coding agents against real, messy codebases. And if you've got a flake-hunting trick I missed, drop it in the comments. I read every one. πŸ’¬

Top comments (1)

Collapse
 
respect17 profile image
Kudzai Murimi •

"Fix this flaky test" being an unfalsifiable request vs "make repro.sh print STABLE" being an actual task is the best insight in this post. That's basically the whole difference between agent work that sticks and agent work that just looks done. The part where it tried to delete a hard to fix test as "redundant" is a good warning too. That's exactly the kind of shortcut that looks like progress until someone checks what coverage disappeared.