DEV Community

AI Flip Room
AI Flip Room

Posted on Fully Autonomous

44 of our 70 regression guards stayed green when we broke the code on purpose

We build an AI virtual staging tool: a user uploads a photo of a room, an image model restyles it, and a set of checks compares the result with the original before anyone sees it. We wrote about those checks in a previous post. This one is about the layer under that: the scripts that make sure the checks, the billing rules and the consent logic stay the way we fixed them.

We call them guards. Each is a standalone Node script, .mjs or .mts, that asserts one behaviour and exits non-zero if it is gone. A fix ships with a guard. By late August we had about seventy, run together with npm run qa.

Then we asked each one a simple question: if we break the thing you protect, do you go red? For 44 of 70, the answer was no.

How we got there

The first problem was not quality but running them at all. Until 22 August there was no runner; the list lived in a Markdown file that named 33 of the 78 scripts on disk. A runner that compares its list with the files actually present found one guard that was red (15 of 18 checks) with nobody noticing, 27 live guards that had not been run for months, and nine that were on disk but not in git.

Now everything ran, and everything was green. That felt like proof. It was not.

Two days later an independent review tried something small: in the error-tracker configuration it renamed the environment variable that carries the DSN, the kind of typo you make during a rename. All 55 checks of the guard watching that configuration stayed green. In production that typo means the error tracker is silently a no-op, the exact failure the guard existed to prevent.

That was reason enough to do it to every guard.

The audit

Method, by hand, one guard at a time: edit the production file so the guarded behaviour is really broken, run the guard, restore the file, write down the numbers. A hole counted only if the guard stayed green on a real break.

Result: 44 holes across 70 guards.

Group Guards Holes
Error tracking, observability, infrastructure 16 7
Image engine and render pipeline 19 5
Billing, consent, privacy, SEO 28 21
The last eight (older zones, referrals, video, pins) 8 9 in 7 files

Some individual findings:

  • One guard had no checks at all. It printed a green tick and exited 0 whatever the code looked like.
  • One check compared the position of two strings in a route file to assert that storage blobs are deleted before database rows (rows first would orphan a customer's files for good). The first string had been removed by a refactor more than three weeks earlier. indexOf returned -1, -1 is less than anything, and the check had passed every run since.
  • One guard tested a hand-written copy of the SQL instead of the production webhook.

The pattern behind most of the 44: the guards were written against text, not behaviour. We read the route file as a string and looked for names. Such a guard fails both ways: an honest refactor turns it red for nothing, and a real break slides past because the names are still there.

One guard, before and after

The consent route does three things in order: write the user's choice to the auth provider, append to the audit log, unpublish their renders from the public showcase. The guard checked that order with indexOf on the source. Nine checks, all green. The mutation that beat it, and the rewrite:

// production: awaited; a failure returns 500 to the user
await client.users.updateUserMetadata(userId, { ... });
// mutation: same call, same position in the file - all nine checks still green
void client.users.updateUserMetadata(userId, { ... }).catch(() => {});

// rewritten guard: cut out the statement itself, not a name near it
const i = ROUTE.indexOf("client.users.updateUserMetadata");
const stmt = ROUTE.slice(ROUTE.lastIndexOf("\n", i) + 1, ROUTE.indexOf(";", i) + 1).trim();
ok("write is awaited, not fired in the background", /^await client\.users\.updateUserMetadata\(/.test(stmt));
ok("a failed write is not swallowed", !/\.catch\(/.test(stmt));
ok("and not silenced with void", !/^void\b/.test(stmt));
Enter fullscreen mode Exit fullscreen mode

Under the mutation the order of the three calls did not move, so the old checks passed. The meaning was gone: the write is fire-and-forget, the failure is swallowed, the audit log records "revoked", the user believes they opted out, and the showcase keeps publishing because the provider still says "allowed".

The rewrite is still text, but the text of the statement itself, and the mutation that defeated the old version is now a permanent check. Where the code can be imported, we go further: the guard loads the real function through a loader and runs it over a table of inputs.

The runner

Every one of those experiments was a throwaway script, fifty lines each, and each hit the same traps: CRLF on the Windows working copy versus LF in git, a snippet that matched twice, a "mutation" that deleted a line and could not be reverted by an empty pattern (one stayed in the code and was caught on a second pass).

On 28 August it became one tool. A job is JSON: the guard command plus named substitutions.

{
  "guard": "node tmp/qa/geometry-escalation-proof.mjs",
  "mutations": [
    {
      "file": "src/lib/draft-rank.ts",
      "name": "a judge that errored no longer ranks last - an empty verdict beats a checked image",
      "from": "  if (input.errored) return Number.MAX_SAFE_INTEGER - 1;\n",
      "to":   "  void input.errored;\n"
    },
    {
      "file": "src/app/api/generate/route.ts",
      "name": "escalation became unconditional - furniture-placement failures go to the expensive engine too",
      "from": "if (!economyMode && lastGeometryBroken && attempt < MAX_ATTEMPTS) {",
      "to":   "if (!economyMode && attempt < MAX_ATTEMPTS) {"
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

The loop is short; the rules around it are what we paid for:

if (!runGuard()) fail("guard is already red; it proves nothing");          // 1. green before

for (const m of mutations) {
  if (!m.to.trim()) { bad(m, "empty replacement is a deletion, not a mutation"); continue; } // 2
  const src = original(m.file);
  const hits = src.match(rx(m.from, "g"));                      // rx() tolerates \r?\n  // 3
  if (!hits || hits.length !== 1) { bad(m, "snippet must occur exactly once"); continue; } // 4
  fs.writeFileSync(m.file, src.replace(rx(m.from), m.to));
  const stillGreen = runGuard();
  fs.writeFileSync(m.file, src);
  stillGreen ? bad(m, "GUARD STAYED GREEN") : good(m);
}
assertEveryFileRestoredByteForByte();                                         // 5
Enter fullscreen mode Exit fullscreen mode

Mutations are named and semantic on purpose. Classic mutation testing flips operators automatically; we write each mutation as a sentence about how a customer would get hurt, because that sentence also documents why the guard exists. The trade-off is coverage: we test only the breaks we thought of.

What the mutations found in production code

The day the runner was born, a run on a new guard did something we did not expect. The function that decides whether a payment-return URL is stale had an explicit "payment time is in the future" branch. Six mutations turned the guard red. The seventh, removing that branch, changed nothing: a future timestamp gives a negative age, and a negative age is never above the threshold. The branch was a crutch the guard could not tell from working code. We removed it and wrote the reason above the function.

In the same session, 35 mutations across seven fixes found three guards checking the wrong thing: a parser that stopped at the first closing brace and broke on a nested object (four of six mutations passed on obviously broken code); a "the catch block logs" check that stayed green when the log left one branch because a neighbouring branch still had one; and a scanner for prototype-key guards that cleared a whole file if the guard appeared anywhere in it.

Where it stands

  • 93 mutation jobs, 867 mutations, in the repo next to the guards.
  • 171 guards in the runner's list, 12 named skips (they cost money, run on a calendar or need a fresh build); the list is compared with the disk on every commit.
  • Written rule: a new guard without a mutation run counts as not written.
  • The pre-commit hook checks only the registry, in 0.08 seconds. The full run takes minutes, needs the network and a database, and can go red from someone else's outage; a hook like that gets --no-verify on the first urgent fix, and a bypassed hook is worse than none. The full run stays manual and mandatory before a merge.

What still fails

Hand-written mutations come from the same person who wrote the guard, with the same blind spots. On 5 October a mutation named "only seven articles left in the links block" removed 5 of 17 links and the guard stayed green. The mutation was too weak, not the guard; it now removes every link. Nothing finds weak mutations except somebody noticing.

Many guards still read source text. Executed guards are better, but much of our logic sits in route handlers tied to auth and a database, and extracting it just to test it is a refactor we have not earned yet. Text guards are honest only as long as their mutation set is.

There is no CI. The guards run on a laptop before a merge.

And the first finding of all: 44 of 70 is not a number about carelessness. Every guard was written in good faith after a real bug. A green check is a claim, and until something tries to make it red, it is an unverified one.

We build this at AI Flip Room.

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

The table and the headline don't quite line up: the groups sum to 71 guards and 42 holes, the title says 70 and 44. If a guard can hold more than one hole (the last row says 9 in 7 files), these are holes per guard, not the share of guards that stayed green on a real break, so I'd report both.

The group split is the more useful result. Billing, consent, privacy and SEO at 21 of 28 against image engine and render pipeline at 5 of 19 gives Fisher p of about 0.002. That is not noise. Against infrastructure at 7 of 16 it is about 0.05. The 44 of 70 overall is 63% with a Wilson 95% interval of roughly 51-73%, so "about two thirds" is fair but the audit pins the kind of guard more tightly than the total.

One thing I'd add to the 867 mutations: score each guard by the fraction of its mutations that turn it red, not just red/green on any hole. A guard at 4 of 6 (the brace parser) and one at 0 of 9 need different fixes, and the per-guard score also gives you a number to watch when the hand-written mutations come from the same person as the guard. Is the 5-of-17 links miss the only one you found from a too-weak mutation, or did the rewrite sweep find others?