From September 23 to October 2, a customer kept asking our AI app builder for changes that never held.
They had built a sizeable section of their app with it in early September: a form-heavy flow, an admin view, a database table behind it. About seventy edits over three weeks, each one asked in plain language, and the generated JavaScript grew to roughly 70,000 characters without breaking. On the evening of September 23, the day we moved the builder to a new model, the next edit came back with 8,000 characters of JavaScript. The one after that, 1,400. Then 874. Then 510. Each time, the customer restored the previous version and asked again. Thirty-seven restores and manual saves later, the report reached me on October 5: the code is gone, why?
One line of context so you know where I stand: I run engineering at GoodBarber, an app platform. Our AI Extension Builder lets a customer describe a section of their app; a model writes the HTML, CSS and JavaScript; then the customer asks for changes, one at a time. We call each change a refine. On September 23 we switched the builder to GPT-6 Luna, for every user at once.
This post is the investigation, the check we added, and how we calibrated it on 5,960 production refines. The short version: the model changed how it answered, and our pipeline had never checked what an edit kept.
A word of caution before you read on. This is a bug fix, written and pushed the day the report came in. The thresholds come from five weeks of production data, the check has not had time to prove itself in production yet, and the open items at the end are real. It probably deserves more hindsight and more work than one day gave it. Read it as a field report, not a recipe.
How a refine works
A refine runs in two calls. The first one plans: it reads the current app and the request, and declares which fields the change must touch, for example ["js"]. The second one writes the code. It must return each field in that scope in full, and mark the others with a sentinel, @@REFINE_UNCHANGED@@, so we don't pay for 30,000 characters of HTML the model would copy verbatim. The server swaps the sentinels for the parent's fields, validates the result (well-formed fields, our JavaScript API used the way it should be, a list of rules models tend to forget), and publishes.
When validation fails, we retry, up to twice, with the validation error as feedback. Most failures are a forgotten rule, and one corrective prompt fixes most of them.
Look at what the validator judges: the output. Is this a valid app? Not: is this still the app the customer had?
The first diagnosis blamed the model
I ran the investigation with Claude Code, with read-only access to the production database and the code. Within five minutes it had a tidy story in three links:
- Targeted refine was on in production.
- The new model, asked to write the JavaScript in full, often wrote a patch instead. One real output kept the parent's HTML and CSS through the sentinel, and replaced 68,000 characters of JavaScript with a single new function plus a call to
window.initExistingMicroApp(), a function that exists nowhere. The merge assembled it, the validator found a well-formed app, the finalizer published it. 68,000 characters of live code became 1,400. - The retry made it worse: when a fragment failed validation, the retry started from the fragment.
Output tokens agreed: between 1,200 and 8,000 per refine for the new model on this app, between 25,000 and 45,000 for the previous one, which rewrote the whole file every time.
Every link was true. The weighting was not. The model was the cause, the retry an aggravating factor, and the first move proposed to stop the bleeding was to put the refine step back on the previous model, or to turn the targeted mode off. I typed back: "you're blaming the model. Explain." And I asked whether the real fix wasn't a deterministic check on the process.
What the data said once we looked at all of it
Second pass, on every two-phase refine in production since September 9 whose parent had more than 5,000 characters of JavaScript: how many shipped with less than 40% of it?
| period | refines | lost more than 60% of their JS | sections affected |
|---|---|---|---|
| previous model, Sept 9 to 23 | 3,157 | 45 (1.4%) | 37 |
| GPT-6 Luna, Sept 23 to Oct 5 | 1,468 | 111 (7.6%) | 67 |
Two things in that table. First, the previous model went through the same hole, five times less often. Some of those 45 may be rewrites users asked for; nothing in the pipeline would have stopped them either way. Second, the 111 did not look like the story above. Only 13 had the patch shape, and 6 of those went on through the retry. 78 had gone through our retry in all. The other 26 were shorter rewrites on the first try, published without a question.
So the retry deserved more than an "aggravating factor" line. Simplified, it chose its starting point like this:
if failed_output_is_missing_a_field and parent_spec:
base = parent_spec
elif isinstance(failed_output, dict):
base = failed_output # "here is your answer, fix the mistake"
That is right for the job it was written for: you forgot one rule, here is your output back, fix that. It breaks when the output is a fragment. The field is not missing, it is present and nearly empty, so the retry is built on the fragment and the model never sees the original JavaScript again. The token counts show it. On the customer's app, the retry received 35,708 input tokens against 56,886 for the refine before it: about 21,000 fewer, roughly the size of the JavaScript that had vanished. With nothing left to preserve, the model did the sensible thing and wrote a small, complete app. 405 characters of HTML.
So the model was the trigger. The cause was ours: a validator that only judged the output, and a retry that could turn an incomplete answer into the new baseline.
The check: count what survives
The fix is a pure function. It compares the output to the spec the model was given, with no model in the loop. It had to catch mass disappearance without blocking legitimate edits: removing a feature, refactoring, renaming.
A diff would not do; a refactor produces a huge diff and loses nothing. So the check counts names. In JavaScript, declared function names. In HTML, element ids. Those are what the rest of the code refers to. And it measures size.
A field counts as destroyed when it keeps less than 30% of the parent's names and less than 40% of its size. Both conditions are required. A refactor that renames every function keeps its size. A cleanup that shrinks the code keeps its names. Fields under 3,000 characters, or with fewer than 3 names, are skipped: ratios on tiny fields mean nothing.
_JS_NAMES = re.compile(
r"function\s+([A-Za-z_$][A-Za-z0-9_$]*)\s*\("
r"|(?:const|let|var)\s+([A-Za-z_$][A-Za-z0-9_$]*)\s*=\s*(?:async\s+)?"
r"(?:function|\(|[A-Za-z_$][A-Za-z0-9_$]*\s*=>)"
)
_HTML_IDS = re.compile(r"\sid\s*=\s*[\"']([^\"']+)[\"']")
MIN_PARENT_CHARS = 3000
MIN_PARENT_NAMES = 3
MAX_KEPT_RATIO = 0.3
MAX_SIZE_RATIO = 0.4
def destroyed_fields(parent_spec: dict, spec: dict) -> list[FieldLoss]:
"""Fields of ``spec`` that lost most of what ``parent_spec`` had."""
losses = []
for field_name, unit, extract in _CHECKED_FIELDS: # js -> functions, html -> ids
before = str(parent_spec.get(field_name) or "")
if len(before) < MIN_PARENT_CHARS:
continue
names = extract(before)
if len(names) < MIN_PARENT_NAMES:
continue
after = str(spec.get(field_name) or "")
kept = len(names & extract(after))
size_ratio = len(after) / len(before)
if kept / len(names) < MAX_KEPT_RATIO and size_ratio < MAX_SIZE_RATIO:
losses.append(FieldLoss(field_name, unit, kept, len(names), size_ratio))
return losses
A regex is not a parser. This one misses class methods and functions declared inside object literals. It doesn't need them: it is a smoke alarm for mass disappearance, not an understanding of the code. An app written entirely in classes yields fewer than 3 names and is skipped, which is the safe direction to fail.
Where it runs:
- On every refine step: the first output, the repair pass, each retry. It compares with the nearest ancestor that has code, because a chat answer in between ("what does this form do?") has no code of its own.
- A destroyed output is never the retry's base. The retry starts from the parent's code. The validation error doubles as the retry's instruction, so it carries the user's request: Your output dropped most of the existing app (js keeps {kept} of its {total} functions and {n}% of its size). This request changes an existing app: keep every existing element, function and behavior, and return the COMPLETE html, css and js with only this change applied: {user request}
- If the retries destroy too, the generation fails. Nothing is published, and the live app stays as it was.
Calibrating without labels
Thresholds picked by feel would have been a guess, and we had no labeled set of destroyed edits. Our users had labeled them for us, with what they did next.
We took every two-phase refine since September 1 whose parent had more than 3,000 characters of JavaScript: 5,960 of them, pulled in 25 seconds. 4,207 ran on the previous model and 1,576 on the new one; the other 177 ran on an older model and stay out of the comparison. For each one we looked at the next action in the same section:
- continued: the next edit was built on this one. The user accepted it, at least for a while.
- restored: the next action was restoring an earlier version.
Then, for each refine: the share of the parent's function names kept, the share of its ids kept, the size ratio.
The first number that came out had nothing to do with thresholds. The share of refines followed directly by a restore was 0.45% with the previous model (19 of 4,207) and 3.81% with the new one (60 of 1,576). More than eight times as often. That number sat in our own table for twelve days, and nobody computed it.
Then the candidate rules, on JavaScript:
| flag when names kept < and size < | kept edits flagged, previous model (of 3,517) | kept edits flagged, new model (of 1,257) | restored edits flagged (of 79) |
|---|---|---|---|
| 20% and 30% | 15 | 60 | 25 |
| 30% and 40% | 25 | 67 | 27 |
| 30% and 50% | 33 | 76 | 27 |
| 50% and 50% | 46 | 87 | 29 |
We kept 30% and 40%, and added the same rule on HTML ids, which is why the numbers that follow are slightly higher than the table's 25 and 67. Among the 30 restored edits that had lost more than half of their parent's functions, it flags 27. Among the edits users kept building on, it fires on 0.8% with the previous model (27 of 3,517) and 5.8% with the new one (73 of 1,257).
The 5.8% is not all false positives. "Continued" means the user built on it, not that it was fine; some people kept going on a broken app before they noticed. Reading a sample of 40 of those flagged edits (other customers' prompts, so no examples here), roughly one in four was a rewrite the user had really asked for.
No exception for "start over"
That one in four forced the only real design decision. A user who writes "drop all of this and build me a simple contact form" wants exactly what the check blocks.
Two options were on the table:
-
Let the planner declare it. A field like
replace_app: truewhen the request is an explicit rewrite, and the check is skipped. - No exception. To start over, the user clears the history. The next generation has no parent, so there is nothing to compare with.
I took the second. The first hands the exception to the component whose judgment we are guarding against. If the model errs in the "rewrite" direction, the check is bypassed exactly as before, now with a flag saying it was intentional. A deterministic guard that asks the model for permission is not deterministic anymore.
The price is real: some legitimate rewrites will now fail where they used to go through. That gets fixed in the failure message, not in the check.
Replay on the real thing
Before shipping, we replayed the check on the customer's actual refines:
- all 13 refines from the previous model pass the check;
- 12 of the 13 destructive refines from the new model are blocked;
- one still passes. It kept more than 30% of the function names and lost 70% of the JavaScript: the names survived, the bodies did not. The rule as written lets it through, and I would rather know that than tune a threshold to a single sample.
Eleven new tests, including the real shape: HTML and CSS through the sentinel, a JavaScript fragment that calls a function defined nowhere. The four end-to-end ones fail without the fix and pass with it, including one that asserts nothing gets published when the retries destroy too. No new failure in the rest of the suite.
The agent wrote the check, the retry change and the tests. The calls were mine: no exception for rewrites, keep two retries, ship.
A second bug on the way: the save that forgot the database
Reading the customer's history surfaced a bug that had nothing to do with the model. Saving from the Code tab rebuilt the stored spec field by field from the editor's content, and the list of fields it copied was short. backend_requirements, the declaration of the app's database tables, was not on it. After one save from the editor, this app declared zero tables. design_spec and technical_summary went the same way; since September 1, 322 of the 1,189 saves in production had dropped fields like this. The same save also froze into the source the runtime glue the platform injects at display time, and appended it again at each save. Three copies, in this app.
The fix inverts the construction. Before, the save listed what to keep:
return MicroAppSpec(
name=source_spec.name,
description=source_spec.description,
html=html_content or source_spec.html,
# ... nine more fields, and nothing else
)
After, it copies everything and overrides what the editor changed:
return dataclasses.replace(
source_spec,
html=html_content or source_spec.html,
css=css or source_spec.css,
js=strip_proxy_runtime(js or source_spec.js or ""),
)
Copy, then override. A field someone adds to the spec next month will survive saves without anyone remembering this function exists.
The part that stung: the bug was known. Three tests in the suite described it, marked as expected failures. They pass now, and the markers are gone. An expected failure is a bug report that stopped asking to be read.
What is still open
- The failure message is generic: "We couldn't generate your component. Please try again." After a rejected destruction, trying again will probably loop. It should say what happened, and point to Clear history for someone who really wants to start over.
- The save still bakes in the display-time glue. It strips one runtime block, not the others. The copy-then-override fix stopped the lost fields; the duplicated glue is still on the list.
- The check prevents new losses. It does not repair what was already saved. Versions published before the fix keep what they lost; for the customer who reported it, that means a restore of their last good version, by hand.
- We had no canary. Every user moved to the new model at once. On the next switch, a slice of apps goes first, and the restore rate is the number I watch.
What I'd take away
- For an edit, validating the output is not enough. Check what it kept from the input. Counting names and size is crude, and it catches what matters.
- Never build a retry on the output you just rejected, unless you know the rejection is about a detail. "Missing" and "present but gutted" are different failures.
- Your users already label your outputs. Restore, undo, regenerate, abandon: that is ground truth for calibration, and it costs one SQL query.
- A model switch runs a test suite you never wrote. The new model broke a contract our prompt stated ("write these fields in full") and our code never enforced. The previous one happened to honor it. That is not the same as the pipeline being correct.
Top comments (7)
The passing destructive refine that preserved names but lost bodies makes a valuable permanent regression fixture. I would keep it separate from threshold calibration and add a few legitimate shrinking refactors as controls, so tightening the smoke alarm does not merely fit that one incident.
Your retry change is the stronger invariant: every rejected attempt must recover from the original accepted ancestor, including a failure on the second repair. An end-to-end test that records each retry's parent version and asserts the published version never changes after exhausted retries would protect that behavior even if the disappearance heuristic evolves.
Thanks. I checked both against the test suite, and each one found something.
On the first: our only shrink control empties every function body, keeps every name, and asserts the check lets it through. It was written to prove the rule needs both conditions. It is also the very failure you describe, filed as a pass. Your version is better: real shrinks as controls (dead code removed, a feature dropped on request), and the refine that got through pinned as a known miss, asserting it passes today with a comment saying why. Whoever tightens the rule then sees which side each test lands on. Telling those cases apart will take another signal than names and total size, not a tighter threshold.
On the second: the exhausted-retries test checks the call count, the failed status and that nothing gets published. It checks the base only on the first retry; the second retry's base is never asserted. Your test closes that; I'm adding it.
One nuance on the invariant: only destroyed attempts restart from the accepted version. An attempt rejected for a forgotten rule still retries on its own output, since that's the cheap fix. What keeps it safe is that the anchor never moves: every attempt is measured against the last accepted version, not against the base it was handed, so drift can't compound past the threshold.
That is a fair correction: I overstated the restart invariant. The execution base can remain the rejected output for a rule fix; the comparison anchor must remain the last accepted version. Those are separate properties.
A mixed-failure sequence would make the distinction explicit: first reject an otherwise intact attempt for a forgotten rule, then produce a destructive attempt during that correction. Assert both the base chosen for each retry and the unchanged comparison anchor. That would test the transition between the two recovery paths, rather than only repeated failures of one kind, while keeping legitimate shrink cases separate from the known miss.
Reading the loop again to answer you, I had the base part slightly wrong too. Bases don't chain. The base is chosen once, from the first rejected output: that output itself for a rule fix, the last accepted version if it came back destroyed or missing a field. Every retry starts from that same base; only the error message moves forward.
So in your sequence, the second retry starts again from the intact first attempt, with the destruction message as its only feedback. The forgotten rule is no longer mentioned anywhere. It fails safe (nothing ships unless it passes both the validator and the check), but it can spend the last retry fixing the wrong thing. That is a better reason to write your test than the one I had: assert the base of each retry, the fixed anchor, and that the last feedback still carries every error seen so far. Today that last assertion would fail.
The output collapsing from 70,000 characters to 510 after the GPT-6 Luna swap is a brutal way to surface a latent bug. Your point lands: the validator checked the output but never what the edit kept, and I hit that same blind spot grading agent edits, so I now diff each result against its parent. How did you set the preservation threshold without flagging a legitimate rewrite?
The line about the validator only judging the output stuck with me. Counting surviving function names and size is a clever invariant, and it is cheap. In my world that is the same lesson as verifying LLM outputs against something outside the pipeline itself.
Agreed, with a twist: the reference here wasn't outside the pipeline. It was the input, which the pipeline already had and never looked at. The validator read the output; nothing compared it to what came in. The outside part came later, from users: their restores are what told us where to put the thresholds.