DEV Community

Robert
Robert

Posted on Originally published at neuragrowth.co

Twelve Rejections, And The Gates Were Wrong Four Times

Overlapping gates, on purpose

A model writes an episode as structured data and a renderer turns it into video. Between them sit three layers: types at parse time, a checker that returns a list of complaints in plain language, and a set of mechanical repairs for the faults that need no judgement. Anything the repairs cannot fix goes back to the model with the complaints attached.

That design has earned its keep. It is why a prop that does not exist cannot reach the renderer, and why the narrator cannot claim a flower starts with the letter P. The gates are not the problem. Four specific gates were.

A proxy that outlived its reason

The long-format episode had a duration floor of 240 seconds. Six consecutive drafts came back between 205 and 230 seconds. Each time, the rejection message did its job well and told the model exactly what to do:

too short by 35s (205s, wanted 240-330s). Add 1049 frames:
raise minFrames on several beats, or write about 412 more
characters of speech
Enter fullscreen mode Exit fullscreen mode

Which is a polite way of asking for padding. The drafts were not thin, they were finished: a beginning, a complication, a false attempt, a resolution. What the floor had originally stood in for was "is there enough in this film", asked at a time when literally nothing checked what was in it.

By the time it was rejecting these drafts, six real content checks existed: this week's colour, this week's animal, this week's letter, the count of character dialogue, the change of location, and a rule that the narrator may not name anything not on screen. The proxy had been replaced by the thing it was proxying for, and nobody retired it.

I lowered the floor to 195 and wrote the reason next to it, because "we lowered a threshold until it passed" is exactly what this looks like from outside, and it is worth being precise about why it is not. Lowering a bar to fit a result is cheating. Retiring a stand-in once you have the real measurement is housekeeping. The test is whether you can name the specific checks that replaced it.

A rule about a number nothing reads

One draft came back with twenty-one complaints. They were all the same complaint:

beat 23: prop t1 is floating at y=1200, put it on the ground
beat 24: prop t1 is floating at y=1200, put it on the ground
beat 25: prop t1 is floating at y=1200, put it on the ground
... nineteen more ...
Enter fullscreen mode Exit fullscreen mode

The rule said an object must stand on the ground, be held, or be explicitly resting on another object. Nothing levitates. It had caught real mistakes before.

Then I read the renderer. For props in the front row, it ignores the written coordinate entirely and lays them out on its own line, computed from how many share the beat. The y value the gate was policing is never read. Worse, I had recently made front the default, so the rule now applied to almost every prop while governing almost nothing.

It was a rule about nothing that could still kill a four minute film. And it had a second cost I did not see at first: twenty-one identical lines drowned the three real complaints in the same list, so the repair prompt became a wall of one sentence with the number changing.

A word list that ate a character name

One mechanical repair splits an over-long sentence at a comma. To avoid producing fragments, it checks that neither half opens with a word that leans on the other: subordinating conjunctions, relative pronouns, that sort of thing. In Polish that list includes bo, meaning "because".

The two characters in this series are called Pip and Bo.

So the repair refused to split any sentence mentioning one of the two protagonists, which in a two-character show is most of them. The rule was correct, well motivated, tested, and silently disabled half its own coverage. The fix is one line, and it is the kind of thing you only find by reading the sentences it declined to touch:

# Capitalised "Bo" is the character. Lower case "bo" is "because".
if head != "Bo" and head.lower() in NOT_A_SENTENCE_HEAD:
    return False
Enter fullscreen mode Exit fullscreen mode

Two rules that made the task impossible

The nastiest one. A new rule said the weekly summary episode must show the letter that week taught, as a symbol on screen. An older rule said that series never puts a symbol on screen, because a stray digit floating over a snail is a common artefact.

Between them, no episode could ever satisfy both. Not a hard failure, not a crash: a brief with no solutions, which the generator would have failed on for as long as it existed.

The older rule was written against accidental symbols, and said so in its own comment. So the fix was to narrow rather than remove: a symbol is allowed in this series only in a beat whose text actually mentions it, decided by the same parser that already validates "P as in ball". One definition, two readers.

Run every new gate against work you already accepted

The check that would have caught all four is cheap and mechanical: before believing a new rule, run it over the artefacts you have already approved.

for name in ACCEPTED_EPISODES:
    problems = check(load(name))
    print("OK " if not problems else "BAD", name, problems[:1])
Enter fullscreen mode Exit fullscreen mode

A new rule that fires on known-good work is wrong, whatever its reasoning sounded like. A new rule that fires on nothing at all is also suspect, and that is the second half of the habit: run it against a known-bad artefact too, and watch it catch what you wrote it for. I once shipped a regex that could never match anything because an escape had been mangled into a control character, and it looked like a working rule for a week.

Both halves take a minute. I ran them on the last three gates in this story and each one changed what I shipped.

If you gate generated output

  • Write down what each threshold stands in for. Proxies are fine. Proxies that outlive the measurement they were substituting for are not, and without the note nobody can tell which one they are looking at.
  • Check that the consumer reads the value you are policing. A rule about a field the renderer ignores is a rule that can only cost you.
  • Collapse repeated complaints. One fault reported twenty-one times buries the three that matter and turns your repair prompt into noise.
  • Watch for collisions with your own vocabulary. Generic word lists meet domain nouns, and the failure is silent: the rule keeps passing, it just quietly stops applying.
  • Two rules can be individually right and jointly impossible. If a class of input starts failing forever, check for a contradiction before assuming the generator got worse.
  • Test new rules both ways. Against accepted work, so it does not fire. Against known-bad work, so it does.

Originally published at neuragrowth.co. I run a one-person digital-products studio and write up what breaks in production.

If you write CLAUDE.md files, I keep a set of working templates here: neuragrowth.co/free/claude-md-templates.

Top comments (2)

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

The two-way habit at the end is the right instrument, and this story shows where it stops: it tests precision. Of the four failures here it catches the proxy, which fires on accepted work, and the contradictory pair, which fires on everything. It is structurally blind to the other two, because those are failures of reach rather than of firing.

The floating-prop rule is the clear case. Run it over accepted episodes and it passes: the coordinate it complains about is written, plausible and wrong, and the renderer lays those props out from its own line, so nothing in the accepted set ever looks broken. What an accepted-work run cannot ask is whether the value in the message is on the path to the output. That question has a mechanical form and it is not a pass/fail run: perturb the policed value and compare what the consumer produces. Render the same beat twice, y at 1200 and y at 0, diff the outputs, and if the two are byte-identical then the gate is reporting on a coordinate nothing reads. One ablation per policed field, no judgement involved. You found that one by reading the renderer, which is the expensive way.

The second blind spot is the repair that declines. A rule that may not act has two outcomes, and both of your halves are satisfied when it always declines: it does not fire on accepted work, and it does not fire on the known-bad artefact either, if that artefact happens to be one it silently skips. So the instrument is a counter on the refusal itself. Print declined-with-a-reason beside applied, per rule, and the Polish case becomes visible without reading a sentence: a repair that declines nearly every sentence naming either protagonist has a refusal rate that looks nothing like a rule which occasionally finds nothing to do. The same number is what tells you whether the collision is rare or is the rule.

Both cost one measurement each, and both read a different thing from the pair you already run. Your two ask whether the rule fires correctly. These ask whether it is attached to anything. A gate can pass both of yours while governing nothing, and in this story two of four did.

Which is also why collapsing repeated complaints is more than prompt hygiene. Twenty-one identical lines are the signature of one variable reported over twenty-one objects, and that count is what an ablation would have contradicted: a value that moves twenty-one complaints and no output.

Collapse
 
arhancanli profile image
Arhan Canli •

Two rules that make the task impossible is the case I'd put first, because it's the one that looks like the work is bad rather than the gates. We hit the same shape in a research pipeline: two acceptance thresholds, each reasonable on its own, that no candidate could satisfy together, so every failure read as "the candidates aren't good enough". The check we added sits next to your two: before trusting a set of gates, build the best artefact you can (by hand, or by loosening the generator) and confirm the whole set can pass it at all. If the ideal input fails, the gates disagree with each other, and no amount of regeneration will fix that. Your known-bad half matters just as much; a gate that never fires looks exactly like a gate that's working until you plant the fault it was written for.