DEV Community

Seth Wheeler
Seth Wheeler

Posted on Originally published at sethwheeler.dev

Taking the Language Model Out of a Build Planner

I have a tool called appgen that turns a sentence into a running application with no language model anywhere in it. You type "a support desk system with priorities, comments, search and closing tickets"; it writes a dependency-free Python file, starts it on a private port, exercises every feature you asked for over real HTTP, and only then shows it to you. Around it sits a larger stack called olo, which adds a planner in front. The planner is a local language model. Its job is to turn your sentence into a build plan: seven fields, six of which are a choice from a closed vocabulary (the program kind, the language, the shape, the noun the app is about, and so on) and one of which is the request passed through unchanged.

That planner is the only place in the architecture where a neural network does anything load-bearing. The repository's own README puts it as a negative: of at least six separable jobs a language model does, five have measured non-neural replacements in this tree and the sixth, proposing a program in the open domain, does not. This post is what happened when I measured what the seat buys, by taking the model out and putting nothing in its place.

Over 88 requests taken complete from florinpop17/app-ideas, with every line of code downstream of the plan held constant, the model's plan rescued zero requests that the empty seat lost. That holds under all four ways I can score the run, and it is the only thing about this round that does.

No link to the code: the repo is private research. The figures all come out of results files or commands, named as they appear.

Why this arm did not already exist

Two earlier rounds had measured this seat, and they have the same shape: one set the model alone against the whole stack, the other re-ran the stack through repaired plumbing. Both have a model on both sides. Neither says what the planner buys. olo differs from appgen in more than its proposer; there is a router, a plan validator, an entity read-back, and the code that assembles appgen's command line, all sitting in between. Any margin you attribute to the planner is the planner plus all of that.

A third figure had been sitting in the tree beside them, unread against either. An older round ran bare appgen over this same corpus and published 21 of 88, above the 15 the full stack scored. Nothing in the repo had ever read those two numbers together.

The one edit that decides the experiment

olo could not leave the seat empty, so the round starts by building a seam. --planner defer declines every field and passes the sentence through. --plan PATH reads a plan computed elsewhere off disk. Both run through the same validator and the same command-line assembly the model's plan does.

The load-bearing change is a single branch in that assembly: a declined field is now absent from the command line rather than filled with appgen's default. That is not tidiness. appgen reads the sentence for its program kind and its language only when the flag is absent, so --kind web is not "the default, stated"; it suppresses a reading the layer below already does. A version that helpfully filled in web would have measured appgen with its own router disabled and reported it as appgen.

Three tests guard that, and each was shown to bite by planting the defect it names in the source and checking the test goes red, rather than by being read and believed. The first plant found something I would not have: the tests had been appended below sys.exit(main()), so they had never run. A file of assertions that cannot execute looks exactly like a file of assertions that pass.

The arms

Each arm differs from its neighbour by one thing. bare is appgen driven with the request and nothing else; defer is the full olo stack with the seat empty; model is the full stack with a local qwen3-coder:30b planning. Delivered counts the requests where something was built at all; the rest are refusals, which is appgen declining a request outside its grammar rather than guessing. Precision is working over delivered.

arm working, of 88 delivered precision
bare 29 31 0.9355
defer 29 31 0.9355
model 22 32 0.6875

Those two 29s are not a coincidence of the marginal totals. Row for row, bare and defer return the same 88 verdicts, with nothing discordant. Everything olo does that is not the planner costs nothing and buys nothing once the seat is empty.

Marginal counts throw away the pairing, so the comparison that matters is the discordant one: of the requests where the two arms disagree, how many did each win? A 7 – 0 means seven requests where one arm produced a working app and the other did not, and none the other way.

reading n defer model discordant p
published scorer, all rows 88 29 22 7 – 0 0.0156
entity clause forgiven 88 31 27 4 – 0 0.125
undrivable rows dropped 83 25 22 3 – 0 0.25
both corrections 83 27 27 0 – 0 1.0

The two corrections are real and neither is a thumb on the scale. The scorer fails an entity (the noun the generated app calls a row, which becomes its table name and its URL path) when that noun is not a word of the request. That is a rule an earlier round paid for and one I would keep. It is lexical, though, and the gold labels for the same task are semantic. A proposer answering post for "A clone of Facebook's Instagram app" is therefore broken by one instrument and correct by the other. The second correction is that an artifact the harness could not drive at all is recorded as the harness's failure by this tree's own definition, and the model arm carries five of those against defer's zero.

What survives every reading is that the model wins nothing. What does not survive is any claim that the empty seat beats the model: two of the four readings have them tied or nearly so.

The number that did resolve was measuring the harness

I wrote the first draft of this with the 7 – 0 and p = 0.0156 at the top, and it did not survive its own follow-up.

Two of those seven rows are Create a callable engine to play the Battleship game and Query NASA's Exoplanet Archive. The model planned both as a JSON API, and the driver at the time could only exercise an artifact by fetching / and finding an HTML form, so an API build came back as could not be driven whatever it did. Once a driver existed that exercises an API on its own published endpoint, replaying the stored command line from those two rows, contacting no model, returns working for both. That is what bare and defer also scored.

So they were ties, not model wins, and the headline is untouched. But hold the other 86 rows fixed and the count becomes 5 – 0, p = 0.0625, which is above 0.05. The only one of four readings in which this round resolved anything was resting on two rows its scorer could not see, and I had already recorded that reading as the one that held.

Where the five remaining rows went

I designed this round on the theory that a planner earns its keep on the enum fields, which is what a plan mostly is. The rows it lost say otherwise. After the two API ties above, five remain, and all five carry the plan web/python, which is exactly what appgen resolves to on its own. The enums did not lose them:

request the model's plan what happened
A clone of Facebook's Instagram app web/python entity post
Browse, Find Ratings, Check Actors… web/python entity movy
Click list item to display item details web/python entity plant
Keyboard Event Values web/python exit 1, no artifact
Review and test your knowledge through Flash Cards web/python exit 1, no artifact

Three of the five are the entity field, and it is the same shape as the command-line finding one layer over: --entity overrides a derivation appgen would have done from the sentence itself. A field filled by the layer above suppresses the layer below's own reading of the request. That is the mechanism this whole round is about, and it shows up twice in one stack, once in the plumbing and once in the result.

That mechanism also gives the fix, and it shipped: when the plan names an entity the request does not contain, olo now withholds the field instead of passing it, and appgen derives its own. Two of those three rows come back working with nothing else moving. The third, movy, is not a reading failure at all. The model answered movies, which stems to movie and is therefore a word of the request by the rule the scorer applies; appgen then singularised it into a non-word before rejecting the result. That row has a separate cause and a separate fix, and it gets its own write-up.

Both published incumbents had expired

Re-running the two baselines rather than quoting them is what turned up the last thing worth reporting:

figure published re-run here drift
bare appgen 21/88 29/88 +8
the full stack 15/88 22/88 +7

Neither number is wrong. Both are expired: true at the commit they ran at, and false on a tree that has moved since. A round quoting either would have compared its new arm against a stack that no longer exists, and it would have looked completely reasonable doing it.

What this does not say

It does not say planning is useless. It says this stack's planner fills fields the layer below already derives, and that filling them suppresses that derivation. A planner over fields appgen does not read at all, like CSV export or email or validation, is untouched by any of this.

It does not say the empty seat is better than the model. That holds under the published scorer and dissolves under either correction. The measured claim is the weaker and more robust one: removing the model costs nothing here.

And it is one seat of two, one corpus of 88, one local model at one temperature, one run. The planner has not been removed from the shipped tool on the strength of it; what has changed is the one guard above, and the seam that made the empty seat runnable at all now exists for the next round to use.

What generalises

The part I would carry somewhere else is not the negative result. It is that a default supplied from above is not neutral. Filling a field with the same value the layer below would have computed looks like a no-op and is not one, because the layer below stops computing it. Three of the five losses in this round are that. The command-line branch I had to write before the experiment could run at all is the same bug in the opposite direction: had I filled in web "for consistency", I would have measured appgen with its own router switched off and published the number as appgen's.

The other half is about which number I led with. The reading that resolved was the reading my instrument was least able to support, and it took building a better driver to find that out. The claim that survived all four scorings was available on day one and it is the boring one: the model won nothing.

Top comments (0)