appgen is a tool of mine that turns a sentence into a running application with no language model in the generating path. Something has to read the sentence first, and one field of that reading does more work than the rest: the entity, the noun the generated app calls a row. It becomes a table name, a form field and a URL path, so getting it wrong means the right program gets built about the wrong thing.
Today that field is read by a hand-written keyword rule. A local language model reads it better, and I am not simply using the model, for two reasons: it costs a 30-billion-parameter forward pass per request against a rule that costs nothing, and when I measured the whole stack end to end the model's plan won no requests the rule's did not. So the interesting question is whether something cheap and learned could sit between them. I put a learner there, and it scored below random units of the same shape.
The learner is growone, a small online model of mine that starts with no hidden units and adds them as it goes. Its selling point is an open key universe: it is built for features that arrive mid-stream, which is exactly what happens when each token of a request becomes a feature key and the corpus keeps reaching words it has not seen. It produces one number per candidate rather than a sequence, so the entity is never generated: every token of the request is scored, the highest wins, and a this request names nothing to track sentinel is carried as one more candidate rather than as a special case. No part-of-speech filter and no stop list, because a hand-written stop list is a rule, and importing a rule into a learner's candidate set is how a learner takes credit for somebody else's work.
The corpus is 75 project titles with frozen hand-written gold answers, drawn from three third-party lists: build-your-own-x, project-based-learning and Project-Ideas-And-Resources. The gold set holds 87 rows; twelve of them name a domain the tool already has a table for, so the reader is never consulted on those, and 75 reach it. The code is in a private research repo, so there is no link to that. Every figure comes from a results file or a command I ran, and I say which.
The comparable arm, and the one that looks respectable
There are two protocols here and only one of them may be set beside the baselines. The cross-corpus arm trains on 47 gold rows from a different corpus, freezes, and evaluates. That is cold on these 75 rows, exactly as the rule and the model are. The prequential arm predicts each row before learning from it, which is growone's native protocol and is not comparable, because by the time a late row is scored the net has already seen fifty earlier rows of the same corpus.
Every learner score below is a mean over seeds, which is why some of them are not integers.
| arm | of 75 | what it is |
|---|---|---|
| the model | 55 | qwen3-coder:30b, one run, transcribed from an earlier round rather than re-run here |
| the shipped rule | 47 | the hand-written incumbent, re-measured on today's tree |
| prequential | 33.5 | not comparable to the two above |
control: random units, same shape and schedule |
17.0 | |
floor: growth switched off, so a plain linear model on the same features |
17.0 | |
| the grown net | 15.5 | the comparable arm |
always answer post, the most common gold noun |
13 |
That 47 deserves a footnote, because this post is partly about stale numbers. The write-up that first scored this rule on these 75 rows published 29; the same scorer returns 47 today, because the rule has been worked on since. Every figure here is the re-measurement.
The 33.5 is the number that looks respectable and it is the one that should not be set against the rule. The comparable figure is 15.5.
And 15.5 is the interesting number for a reason that has nothing to do with the rule. The search is worth −1.5 rows against random units of the same shape and the same schedule. That is growone's own question, did the hidden units do anything, asked on a real task rather than on a generator.
Two seeds each, so be careful with the sign: the grown arm scored 17 and 14, the random-unit control 18 and 16, and the linear floor 17 twice. The round's pre-set band for a real difference on this arm was 2.9 rows, and 1.5 is inside it. So the finding is not that searching for units hurts; it is that it buys nothing measurable, against a control drawn to have exactly the same shape and schedule. The tool's own README reports the same thing on its home ground, where a linear floor beats the grown net at every size and randomly drawn units match the searched ones.
The only arm it clears is the constant. I should say what that constant nearly was. The first version answered with the last token of the request, on the reasoning that English noun phrases are head-final. It scored 1 of 75, because these titles end in the framework: …with React Native. An arm that cannot win makes the comparison against it pass by construction, which tests nothing.
The candidate set is not the excuse
The obvious objection is that picking the highest-scoring token of the request is a bad design: a row whose gold noun is not in the sentence cannot be won by it, however good the learner. That is true and it is worth 11 rows.
| of 75 | |
|---|---|
| ceiling: some candidate is accepted by the grader | 64 |
| unwinnable, because no token of the request is the answer | 11 |
| the same ceiling with the rule's own answer added as a candidate | 64, no change |
| rows the rule gets right that are not reachable as a raw token | 0 |
The eleven are not hard rows, they are rows where the noun is simply absent: Fitness App wants workout, A COVID-19 Tracker wants case, Tik Tok Clone - MERN STACK wants video. Supplying those is world knowledge, not extraction, and nothing in this round touches it.
Read against 64 instead of 75, every rate multiplies by the same 1.172 and no ranking moves. What the third line settles is the excuse: adding the rule's own answer to the candidate set raises the ceiling by zero rows, and not one row the rule gets right is unreachable as a raw token. The learner was picking from a set that contained the rule's answer every time it mattered, and picking something else.
Handing it the answer
So I handed it the answer directly. One extra feature key marks whichever candidate the shipped rule chose, and everything else stays. All arms in this section are prequential, so they are comparable to each other and all of them carry the same advantage over the cold baselines: 74 rows of practice the rule and the model never get.
| arm | mean of 75, five seeds | what it is |
|---|---|---|
| ceiling | 64 | oracle over the candidate set |
| the model | 55 | transcribed, not re-run |
rule_only |
48.0 | a net given the rule's pick and nothing else |
| the rule | 47 | the incumbent |
hybrid_floor |
41.0 | hybrid features, growth switched off |
hybrid |
40.4 | all the lexical features plus the rule's pick |
base |
32.6 | the lexical features alone |
The rule's read is worth a real and reliable +7.8 rows to the learner, from 32.6 to 40.4. Every seed gains between six and ten rows and not one seed loses a row, at p between 0.0020 and 0.0312. It is still not enough to reach the rule that supplied it, and that comparison does not resolve: the rule leads the discordant count (the rows where exactly one of the two is right) in all five seeds, and no seed gets below p = 0.096. At n = 75 the ranking is simply unresolved, which is weaker than calling it a loss and is what the numbers support.
Then there is the row I did not expect. rule_only, a net handed the rule's pick and nothing else, scores 48.0. hybrid, which is that same key plus thirteen families of lexical features, scores 40.4. Adding features to a working signal costs about eight rows, and rule_only leads the discordant count in every seed.
The mechanism is not mysterious. The token, suffix and neighbouring-word keys are dense and arrive on every candidate; the rule's pick is a single key arriving on one candidate; and the net has 75 rows in which to learn that the rare key outranks the common ones. It is the same shape as something one layer up in the same stack, where a planner filling a field suppressed the reading the layer below was already doing. Here it is a token vocabulary drowning the rule's.
rule_only beating the rule itself, 48 against 47, is not a result. It splits 2–1 discordant on every seed at p = 1.0. It is a pinned arm landing where a pinned arm should, and its value is precisely that it did: a harness that could not reproduce the rule from the rule's own answer would have invalidated every other number on this page.
Three arms separated by one key will report a margin whether or not the key does anything
So the key was blinded. The rule's accessor is rebound in memory to return nothing, and each arm is required to move the way its name demands.
| arm | with the key blinded | unblinded | what it must do |
|---|---|---|---|
base |
33.3 | 32.6 | must not move; it never had the key |
hybrid |
33.3 | 40.4 | must collapse onto base
|
rule_only |
21.0 | 48.0 | must collapse to no signal |
All three did. The load-bearing one is base: a check that every arm responds to shows only that the harness is live, and the arm that must not respond is what scopes it.
Growth earns nothing, and the totals were order-dependent
| mean | per seed | |
|---|---|---|
hybrid, growing |
40.4 | 41, 41, 39, 39, 42 |
hybrid_floor, growth switched off |
41.0 | 41, 41, 41, 41, 41 |
The floor is higher and it is identical on every seed. On this task, growth buys seed variance and nothing else.
A follow-up round found something else in those totals. The corpus arrives in source order, and the three lists are fully contiguous: 4 rows from one, then 38, then 33. So a prequential learner trains on one register, meaning one house style of writing project titles, and is then tested inside it. Shuffle the order and hold the seed count fixed, and the learner drops 3.7 rows:
| rule | learner | hybrid | |
|---|---|---|---|
| shipped order, 5 seeds | 47 | 32.6 | 40.4 |
| shipped order, 20 seeds | 47 | 32.9 | 40.6 |
| shuffled, 20 seeds | 47 | 29.2 | 37.9 |
The middle row is why the bottom one means anything. Seed count moves 0.3 rows; the order moves 3.7. Comparing five unshuffled seeds against twenty shuffled ones would have confounded the two.
Another label does help, and it never changes the ordering
The prequential protocol scores row i with a net that has seen rows 0 to i−1, so accuracy by position is the label curve. Over twenty shuffles, by thirds:
| third | rule | hybrid | learner |
|---|---|---|---|
| first | 0.648 | 0.384 | 0.302 |
| middle | 0.606 | 0.544 | 0.404 |
| last | 0.626 | 0.586 | 0.462 |
The rule is the control here, not a baseline. It does not learn, so its per-third rate is row difficulty alone, and it is flat to slightly down. Every row of the learner's rise is therefore learning rather than an easier tail, and it comes to +4.55 rows over that control, rising at every step.
That is a real learning signal, and it changes nothing about the ordering: rule, then hybrid, then learner, in every third, with the gap in the last third still about four rows in twenty-five. The stream runs out of labels before anything crosses.
A straight line through gained 0.16 accuracy over roughly fifty labels, still 0.164 short says another fifty would close it. That line is not evidence and I am not drawing it. Learning curves flatten, the reachable maximum for this design is 64 of 75 rather than 75, and nothing here measures the shape past label 75. What is established is narrower and still useful: the curve had not flattened when the labels ran out, so more labels would not help is refuted, and the label budget is the live lever.
What generalises
Three things, in the order I would want to be told them.
A learner's own control is the comparison that matters, and it is the one that gets left out. Against the incumbent rule this arm loses by 31.5 rows and the honest summary is "not close". Against random units of the same shape and schedule it is level, and that is the number saying the search bought nothing, which is a claim about the method rather than about the task.
Adding a feature to a working signal can subtract. I would have said more features is weakly monotone, and it cost eight rows here, for the ordinary reason that dense keys outvote a rare one when there are only 75 rows to sort it out in.
A pooled rate hides whether a learner is failing or merely starving. The same 32.6 comes out of a method that does not work and a method that has not finished; only the per-position curve separates them, and here it says the second one, while also saying the ordering never moved.
Top comments (0)