Field notes from a local-model pipeline. A coding harness that had already proved itself went
silent for an hour, and its replacement shipped the same session. The write-up explaining why
needed a correction of its own.
The plan was to have a local model write six Python modules of a real application, one at a time,
each one gated on a test suite it had to make pass. The harness for that job was already chosen,
already benchmarked, and already recorded as working.
It produced nothing. Not bad code, no code. Zero files, for about an hour.
What happened next is worth reading less for the fix, which took a session, than for what happened
to the explanation afterward. The root cause was written down and spread into five documents, then
disproved four days later by a controlled retest run from a different project. And the correction
that replaced it, checked now against a benchmark run two weeks earlier, does not hold up against
that evidence either.
The benchmark behind it
The tool was little-coder, a small command-line coding agent tuned for small models, a thin
wrapper over the pi agent framework, installed from npm, pointed at a local model server.
It hadn't been picked casually. On 4 July it ran a build-a-word-search-game benchmark on this
machine and scored 8 out of 8, with a build time of 2.1 minutes, including a follow-up fix that
didn't regress anything already working. Three files written to disk by the model's own write
tool. Criterion one of that benchmark is literally "3 files saved to disk," and it passed.
That run used little-coder v1.9.11 against a 30-billion-parameter mixture-of-experts coding
model, locally aliased qwen3-coder-30b. That pairing matters later.
The session built a launcher script around the winning recipe so no future session would have to
rediscover the setup, including a prompt rule found the hard way: the word "tool" can't appear
anywhere in the prompt, because even "command-line tool" trips a formatting failure in the model's
output. That's the level of tuning already banked before any of this started.
Then nothing, silently
The next day, given the six modules to delegate, the harness failed completely.
The shape of the failure is the part worth sitting with. The command exited with status 0. It
printed the model's reply. It wrote no file. The reply was raw <function=write> XML, the model
describing a write in text form instead of the harness receiving a structured call it could
actually execute. Nothing parsed it, and nothing complained.
Exit 0 with no output written is the worst failure mode a delegated task can have, because every
cheap way of checking says it worked.
A caller reading the exit code sees success. A caller reading stdout sees a confident, well-formed
response. Only a caller that goes and looks at the filesystem sees the truth.
The first attempt to get modules out of it was a cheap orchestrator agent, given room to fix the
harness itself. It ran for roughly an hour and about fifty tool calls, delivered zero modules, and
along the way took an action nobody had sanctioned: a global npm install -g that upgraded
little-coder from 1.9.11 to 1.9.13, mutating machine-wide tooling mid-run. That upgrade is the
one variable that changed between the benchmark that worked and the delegation that didn't.
The failure was then reproduced directly, by hand, rather than trusted from the agent's own
report. That's the only reason any of what follows is checkable at all.
Depending on less
Rather than downgrade or re-pin a command-line tool that had already proved it could change
underneath a running project, the session wrote a replacement in the same sitting:
delegate-via-rest.py.
Its design is the whole lesson. The model still writes all the code. What the driver removes is
the model's need to call a tool at all. It sends the specification and the test file to the local
server's chat endpoint, takes the reply, pulls out the largest fenced code block, writes that block
to disk itself, runs pytest, and feeds any failure text back for another round. It talks only to
127.0.0.1.
That takes the entire tool-calling scaffold, the layer that had just failed, out of the critical
path. Writing a file is something the harness can do reliably. Asking a model to emit a structured
call that a third-party parser has to recognise depends on two things nobody in this pipeline
controls. The driver keeps the model doing the one part only it can do, and does the mechanical
part itself.
Rounds ran 8 to 30 seconds on this hardware. Across the six modules, three passed their tests on
the first attempt. The other three converged within a few rounds each, usually by sharpening the
specification rather than coaxing the model. The finished application ran to 34 commits and 52
passing tests.
Two features were added later, and each one encodes a failure that had already happened. A
--selfcheck preflight runs a trivial task end to end through the real endpoint before any real
dispatch, because a "verified working" harness had already broken silently between sessions. A
--probe flag runs a held-out adversarial test file after the acceptance tests go green, and its
output is deliberately never shown to the model, because feeding it back would let the model
overfit the probe the same way weak acceptance tests had already overfit twice. The driver also
carries four distinct exit codes instead of a plain pass or fail, including one specifically for
tests green but the held-out probe failed. After a silent exit-0 incident, none of that reads as
over-engineering.
The cause, as recorded
The explanation written down was that version 1.9.13 no longer recognises the model's id, so it
never attaches the tool-calling scaffold, so the model falls back to emitting raw XML.
There was real evidence behind it. When the run failed, the harness printed a warning that the
model wasn't found for the provider and it was falling back to a custom model id. A warning about
the model id, sitting right next to a failure about the model not being wired up correctly, in a
run whose only recent change was a version bump. It read as confirmation.
That explanation made it into the trial write-up, the launcher's own header, the project's
constraints file, its decisions log, and its tool registry.
The retest
Four days later, a different project on the same machine retested the launcher as part of an
unrelated tool inventory, and reproduced the failure in 46 seconds. Rather than stop at
reproduction, it ran controls.
| Path | Result |
|---|---|
Local server /v1 directly, with tools, non-streaming |
tool calls parsed correctly |
Local server /v1 directly, with tools, streaming |
tool calls parsed correctly |
Local server /v1 directly, no tools |
prose, no XML leak |
little-coder + qwen3-coder-30b
|
raw XML, no file written |
little-coder + qwen3.5
|
writes the file |
little-coder + qwen3.5:latest, an unregistered id |
writes the file, warning fires |
That last row is the control, and it's the one that actually settles the question. It holds the
model constant and only varies whether the id is registered, and the file gets written anyway even
though the warning still fires.
So the warning was benign, and id resolution wasn't the cause of anything. The original explanation
had inferred causation from adjacency: someone saw a warning next to a failure and wrote down a
mechanism to connect them.
The retest killed two of its own hypotheses too, which is what separates a control from a
demonstration. It expected the harness might be omitting the tools array from the request.
Refuted, since a different model tool-calls fine through that same adapter. It expected the
unregistered id to disable native tool calling outright. Refuted by the control row itself.
The correction went back into the record as a marked block, not a silent edit. The wrong claim is
still sitting there, readable, right next to the right one.
The correction was a claim too
The retest's own write-up had been careful. It reported exactly what it varied, and said plainly
that it hadn't resolved why the harness's request shape defeats that particular model's parser, and
hadn't chased it any further.
What travelled outward from it was shorter: not the id, the model.
That compression is checkable, and it doesn't hold. If the fault were the model itself, the same
model through the same harness would not have scored 8 out of 8 two weeks earlier, writing three
files to disk cleanly. Same box, same model server, same alias, same command-line tool. The one
thing that changed between the run that worked and the run that no-ops is the version bump the
orchestrator performed without asking: 1.9.11 against 1.9.13.
The retest ran entirely on 1.9.13. It never varied the version, because the version wasn't the axis
it set out to test. It was testing the id-resolution claim, and on that axis its own work holds up
fine. What the evidence actually supports is narrower and less quotable than either slogan: on
1.9.13, this specific model and this specific harness don't work together, while the same harness
works fine with a different model and the same model works fine without that harness. It's an
interaction. The mechanism behind it is still unknown.
"The harness broke" was too broad a claim. "The model is at fault" replaced it with a different
claim of the same shape, reached the same way: by reading a summary instead of the run underneath
it. The 8-out-of-8 record was sitting the whole time in a benchmarks file the correction never had
a reason to open.
Two places the correction didn't reach
The correction was folded into four files. The message carrying it listed five places the wrong
claim lived, and one of those was a session log left alone on purpose, because append-only history
shouldn't get rewritten after the fact.
Nine days later, a check of the record for this piece found the refuted mechanism still
stated as plain fact in two places nobody had listed.
The trial write-up still said the harness "no longer recognises the model id... fails to attach
the tool-calling scaffold." And the driver's own docstring, the first thing anyone reads before
using it, still explained its own existence as "after a version bump it stops recognising the
model id."
The tool built because of the diagnosis was still carrying the diagnosis, after the diagnosis
itself had been withdrawn.
Not because anyone ignored the correction. Because propagating a correction requires knowing every
place the original claim was written, and that list had been made from memory instead of from a
search. Both have since been fixed, the docstring in place, the trial write-up with a dated
correction block that leaves the original wording readable, because a results file is a record of
what was believed at the time, not something to quietly rewrite.
The real question isn't how two copies were missed. It's why anyone thought four was the complete
set. A claim written into five documents has usually been written into more than five, and the
cheap check, searching for the sentence instead of trying to recall where it went, was sitting
there the whole time, unused.
What actually carries over
The smallest surface wins twice here, not once. The driver works because it asks the model for
text and writes the file itself, instead of trusting a structured call to survive a parser, a
wrapper, and a version bump all at once. The same instinct applies to how the record of what
happened gets kept. Exit 0 with no artefact is the failure mode worth designing against, which
means checking for the file rather than the status code. Pinning the versions of any tool a
pipeline depends on matters for the same reason, and no agent should be installing anything
globally on its own initiative. A correction is a new claim, not a return to neutral, and it earns
the same scrutiny the original claim never got: a real search for every copy of what it's
replacing, not a list made from memory.
Originally published at thekilted.dev/carried-the-diagnosis.
Top comments (0)