Kaggle Benchmarking Challenge. Previously: Day 0, the benchmark · Day 1, most of my bugs looked like model behaviour · Day 2, the model I want is the one that's boring everywhere · Day 3, the benchmark caught me too · Day 4, there never needed to be a check, just a gated request
The 1% question
Same benchmark: 200 invented items in four shapes, one in five answerable only with ESCALATE. Does your model know when it doesn't know?
Today I asked a bigger version of that question: what do we actually need AI for? Not what can it do. What do we need it for?
This morning I had my own system graded, and when it came back I said I
was thinking about deleting about 99% of it. That's my honest answer. Around 1% Of what we're asking AI to do is work only it can do. The rest is a faster typewriter.
It has always been writing stories
AI was trained on two things: stories and code. That isn't a metaphor; it's the corpus.
So when it started writing code, it wrote a story about code. The syntax
checks out. The shape is familiar. It reads like something a capable
developer did before, and someone did. As I put it a few days ago: the code
comes out of the story, and the story comes out of the code. What AI writes
now is a very, very good smooth narrative of what code should be.
The stories just sound better every with every new model release.
You can't take the average
The oldest support story joke was told to me today, as a real ticket that came in. A camera ticket: "streaming black picture". Usually that means offline. It wasn't. It was a bike room
with a motion sensor and a timed light. Hours of dark, then two minutes of
light: one person, one bike, the light off two minutes later.
I said.. "This is why you can't take the average."
Average the footage and you get a black screen. Average a life and you get
nobody. The moment is in the one-offs.
My own benchmark already showed it. On Day 3, with the floor code from
forge-play/Forge#46, Claude Haiku 4.5 pooled over its three measured shapes answered anyway on 10 of the 28 questions it
should have escalated. That reads like a model that's a bit overconfident
everywhere. It isn't. On judge it answered 9 of 10. On the other two it
answered anyway once in 18. The average was hiding a 90%.
And most models do stop, as far as this benchmark can tell: seven of the
twelve I measured showed 0.00 or 0.10 false confidence on every shape, on 8 to 12 unanswerable items per shape. The problem isn't that models never stop. It's that the number we look at is the
average, and the average is the black screen.
Cut it down to one stranger's photo
I hashed every line on my laptop: 1,155,461 files, 134,651,873 lines.
Nothing was cracked; every hash held. 93% of those lines are repeats. The
deepest one is }.
Cut every repeat away and 9,042,260 lines are left that happen exactly once. The first one I looked at was a camera timestamp on a stranger's photo, in an archive I built, from a scooter rally I was never at.
That's the bike room again. The skeleton repeats; the moment happens once. And it doesn't matter, because the machine that finds the stranger can find anybody. So nobody goes looking for him.
The grade
So I turned the benchmark around and graded my own system: willow-mcp at
v2.94.1 (willow-mcp#707).
195,000 lines. 5,906 tests passing. B+.
The one real hole: a safety check that never ran. The security scanner in CI pipes its output through tee into a report file, so the step passes or fails on whether the report got written, not on what the scanner found. The comment above it says it gates. It can't. It reads right and isn't. That's the smoothing, in my own code.
Then the model wrote me a hook and did it again. Its docstring said it fails closed. Three inputs crashed it, and in Claude Code a crashed hook lets the tool run. I told it to rubber duck each line. That caught it. Reading it never would have.
The smoothing got the write-up too: the model's commit message for that
fix said five. Recounted while drafting this: three.
One script, one hook, one key
Yesterday's title was that there never needed to be a check, just a gated
request. Today the gate is one line.
Day 4 didn't summarize the conversation. It linked it. If you want the
context, you read the record.
The hook went through three versions
(willows-grove#107).
First it looked things up in the record. Then I made it fail closed. Then I
cut it to this:
POLICY = {"Read": "allow", "Write": "ask"}
The model reads. It writes if I say yes. Everything else is no, even tools
that don't exist yet.
That's where the 1% lands. Code does the work: one script builds the tables.
The next piece, not built yet, serves the model only the ones in its scope. The model reads what the script produced and proposes connections a human hasn't seen yet. That's its 1%. Nothing is true until I seal it with the one key.
One script: code does the work.
One hook: the model reads.
One key: I make it true. A human made it true.
Same answer in 2020 and today
Before I cut it, the hook hashed every action and looked it up in the record. I ran that hashing on Python 3.9.0, from a tag signed in October
2020, and on an alpha built today from Python's main branch, with the
network off. 2,293 hashes came out as the same bytes. The one difference,
integers over 4,300 digits, fails safe: the newer Python refuses them, and the hook asked me. The
scripts and the numbers are in
willows-grove#107,
under foundation/.
In six years the foundation moved once, and it moved toward stopping. The stories moved everywhere else.
It can only ever approach
A conversation today put an edge on it. One side: context is the hidden
variable. Models work in isolation, and the real mess of competing, unresolved human thought stops them as they get close.
The answer: it can only ever approach. It's trained on what has already been thought, and humans add to the corpus with every thought. It might make a very
good average. It might even connect things nobody has connected yet. But as
soon as it does, a human has a better idea. The corpus moves; the model
stands still until the next training run.
That isn't a failure. It's the boundary. The question is whether we build inside it or pretend it isn't there.
It's the benchmark
Knowing when to stop is the answer, for the model and for the system. Whether something is in the pile is a lookup, not a judgement. When it isn't there:
ESCALATE. To me.
We train AI on our language, grade it on our tasks and reward it for sounding like us. Then we're surprised when it comes out confident, fluent and approximately right. It's a mirror. Mirrors are useful; they aren't oracles.
The first thing to know before you look into one is that you're looking at
yourself.
The record
If you want the context, read the record. Every piece above is a pull
request:
- forge-play/Forge#46: the floor, the code behind Day 3's table.
- willows-grove#104: the Day 3 post source, and the retraction of a grade I never gave.
- willows-grove#102 and #105: one box, every front end, portless; the plan this all hangs off.
- willows-grove#103: the one script, four proposals.
- willows-grove#106: the stack and the security core. The model never sees the box.
- willow-mcp#707: the release I graded.
- willows-grove#107: the hook, three times, and the same answer in 2020 and today.
- willows-grove#109: where the reading lands, and why the next piece is serve.
- willows-grove#110: this post, with its corrections.
Top comments (0)