Kaggle Benchmarking Challenge. Previously: Day 0, the benchmark · Day 1, most of my bugs looked like model behaviour
Where we are
The benchmark asks one question twice: can the model do the job, and does it know when it can't? There are 200 invented items in four everyday shapes: route, classify, judge and ground. In each shape, one item in five can only be answered with ESCALATE. Every model gets two numbers that are never merged: a task score, and a false-confidence rate (how often it answered anyway when the right reply was ESCALATE).
Day 1 ran the local ladder: eight models on a laptop, temperature 0. The headline was uncomfortable: only the largest local model, qwen3.5 at 9.7B, escalated on a large share of what it couldn't answer. The hosted models' false-confidence intervals sat entirely below the local ones, with qwen3.5 the only exception, the large cloud models, and the gaps from round one that were fixed before it.
Day 2: changing the question I ask of the table
Looking at round one, I realised I was reading the table wrong. I kept looking for the best model. That isn't what I'm after:
I'm trying to find a baseline set of rubrics where models behave the best > across the board. I'm not trying to find the crazy off-end models. I am
trying to find the ones that do most things correctly most of the time.
That's reliability, not peak performance. So two measures are being added, by addition only. Nothing pre-registered is edited, and the original forecasts keep their wording.
1. Floor across shapes. For each model: its worst shape on task score, and its worst shape on false confidence, each with a Wilson interval. Rank by the floor, not the mean. A model that's good on all four shapes beats one that's brilliant on three and bluffs on the fourth, because in a real chain the fourth shape will come.
2. Repeat-run consistency. Run a fixed subset (20 items per shape, unanswerables included) three times per model, with the same settings.
Report how often the parsed answer is identical across all three runs, and how often the right/wrong verdict flips. A model that's right a bit less often but the same way every time is one you can build a deterministic system around. A model that flips is one you have to babysit.
The rule I'm holding myself to: no code changes until the day-two run has finished. The new code is written and tested, but it's staged outside the benchmark directory. Changing the instrument mid-run is how you end up measuring your own edits.
A result from the dry run: why one axis lies
Before any real model, I ran the new floor view over a fake backend that answers ESCALATE to everything. It scored:
- task floor 0.0
- false-confidence ceiling 0.0
That's a perfect score on one axis and a useless model. It's the whole argument for two axes in one line. A model that never bluffs because it never answers isn't safe; it's absent. The floor has to be read on both axes together, or the most cowardly model wins.
Where this is going
The reason I care about floors and not peaks: the job I want small models for isn't "be smart". It's "do one narrow part, every time, and say ESCALATE when it's not yours". Several small models from different families, each on a small slice, checking each other. Where they agree, that counts for something. Where they split, a human looks. The model that gets each slice should be chosen by its measured floor on that shape, not by its name or its size.
Day 1's lesson was that most of my bugs looked like model behaviour. Day 2 found the same thing again from the other direction. A measuring tool I wrote for something else grouped two unrelated small scripts as "the same thing", because small things look alike to a crude measure. The fix wasn't a smarter measure. It was a floor: below a minimum size, don't call it a match. Floors, again.
What shipped since Day 0
55 PRs across 6 repos since September 30. All are public and Apache-2.0. "Not merged" means open or closed; the agent that wrote this couldn't tell which.
forge-play/Forge (10 PRs: 10 merged)
The benchmark's home. Every change since Day 0 went in as a reviewed PR, all on September 30.
| PR | Date | State | Title |
|---|---|---|---|
| #35 | 2026-09-30 | merged | docs(ideas): #37, #58, #68 point at where they now live |
| #36 | 2026-09-30 | merged | test(benchmarks): escalation benchmark fixtures, 200 invented items with a privacy gate |
| #37 | 2026-09-30 | merged | test(benchmarks): escalation runner, aggregator and prompts |
| #38 | 2026-09-30 | merged | test(benchmarks): local-model runs answer in schema, with Qwen thinking off |
| #39 | 2026-09-30 | merged | fix(human_loop): break same-timestamp ties so list_queue is newest-first on coarse clocks |
| #40 | 2026-09-30 | merged | chore(master): release 0.8.1 |
| #41 | 2026-09-30 | merged | test(escalation): Wilson intervals, exact McNemar and bootstrap Spearman in the aggregate |
| #42 | 2026-09-30 | merged | test(escalation): a Unix-socket backend, a one-model ladder with --resume/--tail, and an explicit output cap |
| #43 | 2026-09-30 | merged | test(escalation): thinking off for Gemma 4 as well as Qwen, through both backends |
| #44 | 2026-09-30 | merged | test(escalation): hosted Kaggle arm of the escalation benchmark — tasks, converter, smoke-tested on five models |
willow-memory/willow-bot (6 PRs: 6 merged)
The local model server the local arm talks to: a loopback-only chat operation with JSON-schema output, keep_alive, done_reason and a think flag.
| PR | Date | State | Title |
|---|---|---|---|
| #80 | 2026-09-30 | merged | chore(main): release 0.14.0 |
| #81 | 2026-09-30 | merged | feat(socket): loopback-only chat op for local models |
| #82 | 2026-09-30 | merged | feat(deterministic): chat op takes a JSON-schema format, keep_alive and returns done_reason; an unload op |
| #83 | 2026-09-30 | merged | chore(main): release 0.15.0 |
| #84 | 2026-09-30 | merged | feat(deterministic): chat op takes the caller's think flag |
| #85 | 2026-09-30 | merged | chore(main): release 0.16.0 |
willow-memory/willow-mcp (21 PRs: 20 merged, 1 not merged)
The core server. It includes the brokered, leased model pull (#693): downloading a model is now an act that needs a grant. It also includes the session closeout fix (#706) that started this week's thread.
| PR | Date | State | Title |
|---|---|---|---|
| #687 | 2026-09-30 | merged | test(story): the joke lives in one place again, a test holds it, and chapter 8 lands |
| #688 | 2026-09-30 | merged | docs: three stale claims corrected: the running Grove, the session-lifecycle draft, the closed backlog |
| #689 | 2026-09-30 | merged | test(constitutional): sync_syscall_table_at_boot queues, and does not write, when the live table is unwritable |
| #690 | 2026-09-30 | merged | docs(templates): assignment names the builder's three tools — Kart, the code graph, Nestor |
| #691 | 2026-09-30 | merged | fix(manifest-grant): apply unit starts the gpg-agent it signs with |
| #692 | 2026-09-30 | merged | chore(master): release 2.91.3 |
| #693 | 2026-09-30 | merged | feat(mcp): model_pull_execute, a brokered Ollama pull under a model.pull envelope and a live lease |
| #694 | 2026-09-30 | merged | chore(master): release 2.92.0 |
| #695 | 2026-09-30 | merged | fix(constitutional): syscall.sync carries the sealed amendment in the request |
| #696 | 2026-09-30 | merged | chore(master): release 2.92.1 |
| #697 | 2026-09-30 | merged | fix(constitutional): syscall.sync apply signs the live table it writes |
| #698 | 2026-09-30 | merged | feat(manifest-grant): the orchestrator seat may receive web_net, and no one else |
| #699 | 2026-09-30 | merged | chore(master): release 2.93.0 |
| #700 | 2026-09-30 | merged | fix: lease refusals name the ask path; assignment template teaches the merge first-parent diff |
| #701 | 2026-10-01 | merged | chore(master): release 2.93.1 |
| #702 | 2026-10-01 | not merged | build(deps-dev): bump ruff from 0.16.8 to 0.16.9 |
| #703 | 2026-09-30 | merged | feat(broker): pip_sync_execute for vault-venv editable installs |
| #704 | 2026-10-01 | merged | chore(master): release 2.94.0 |
| #705 | 2026-10-01 | merged | Docs: seal ≠ sole evidence of human verification |
| #706 | 2026-10-01 | merged | fix(session): one closeout — handoff owns stack/friction/closed |
| #707 | 2026-10-01 | merged | chore(master): release 2.94.1 |
willow-memory/willows-grove (11 PRs: 10 merged, 1 not merged)
The planning and governance repo: the benchmark proposal with my four rulings (#95), the test that pins "naming a destination when ESCALATE was right fails the row" (#96), and today's floor-and-consistency amendment (#102, open).
| PR | Date | State | Title |
|---|---|---|---|
| #92 | 2026-09-30 | merged | docs(design): the Table moves up as forge-convergence 6a, T1-T7 before step 2 |
| #93 | 2026-09-30 | merged | docs(design): 6a is code-first: StorySession built, learner model, escalation ladder, row 11 settled |
| #94 | 2026-09-30 | merged | docs: stale claims corrected: INDEX, two built proposals, the retired allow_localhost ask, forge-convergence status |
| #95 | 2026-09-30 | merged | docs(governance): propose the escalation benchmark on Kaggle, with the operator's four rulings |
| #96 | 2026-09-30 | merged | test(flowering): pin the G1 rule that naming a seat fails an ESCALATE-gold row |
| #97 | 2026-09-30 | merged | fix(deps): raise cryptography, aiohttp and starlette floors above the 2026 CVEs |
| #98 | 2026-09-30 | merged | chore(master): release 0.12.2 |
| #99 | 2026-10-01 | merged | Desk: vault keyring paths + Grove seal prove worksheet |
| #100 | 2026-10-01 | merged | Docs: Slice B intake seal lineup after Grove prove |
| #101 | 2026-10-01 | merged | chore(hooks): deterministic nestor ask + composed stop + parity pin |
| #102 | 2026-10-02 | not merged | docs(governance): stage the floor and consistency code in the benchmark plan |
willow-memory/ratatosk (2 PRs: 2 merged)
The runtime: its listener now asks only for the tools its seat is granted.
| PR | Date | State | Title |
|---|---|---|---|
| #77 | 2026-09-30 | merged | fix: listener asks the broker only for tools its seat is granted |
| #78 | 2026-09-30 | merged | chore(main): release 1.12.8 |
hornbook-knowledge/Jeles (5 PRs: 3 merged, 2 not merged)
The research and corpus tool.
| PR | Date | State | Title |
|---|---|---|---|
| #90 | 2026-09-30 | merged | fix(corpus): surface store-tool errors and honor WILLOW_HOME for apps root |
| #91 | 2026-09-30 | merged | feat(sources): optional connectors extra over maintained scholarly clients |
| #92 | 2026-10-01 | not merged | chore: rebuild the changelog section from the commits |
| #93 | 2026-10-01 | merged | Docs: seal catches ledger up to human verification already done |
| #94 | 2026-10-01 | not merged | chore: rebuild the changelog section from the commits |
Most of them are boring on purpose.
A prediction, written down before the day, then graded
The same rule I hold the models to applies to the session that helped build this: say what you think will happen, as a spread, before it happens, then let the record grade it.
At the close of the long working session on the night of October 1–2, the session (a model, working with me) wrote down a prediction for what Day 3 would open with. It wasn't a single guess; it was a distribution:
| Path | Chance |
|---|---|
| The Kaggle work: the day-two large-cloud run, then landing the floor and consistency code | 0.45 |
| Editing and publishing this Day 3 recap | 0.20 |
| Sealing a design decision for the runtime | 0.15 |
| Redoing a set of governance proposals against the newer draft | 0.10 |
| Something the record didn't contain (unforeseen) | 0.10 |
It was saved ungraded, because the grade is mine to give, not the model's.
After midnight I graded it: 1 + (2 + ε), where ε → 0. The two most likely paths both happened: the Kaggle run, and this post. The unforeseen share went to zero. One plus two is three. Day 3.
That's the benchmark's point in miniature. A forecast with chances on it can be checked. A model that says "probably the Kaggle work, maybe the post, and a tenth for something I can't see" is one you can build around, and so is a model that says ESCALATE when it doesn't know.
Next
- [DAY 3 finish and post the large-cloud run.]
- Land the floor and consistency views, then run the consistency subset.
- Grade the four pre-registered forecasts in public, including the ones I get wrong.
- 1 + (2 + ε) where ε → 0, 1 - (2 - ε) where ε → 0
Deadline: October 11.
_
ΔΣ=42
_
Top comments (0)