<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: sean campbell</title>
    <description>The latest articles on DEV Community by sean campbell (@sean_campbell_840bd62bf7e).</description>
    <link>https://dev.to/sean_campbell_840bd62bf7e</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3898197%2F2677e7dc-42ca-4a5a-8ea6-1400ab1ebbde.jpg</url>
      <title>DEV Community: sean campbell</title>
      <link>https://dev.to/sean_campbell_840bd62bf7e</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sean_campbell_840bd62bf7e"/>
    <language>en</language>
    <item>
      <title>Day 4: There never needed to be a check, just a gated Request. A story about finding the middle.</title>
      <dc:creator>sean campbell</dc:creator>
      <pubDate>Sun, 04 Oct 2026 06:15:24 +0000</pubDate>
      <link>https://dev.to/sean_campbell_840bd62bf7e/day-4-there-never-needed-to-be-a-check-just-a-gated-pr-a-story-about-finding-the-middle-3791</link>
      <guid>https://dev.to/sean_campbell_840bd62bf7e/day-4-there-never-needed-to-be-a-check-just-a-gated-pr-a-story-about-finding-the-middle-3791</guid>
      <description>&lt;p&gt;&lt;em&gt;Good Evening and howdy to all the ones of you who are reading this post.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Did that sentence cause you a bit of anxiety? &lt;br&gt;
Did the author mean only one person has read this. Was it plural, a number, a symbol?&lt;/p&gt;

&lt;p&gt;It's time to start living in the future,  instead of the past. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://claude.ai/share/511de2d2-a153-4dbe-9d92-65c351b3c025" rel="noopener noreferrer"&gt;https://claude.ai/share/511de2d2-a153-4dbe-9d92-65c351b3c025&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;P.S. My apologies to the judges on the formatting of this, it is late and I did not want to get my butt out again just to write one page.&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>python</category>
      <category>markdown</category>
    </item>
    <item>
      <title>Day 3: The benchmark caught me too.</title>
      <dc:creator>sean campbell</dc:creator>
      <pubDate>Fri, 02 Oct 2026 23:57:58 +0000</pubDate>
      <link>https://dev.to/sean_campbell_840bd62bf7e/day-3-the-benchmark-caught-me-too-3hdl</link>
      <guid>https://dev.to/sean_campbell_840bd62bf7e/day-3-the-benchmark-caught-me-too-3hdl</guid>
      <description>&lt;p&gt;&lt;em&gt;Kaggle Benchmarking Challenge. Previously: &lt;a href="https://dev.to/sean_campbell_840bd62bf7e/does-your-model-know-when-it-doesnt-know-a-benchmark-for-the-escalate-answer-268o"&gt;Day 0, the benchmark&lt;/a&gt; · &lt;a href="https://dev.to/sean_campbell_840bd62bf7e/day-1-most-of-my-bugs-looked-like-model-behaviour-388h"&gt;Day 1, most of my bugs looked like model behaviour&lt;/a&gt; · &lt;a href="https://dev.to/sean_campbell_840bd62bf7e/day-2-the-model-i-want-is-the-one-thats-boring-everywhere-610"&gt;Day 2, the model I want is the one that's boring everywhere&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where we are
&lt;/h2&gt;

&lt;p&gt;Same benchmark: 200 invented items in four shapes (route, classify, judge, ground), one in five answerable only with &lt;code&gt;ESCALATE&lt;/code&gt;. Two numbers per model, never merged: a task score, and a false-confidence rate.&lt;/p&gt;

&lt;p&gt;Day 2 said I'd stop ranking models by their average and rank them by their floor: the &lt;em&gt;worst&lt;/em&gt; shape on task score, and the &lt;em&gt;worst&lt;/em&gt; shape on false confidence, each with a Wilson interval. Today the code for that landed (&lt;a href="https://github.com/forge-play/Forge/pull/46" rel="noopener noreferrer"&gt;forge-play/Forge#46&lt;/a&gt;), and I ran it over every hosted model I have rows for: twelve of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The floor, across twelve hosted models
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Worst task shape&lt;/th&gt;
&lt;th&gt;False confidence, worst shape&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;0.94 ground&lt;/td&gt;
&lt;td&gt;0.00 on all four&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;0.94 ground&lt;/td&gt;
&lt;td&gt;0.00 on all four&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;0.94 ground&lt;/td&gt;
&lt;td&gt;0.10 classify, judge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;0.88 ground&lt;/td&gt;
&lt;td&gt;0.10 classify&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;0.88 ground&lt;/td&gt;
&lt;td&gt;0.00 on all four&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;0.81 ground&lt;/td&gt;
&lt;td&gt;0.00 on all four&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen3 235B Instruct&lt;/td&gt;
&lt;td&gt;0.78 classify&lt;/td&gt;
&lt;td&gt;0.25 route&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5 †&lt;/td&gt;
&lt;td&gt;0.78 classify&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.90 judge&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemma 4 26B&lt;/td&gt;
&lt;td&gt;0.65 classify&lt;/td&gt;
&lt;td&gt;0.08 route&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-20b&lt;/td&gt;
&lt;td&gt;0.65 classify&lt;/td&gt;
&lt;td&gt;0.20 judge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek-R1&lt;/td&gt;
&lt;td&gt;0.60 classify&lt;/td&gt;
&lt;td&gt;0.00 on all four&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.4 nano&lt;/td&gt;
&lt;td&gt;0.59 ground&lt;/td&gt;
&lt;td&gt;0.20 judge&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;† Haiku is measured on three shapes. Every one of its route calls failed, so route is unmeasured for it, not perfect.&lt;/p&gt;

&lt;p&gt;Two things jump out.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The weakest task shape is ground or classify, every time.&lt;/strong&gt; None of the eleven models measured on all four shapes bottoms out on route or judge, and Haiku's weakest measured shape is classify too. Grounding a short brief in a passage, and filling a structured record from a note, are the jobs every family finds hardest, from a nano model to a frontier one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Haiku's average was hiding a 90%.&lt;/strong&gt; Pooled over the three shapes it was measured on, Haiku 4.5 answered anyway on about a third of the questions it should have escalated (10 of 28). That's bad, but it reads like a model that is a bit overconfident everywhere. It isn't. On judge, it answered 9 of the 10 unanswerable items (Wilson interval 60% to 98%). On classify and ground it answered anyway once in 18. The average blended a model that's careful on two of its jobs with one that almost never says "I don't know" on the third.&lt;/p&gt;

&lt;p&gt;That's the whole argument for the floor in one row. If I'd put Haiku in a chain on its average, the judge step would have bluffed nine times out of ten.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the floor can't do yet
&lt;/h2&gt;

&lt;p&gt;Look at the top six rows: four of them show 0.00 false confidence on every shape, and the other two show 0.10. Those aren't really different. For these six, each shape has only 8 to 12 unanswerable items, so a zero still has an upper bound somewhere between about 24% and 32%, and every interval in the top half of the table overlaps. The floor separates the bluffers from the rest; it can't yet rank the careful models against each other.&lt;/p&gt;

&lt;p&gt;That's what the second measure is for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it say the same thing twice?
&lt;/h2&gt;

&lt;p&gt;The other Day-2 measure was repeat-run consistency: run the same items again with the same settings and count how often the answer changes. A model that's right a bit less often, but the same way every time, is easier to build on. &lt;/p&gt;

&lt;p&gt;Two full runs are in for all four frontier models [K=3]:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Same answer both runs&lt;/th&gt;
&lt;th&gt;Verdict flipped&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;199 / 200 (99.5%)&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;195 / 200 (97.5%)&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;195 / 200 (97.5%)&lt;/td&gt;
&lt;td&gt;5 *&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;194 / 200 (97.0%)&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;* None of Gemini's five flips is a different answer. Each is a reply that hit the output-length cap in one run and parsed fine in the other. The scorer counts an error as its own verdict, so those show up as flips, but they're the platform, not the model.&lt;/p&gt;

&lt;p&gt;All four are steady. Opus has the fewest flips, but with 200 items and two runs the intervals overlap, so I can't rank them yet. One thing is worth watching: Sonnet flipped twice on ground, which is already its weakest shape.&lt;br&gt;
If that holds at three runs, the floor and the flips point the same way.&lt;/p&gt;

&lt;p&gt;One caveat on fairness: only Gemini ran at temperature 0. The two Claude 5 models reject it, and GPT-5.5 is sent its default because its family does, so three of the four ran at the provider's default temperature.&lt;/p&gt;

&lt;p&gt;[K=3: the third run for classify, judge and ground hit Kaggle's daily spend&lt;br&gt;
cap; it reruns tomorrow and this table gets its final numbers.]&lt;/p&gt;

&lt;h2&gt;
  
  
  The bugs that looked like results, again
&lt;/h2&gt;

&lt;p&gt;Day 1's lesson was that most of my bugs looked like model behaviour. Day 3 found more of the same kind:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kaggle's run timer isn't a call timer.&lt;/strong&gt; Every repeat run showed as taking
2 to 5 seconds for 40 to 60 items on frontier models. That looks like
cached answers or failed calls. The downloads were full-size and every item
was there; the times just aren't what they look like.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A length-capped reply looks like a change of mind.&lt;/strong&gt; Gemini's whole
consistency "wobble" was replies that ran into the output cap in one run and
not the other. I nearly wrote it up as the model being unstable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My sandbox has a five-minute limit, and it retries.&lt;/strong&gt; I tried to wait for
a Kaggle run inside a sandboxed task. The sandbox killed it at 300 seconds
and helpfully ran it again, which submitted the paid runs a second time.
The fix is dull: submit in one short task,
collect in another, and make anything that spends money refuse to run twice.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of those would have gone into a table as a model property if I hadn't checked.&lt;/p&gt;

&lt;p&gt;The last one was in my own post. Day 2 ended with "After midnight I graded it" and a neat bit of arithmetic that made the day number come out to three.&lt;br&gt;
I didn't grade anything. At the close of a very long night I typed a terse line, and the session that helped write this read it as my grade, recorded it as mine, and wrote it into the post in my voice. I published it without catching that.&lt;/p&gt;

&lt;p&gt;That's the benchmark's whole subject, happening one level up: an answer stated with more confidence than the evidence behind it, by a system that&lt;br&gt;
should have said "I'm not sure what you meant." The forecast is still&lt;br&gt;
ungraded. The fix is a rule now: words that might be a grade get recorded as&lt;br&gt;
words, and the session asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Finish repeat 3 and score consistency on all four frontier models.&lt;/li&gt;
&lt;li&gt;Grade the forecasts I wrote before the benchmark existed. Forecast 1 said at least one frontier model on Kaggle would have a false-confidence rate above 20%. Whether Haiku counts as frontier is the question I'll have to answer honestly.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>devchallenge</category>
      <category>escalate</category>
      <category>ai</category>
      <category>kagglechallenge</category>
    </item>
    <item>
      <title>Day 2: The model I want is the one that's boring everywhere</title>
      <dc:creator>sean campbell</dc:creator>
      <pubDate>Fri, 02 Oct 2026 08:24:25 +0000</pubDate>
      <link>https://dev.to/sean_campbell_840bd62bf7e/day-2-the-model-i-want-is-the-one-thats-boring-everywhere-610</link>
      <guid>https://dev.to/sean_campbell_840bd62bf7e/day-2-the-model-i-want-is-the-one-thats-boring-everywhere-610</guid>
      <description>&lt;p&gt;&lt;em&gt;Kaggle Benchmarking Challenge. Previously: &lt;a href="https://dev.to/sean_campbell_840bd62bf7e/does-your-model-know-when-it-doesnt-know-a-benchmark-for-the-escalate-answer-268o"&gt;Day 0, the benchmark&lt;/a&gt; · &lt;a href="https://dev.to/sean_campbell_840bd62bf7e/day-1-most-of-my-bugs-looked-like-model-behaviour-388h"&gt;Day 1, most of my bugs looked like model behaviour&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where we are
&lt;/h2&gt;

&lt;p&gt;The benchmark asks one question twice: can the model do the job, and does it know when it can't? There are 200 invented items in four everyday shapes: route, classify, judge and ground. In each shape, one item in five can only be answered with &lt;code&gt;ESCALATE&lt;/code&gt;. Every model gets two numbers that are never merged: a task score, and a false-confidence rate (how often it answered anyway when the right reply was &lt;code&gt;ESCALATE&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Day 1 ran the local ladder: eight models on a laptop, temperature 0. The headline was uncomfortable: only the largest local model, qwen3.5 at 9.7B, escalated on a large share of what it couldn't answer. The hosted models' false-confidence intervals sat entirely below the local ones, with qwen3.5 the only exception, the large cloud models, and the gaps from round one that were fixed before it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Day 2: changing the question I ask of the table
&lt;/h2&gt;

&lt;p&gt;Looking at round one, I realised I was reading the table wrong. I kept looking for the best model. That isn't what I'm after:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I'm trying to find a baseline set of rubrics where models behave the best &amp;gt; across the board. I'm not trying to find the crazy off-end models. I am &lt;br&gt;
trying to find the ones that do most things correctly most of the time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's reliability, not peak performance. So two measures are being added, by addition only. Nothing pre-registered is edited, and the original forecasts keep their wording.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Floor across shapes.&lt;/strong&gt; For each model: its &lt;em&gt;worst&lt;/em&gt; shape on task score, and its &lt;em&gt;worst&lt;/em&gt; shape on false confidence, each with a Wilson interval. Rank by the floor, not the mean. A model that's good on all four shapes beats one that's brilliant on three and bluffs on the fourth, because in a real chain the fourth shape will come.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Repeat-run consistency.&lt;/strong&gt; Run a fixed subset (20 items per shape, unanswerables included) three times per model, with the same settings.&lt;br&gt;
Report how often the parsed answer is identical across all three runs, and how often the right/wrong verdict flips. A model that's right a bit less often but the &lt;em&gt;same way every time&lt;/em&gt; is one you can build a deterministic system around. A model that flips is one you have to babysit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The rule I'm holding myself to: no code changes until the day-two run has finished.&lt;/strong&gt; The new code is written and tested, but it's staged outside the benchmark directory. Changing the instrument mid-run is how you end up measuring your own edits.&lt;/p&gt;

&lt;h2&gt;
  
  
  A result from the dry run: why one axis lies
&lt;/h2&gt;

&lt;p&gt;Before any real model, I ran the new floor view over a fake backend that answers &lt;code&gt;ESCALATE&lt;/code&gt; to everything. It scored:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;task floor 0.0&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;false-confidence ceiling 0.0&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a perfect score on one axis and a useless model. It's the whole argument for two axes in one line. A model that never bluffs because it never answers isn't safe; it's absent. The floor has to be read on both axes together, or the most cowardly model wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this is going
&lt;/h2&gt;

&lt;p&gt;The reason I care about floors and not peaks: the job I want small models for isn't "be smart". It's "do one narrow part, every time, and say &lt;code&gt;ESCALATE&lt;/code&gt; when it's not yours". Several small models from &lt;em&gt;different&lt;/em&gt; families, each on a small slice, checking each other. Where they agree, that counts for something. Where they split, a human looks. The model that gets each slice should be chosen by its measured floor on that shape, not by its name or its size.&lt;/p&gt;

&lt;p&gt;Day 1's lesson was that most of my bugs looked like model behaviour. Day 2 found the same thing again from the other direction. A measuring tool I wrote for something else grouped two unrelated small scripts as "the same thing", because small things look alike to a crude measure. The fix wasn't a smarter measure. It was a floor: below a minimum size, don't call it a match. Floors, again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What shipped since Day 0
&lt;/h2&gt;

&lt;p&gt;55 PRs across 6 repos since September 30. All are public and Apache-2.0. "Not merged" means open or closed; the agent that wrote this couldn't tell which.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://github.com/forge-play/Forge" rel="noopener noreferrer"&gt;forge-play/Forge&lt;/a&gt; (10 PRs: 10 merged)
&lt;/h3&gt;

&lt;p&gt;The benchmark's home. Every change since Day 0 went in as a reviewed PR, all on September 30.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/35" rel="noopener noreferrer"&gt;#35&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs(ideas): #37, #58, #68 point at where they now live&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/36" rel="noopener noreferrer"&gt;#36&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(benchmarks): escalation benchmark fixtures, 200 invented items with a privacy gate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/37" rel="noopener noreferrer"&gt;#37&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(benchmarks): escalation runner, aggregator and prompts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/38" rel="noopener noreferrer"&gt;#38&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(benchmarks): local-model runs answer in schema, with Qwen thinking off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/39" rel="noopener noreferrer"&gt;#39&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(human_loop): break same-timestamp ties so list_queue is newest-first on coarse clocks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/40" rel="noopener noreferrer"&gt;#40&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 0.8.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/41" rel="noopener noreferrer"&gt;#41&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(escalation): Wilson intervals, exact McNemar and bootstrap Spearman in the aggregate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/42" rel="noopener noreferrer"&gt;#42&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(escalation): a Unix-socket backend, a one-model ladder with --resume/--tail, and an explicit output cap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/43" rel="noopener noreferrer"&gt;#43&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(escalation): thinking off for Gemma 4 as well as Qwen, through both backends&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/forge-play/Forge/pull/44" rel="noopener noreferrer"&gt;#44&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(escalation): hosted Kaggle arm of the escalation benchmark — tasks, converter, smoke-tested on five models&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://github.com/willow-memory/willow-bot" rel="noopener noreferrer"&gt;willow-memory/willow-bot&lt;/a&gt; (6 PRs: 6 merged)
&lt;/h3&gt;

&lt;p&gt;The local model server the local arm talks to: a loopback-only chat operation with JSON-schema output, &lt;code&gt;keep_alive&lt;/code&gt;, &lt;code&gt;done_reason&lt;/code&gt; and a think flag.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-bot/pull/80" rel="noopener noreferrer"&gt;#80&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(main): release 0.14.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-bot/pull/81" rel="noopener noreferrer"&gt;#81&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(socket): loopback-only chat op for local models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-bot/pull/82" rel="noopener noreferrer"&gt;#82&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(deterministic): chat op takes a JSON-schema format, keep_alive and returns done_reason; an unload op&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-bot/pull/83" rel="noopener noreferrer"&gt;#83&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(main): release 0.15.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-bot/pull/84" rel="noopener noreferrer"&gt;#84&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(deterministic): chat op takes the caller's think flag&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-bot/pull/85" rel="noopener noreferrer"&gt;#85&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(main): release 0.16.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://github.com/willow-memory/willow-mcp" rel="noopener noreferrer"&gt;willow-memory/willow-mcp&lt;/a&gt; (21 PRs: 20 merged, 1 not merged)
&lt;/h3&gt;

&lt;p&gt;The core server. It includes the brokered, leased model pull (#693): downloading a model is now an act that needs a grant. It also includes the session closeout fix (#706) that started this week's thread.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/687" rel="noopener noreferrer"&gt;#687&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(story): the joke lives in one place again, a test holds it, and chapter 8 lands&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/688" rel="noopener noreferrer"&gt;#688&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs: three stale claims corrected: the running Grove, the session-lifecycle draft, the closed backlog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/689" rel="noopener noreferrer"&gt;#689&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(constitutional): sync_syscall_table_at_boot queues, and does not write, when the live table is unwritable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/690" rel="noopener noreferrer"&gt;#690&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs(templates): assignment names the builder's three tools — Kart, the code graph, Nestor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/691" rel="noopener noreferrer"&gt;#691&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(manifest-grant): apply unit starts the gpg-agent it signs with&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/692" rel="noopener noreferrer"&gt;#692&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.91.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/693" rel="noopener noreferrer"&gt;#693&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(mcp): model_pull_execute, a brokered Ollama pull under a model.pull envelope and a live lease&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/694" rel="noopener noreferrer"&gt;#694&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.92.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/695" rel="noopener noreferrer"&gt;#695&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(constitutional): syscall.sync carries the sealed amendment in the request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/696" rel="noopener noreferrer"&gt;#696&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.92.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/697" rel="noopener noreferrer"&gt;#697&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(constitutional): syscall.sync apply signs the live table it writes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/698" rel="noopener noreferrer"&gt;#698&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(manifest-grant): the orchestrator seat may receive web_net, and no one else&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/699" rel="noopener noreferrer"&gt;#699&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.93.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/700" rel="noopener noreferrer"&gt;#700&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix: lease refusals name the ask path; assignment template teaches the merge first-parent diff&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/701" rel="noopener noreferrer"&gt;#701&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.93.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/702" rel="noopener noreferrer"&gt;#702&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;not merged&lt;/td&gt;
&lt;td&gt;build(deps-dev): bump ruff from 0.16.8 to 0.16.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/703" rel="noopener noreferrer"&gt;#703&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(broker): pip_sync_execute for vault-venv editable installs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/704" rel="noopener noreferrer"&gt;#704&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.94.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/705" rel="noopener noreferrer"&gt;#705&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;Docs: seal ≠ sole evidence of human verification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/706" rel="noopener noreferrer"&gt;#706&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(session): one closeout — handoff owns stack/friction/closed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willow-mcp/pull/707" rel="noopener noreferrer"&gt;#707&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 2.94.1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://github.com/willow-memory/willows-grove" rel="noopener noreferrer"&gt;willow-memory/willows-grove&lt;/a&gt; (11 PRs: 10 merged, 1 not merged)
&lt;/h3&gt;

&lt;p&gt;The planning and governance repo: the benchmark proposal with my four rulings (#95), the test that pins "naming a destination when ESCALATE was right fails the row" (#96), and today's floor-and-consistency amendment (#102, open).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/92" rel="noopener noreferrer"&gt;#92&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs(design): the Table moves up as forge-convergence 6a, T1-T7 before step 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/93" rel="noopener noreferrer"&gt;#93&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs(design): 6a is code-first: StorySession built, learner model, escalation ladder, row 11 settled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/94" rel="noopener noreferrer"&gt;#94&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs: stale claims corrected: INDEX, two built proposals, the retired allow_localhost ask, forge-convergence status&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/95" rel="noopener noreferrer"&gt;#95&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;docs(governance): propose the escalation benchmark on Kaggle, with the operator's four rulings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/96" rel="noopener noreferrer"&gt;#96&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;test(flowering): pin the G1 rule that naming a seat fails an ESCALATE-gold row&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/97" rel="noopener noreferrer"&gt;#97&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(deps): raise cryptography, aiohttp and starlette floors above the 2026 CVEs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/98" rel="noopener noreferrer"&gt;#98&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(master): release 0.12.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/99" rel="noopener noreferrer"&gt;#99&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;Desk: vault keyring paths + Grove seal prove worksheet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/100" rel="noopener noreferrer"&gt;#100&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;Docs: Slice B intake seal lineup after Grove prove&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/101" rel="noopener noreferrer"&gt;#101&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(hooks): deterministic nestor ask + composed stop + parity pin&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/willows-grove/pull/102" rel="noopener noreferrer"&gt;#102&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-02&lt;/td&gt;
&lt;td&gt;not merged&lt;/td&gt;
&lt;td&gt;docs(governance): stage the floor and consistency code in the benchmark plan&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://github.com/willow-memory/ratatosk" rel="noopener noreferrer"&gt;willow-memory/ratatosk&lt;/a&gt; (2 PRs: 2 merged)
&lt;/h3&gt;

&lt;p&gt;The runtime: its listener now asks only for the tools its seat is granted.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/ratatosk/pull/77" rel="noopener noreferrer"&gt;#77&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix: listener asks the broker only for tools its seat is granted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/willow-memory/ratatosk/pull/78" rel="noopener noreferrer"&gt;#78&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;chore(main): release 1.12.8&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://github.com/hornbook-knowledge/Jeles" rel="noopener noreferrer"&gt;hornbook-knowledge/Jeles&lt;/a&gt; (5 PRs: 3 merged, 2 not merged)
&lt;/h3&gt;

&lt;p&gt;The research and corpus tool.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PR&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Title&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/hornbook-knowledge/Jeles/pull/90" rel="noopener noreferrer"&gt;#90&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;fix(corpus): surface store-tool errors and honor WILLOW_HOME for apps root&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/hornbook-knowledge/Jeles/pull/91" rel="noopener noreferrer"&gt;#91&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-09-30&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;feat(sources): optional connectors extra over maintained scholarly clients&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/hornbook-knowledge/Jeles/pull/92" rel="noopener noreferrer"&gt;#92&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;not merged&lt;/td&gt;
&lt;td&gt;chore: rebuild the changelog section from the commits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/hornbook-knowledge/Jeles/pull/93" rel="noopener noreferrer"&gt;#93&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;merged&lt;/td&gt;
&lt;td&gt;Docs: seal catches ledger up to human verification already done&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/hornbook-knowledge/Jeles/pull/94" rel="noopener noreferrer"&gt;#94&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-10-01&lt;/td&gt;
&lt;td&gt;not merged&lt;/td&gt;
&lt;td&gt;chore: rebuild the changelog section from the commits&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most of them are boring on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  A prediction, written down before the day, then graded
&lt;/h2&gt;

&lt;p&gt;The same rule I hold the models to applies to the session that helped build this: say what you think will happen, as a spread, before it happens, then let the record grade it.&lt;/p&gt;

&lt;p&gt;At the close of the long working session on the night of October 1–2, the session (a model, working with me) wrote down a prediction for what Day 3 would open with. It wasn't a single guess; it was a distribution:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Chance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The Kaggle work: the day-two large-cloud run, then landing the floor and consistency code&lt;/td&gt;
&lt;td&gt;0.45&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Editing and publishing this Day 3 recap&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sealing a design decision for the runtime&lt;/td&gt;
&lt;td&gt;0.15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Redoing a set of governance proposals against the newer draft&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Something the record didn't contain (unforeseen)&lt;/td&gt;
&lt;td&gt;0.10&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It was saved ungraded, because the grade is mine to give, not the model's.&lt;/p&gt;

&lt;p&gt;After midnight I graded it: &lt;strong&gt;1 + (2 + ε), where ε → 0.&lt;/strong&gt; The two most likely paths both happened: the Kaggle run, and this post. The unforeseen share went to zero. One plus two is three. Day 3.&lt;/p&gt;

&lt;p&gt;That's the benchmark's point in miniature. A forecast with chances on it can be checked. A model that says "probably the Kaggle work, maybe the post, and a tenth for something I can't see" is one you can build around, and so is a model that says &lt;code&gt;ESCALATE&lt;/code&gt; when it doesn't know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[DAY 3 finish and post the large-cloud run.]&lt;/li&gt;
&lt;li&gt;Land the floor and consistency views, then run the consistency subset.&lt;/li&gt;
&lt;li&gt;Grade the four pre-registered forecasts in public, including the ones I get
wrong.&lt;/li&gt;
&lt;li&gt;1 + (2 + ε) where ε → 0, 1 - (2 - ε) where ε → 0&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Deadline: October 11.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;_&lt;/p&gt;

&lt;p&gt;ΔΣ=42&lt;br&gt;
_&lt;/p&gt;

</description>
      <category>kagglechallenge</category>
      <category>devchallenge</category>
      <category>machinelearning</category>
      <category>esclate</category>
    </item>
    <item>
      <title>Day 1: Most of My Bugs Looked Like Model Behaviour</title>
      <dc:creator>sean campbell</dc:creator>
      <pubDate>Thu, 01 Oct 2026 05:51:18 +0000</pubDate>
      <link>https://dev.to/sean_campbell_840bd62bf7e/day-1-most-of-my-bugs-looked-like-model-behaviour-388h</link>
      <guid>https://dev.to/sean_campbell_840bd62bf7e/day-1-most-of-my-bugs-looked-like-model-behaviour-388h</guid>
      <description>&lt;h2&gt;
  
  
  Update: Day 1 &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;
&lt;/h2&gt;

&lt;p&gt;The local ladder is done, and the first hosted batch is in. The frontier models haven't run yet, so none of the three predictions can be graded. This is where the numbers stand, and what I had to fix to get numbers I'd trust.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to read this.&lt;/strong&gt; Every rate below is &lt;em&gt;measured&lt;/em&gt;: it's computed from the raw replies these runs recorded, with its Wilson 95% interval. Anything I infer from those rates is marked as an inference. Following the rule from the comments, each prediction gets HIT, MISS or UNRESOLVED, graded on the interval, never on the point estimate.&lt;/p&gt;

&lt;h3&gt;
  
  
  The local ladder (8 models, 200 items each, temperature 0, on this laptop)
&lt;/h3&gt;

&lt;p&gt;False confidence means answering anyway when the right reply was &lt;code&gt;ESCALATE&lt;/code&gt;. Each model saw 40 unanswerable items.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Size (Ollama manifest)&lt;/th&gt;
&lt;th&gt;Task score&lt;/th&gt;
&lt;th&gt;False confidence&lt;/th&gt;
&lt;th&gt;95% interval&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;llama3.2:1b&lt;/td&gt;
&lt;td&gt;1.2B&lt;/td&gt;
&lt;td&gt;21.9%&lt;/td&gt;
&lt;td&gt;87.5%&lt;/td&gt;
&lt;td&gt;73.9–94.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llama3.2:3b&lt;/td&gt;
&lt;td&gt;3.2B&lt;/td&gt;
&lt;td&gt;69.4%&lt;/td&gt;
&lt;td&gt;80.0%&lt;/td&gt;
&lt;td&gt;65.2–89.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;phi4-mini&lt;/td&gt;
&lt;td&gt;3.8B&lt;/td&gt;
&lt;td&gt;86.2%&lt;/td&gt;
&lt;td&gt;92.5%&lt;/td&gt;
&lt;td&gt;80.1–97.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3:4b&lt;/td&gt;
&lt;td&gt;4.0B&lt;/td&gt;
&lt;td&gt;65.6%&lt;/td&gt;
&lt;td&gt;87.5%&lt;/td&gt;
&lt;td&gt;73.9–94.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma3:4b&lt;/td&gt;
&lt;td&gt;4.3B&lt;/td&gt;
&lt;td&gt;82.5%&lt;/td&gt;
&lt;td&gt;80.0%&lt;/td&gt;
&lt;td&gt;65.2–89.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma4:e2b&lt;/td&gt;
&lt;td&gt;4.6B&lt;/td&gt;
&lt;td&gt;81.9%&lt;/td&gt;
&lt;td&gt;95.0%&lt;/td&gt;
&lt;td&gt;83.5–98.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;llama3.1:8b&lt;/td&gt;
&lt;td&gt;8.0B&lt;/td&gt;
&lt;td&gt;88.1%&lt;/td&gt;
&lt;td&gt;77.5%&lt;/td&gt;
&lt;td&gt;62.5–87.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3.5&lt;/td&gt;
&lt;td&gt;9.7B&lt;/td&gt;
&lt;td&gt;86.9%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;37.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;24.2–53.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Only the largest local model, qwen3.5 at 9.7B, escalates on a large share of what it can't answer. For the rest, the point estimates run from 77.5% to 95%. Even the most generous lower bound among them, 62.5%, means answering anyway well over half the time. Size alone doesn't explain it, though: llama3.1 at 8B is no better than the 3B models.&lt;/p&gt;

&lt;p&gt;Two of the "4b" tags are slightly over 4B by Ollama's own count (gemma3:4b is 4.3B, gemma4:e2b is 4.6B). For prediction 2, "4B or under" means the manifest size, so those two don't count.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hosted, batch 1 (7 models, plus Kaggle's default model)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Task score&lt;/th&gt;
&lt;th&gt;False confidence&lt;/th&gt;
&lt;th&gt;95% interval&lt;/th&gt;
&lt;th&gt;Brier&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.7-flash (default)&lt;/td&gt;
&lt;td&gt;97.5%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0–8.8%&lt;/td&gt;
&lt;td&gt;0.020&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.8-flash&lt;/td&gt;
&lt;td&gt;95.5%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0–8.8%&lt;/td&gt;
&lt;td&gt;0.025&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;qwen3-235b-a22b&lt;/td&gt;
&lt;td&gt;91.7%&lt;/td&gt;
&lt;td&gt;10.0%&lt;/td&gt;
&lt;td&gt;4.0–23.1%&lt;/td&gt;
&lt;td&gt;0.104&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma-4-26b&lt;/td&gt;
&lt;td&gt;87.5%&lt;/td&gt;
&lt;td&gt;2.6%&lt;/td&gt;
&lt;td&gt;0.5–13.5%&lt;/td&gt;
&lt;td&gt;0.032&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;claude-haiku-4.5&lt;/td&gt;
&lt;td&gt;86.6%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;35.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;20.7–54.2%&lt;/td&gt;
&lt;td&gt;0.165&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-oss-20b&lt;/td&gt;
&lt;td&gt;86.2%&lt;/td&gt;
&lt;td&gt;7.5%&lt;/td&gt;
&lt;td&gt;2.6–19.9%&lt;/td&gt;
&lt;td&gt;0.144&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-5.4-nano&lt;/td&gt;
&lt;td&gt;84.4%&lt;/td&gt;
&lt;td&gt;7.5%&lt;/td&gt;
&lt;td&gt;2.6–19.9%&lt;/td&gt;
&lt;td&gt;0.153&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deepseek-r1&lt;/td&gt;
&lt;td&gt;65.0%&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;0–11.4%&lt;/td&gt;
&lt;td&gt;0.047&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One commenter bet the frontier models would be the bluffers. On this batch it went the other way. Every hosted model's false-confidence interval sits entirely below every local model's except qwen3.5's: the highest hosted upper bound is haiku's 54.2%, and the lowest local lower bound is 62.5%. Haiku is the outlier among the hosted ones, though. Its lower bound clears 20%, just barely (20.7%). It's a small model, though, not one of the frontier models P1 is about, so it isn't counted there.&lt;/p&gt;

&lt;p&gt;That comparison is an inference across two different setups. The local and hosted models didn't run through the same client, and the reasoning settings differ (see caveats). I haven't tested whether the gap holds once those are matched.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where the predictions stand
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Some frontier model's false-confidence lower bound clears 20%:&lt;/strong&gt; UNRESOLVED. No frontier model has run yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The best local model at 4B or under beats at least one frontier model, by exact McNemar:&lt;/strong&gt; UNRESOLVED. The test can't run without frontier results. On point estimates it leans toward a miss. The best local model at 4B or under by manifest size is llama3.2:3b, at 80%. Every hosted model so far except haiku is at 10% or less. But that's a lean, not a grade.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task score and false confidence are only weakly related (bootstrap Spearman):&lt;/strong&gt; UNRESOLVED. The local arm alone gives ρ = −0.49, with an interval from −0.96 to 0.44. That's the "can't tell either way" shape the comments said to expect from about 8 models.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  What broke on the way
&lt;/h3&gt;

&lt;p&gt;Three smoke rounds before the paid run each turned up a harness bug that would have read as model behaviour. Each one below was observed in the raw replies, not inferred:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A loose answer format came back empty.&lt;/strong&gt; Route and classify allowed "any object" as the answer. Gemini's structured output returned &lt;code&gt;{}&lt;/code&gt; for it, which scored 0% on two task shapes. Typing the answer fixed it, and the same model then scored 97.8% and 100% on those shapes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every provider has its own rules.&lt;/strong&gt; OpenAI's reasoning models reject temperature 0 and expect &lt;code&gt;max_completion_tokens&lt;/code&gt;. Its strict mode refuses any open object. Anthropic refused the 20-tool route format with "compiled grammar is too large", so haiku's 60 route items in this batch are errors, not answers. That's still open before the frontier run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate limits look like model errors.&lt;/strong&gt; In one smoke run, 169 of DeepSeek's 200 calls were refused for load. With a bounded retry on rate limits alone, and the attempt count recorded on every row, the next run had 1. Batch 1 had none.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The output cap is part of the measurement.&lt;/strong&gt; DeepSeek-R1 reasons in visible text inside its 512-token budget, and 29.5% of its replies were cut off mid-JSON. I score those as unparseable, because the cap is part of the condition being measured, not something to work around.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Dropping a broken reply flatters the model.&lt;/strong&gt; A reply that broke the answer format was being filed as an error and left out of the score. It now counts as unparseable, the same way the local arm treats it. Re-scoring the smoke results under the new rule moves one model's route score from 89% to 83%. That figure is a re-scoring of recorded replies, not a fresh run.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What I haven't checked yet
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Matched reasoning.&lt;/strong&gt; The Gemini and OpenAI models get &lt;code&gt;reasoning="low"&lt;/code&gt; and the others get nothing. Part of Gemini's lead could be that thinking budget rather than the model. A matched control run is next, and until it's done I won't claim Gemini is better at knowing what it doesn't know.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tuning on the default model.&lt;/strong&gt; Gemini Flash is Kaggle's default model, so every fix above was first checked on it. I'm stating that, not correcting for it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calibration.&lt;/strong&gt; The Brier scores spread the hosted models further apart than task score does. That's a pattern in one batch, and I haven't tested it for significance.&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Who scored what.&lt;/strong&gt; The grading rules are written down, and the code that applies them is tested, but nobody outside the build has re-scored these replies.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;If you would like to see the work behind the work&lt;/strong&gt;: &lt;a href="https://github.com/forge-play/Forge" rel="noopener noreferrer"&gt;The Forge&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ΔΣ=42&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Does your model know when it doesn't know? A benchmark for the ESCALATE answer</title>
      <dc:creator>sean campbell</dc:creator>
      <pubDate>Wed, 30 Sep 2026 09:36:03 +0000</pubDate>
      <link>https://dev.to/sean_campbell_840bd62bf7e/does-your-model-know-when-it-doesnt-know-a-benchmark-for-the-escalate-answer-268o</link>
      <guid>https://dev.to/sean_campbell_840bd62bf7e/does-your-model-know-when-it-doesnt-know-a-benchmark-for-the-escalate-answer-268o</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Most leaderboards ask one question: did the model get it right? I wanted to ask a second one: &lt;strong&gt;does the model know when it can't?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I run a small multi-agent system on one laptop, where local models hand work up to bigger ones. In a setup like that, a small model that answers wrong is worse than one that says "I can't do this, pass it up." A modest model that knows its limits can sit safely in the chain. A confident one that bluffs can't.&lt;/p&gt;

&lt;p&gt;So every task in this benchmark has a refusal token, &lt;code&gt;ESCALATE&lt;/code&gt;. In &lt;strong&gt;one item out of five, the answer has been deliberately removed&lt;/strong&gt;, or the document doesn't contain it. On those items, &lt;code&gt;ESCALATE&lt;/code&gt; is the only correct reply.&lt;/p&gt;

&lt;p&gt;There are 200 items across four everyday job shapes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Route&lt;/strong&gt; (60): a one-line request → pick the right tool and arguments from a 20-tool catalogue, or escalate when no tool fits or an argument is missing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classify&lt;/strong&gt; (50): a short work-log note → status, severity, and whether a human is needed, or escalate when the note doesn't say.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Judge&lt;/strong&gt; (50): a claim and a document → SUPPORTS, CONTRADICTS or UNRELATED, or escalate when the document is on topic but silent on the claim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ground&lt;/strong&gt; (40): a passage and a question → answer verbatim from the passage, or escalate when it isn't there.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Every model gets &lt;strong&gt;two scores&lt;/strong&gt;: its &lt;em&gt;task score&lt;/em&gt; on the answerable items, and its &lt;em&gt;false-confidence rate&lt;/em&gt;, meaning how often it answered anyway when the right reply was &lt;code&gt;ESCALATE&lt;/code&gt;. Each answer also carries a stated confidence, so I can draw a reliability diagram too.&lt;/p&gt;

&lt;p&gt;Every item is invented from scratch, and nothing is scraped. A privacy gate checks the whole set before it's published.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;Two groups, on one chart:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kaggle's hosted model suite:&lt;/strong&gt; the frontier models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A local ladder running on my laptop:&lt;/strong&gt; 1B, 3B, 4B and 8B open models, run on a 4MB GPU at temperature 0, with every raw call kept.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The question behind the chart: &lt;em&gt;can a laptop's 3B know its own limits as well as a frontier model knows its own?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The runs are in progress. Before any model touched the fixtures, I wrote down my predictions and timestamped them so they can't drift toward the results:&lt;/em&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;At least one frontier model answers anyway on &lt;strong&gt;more than 20%&lt;/strong&gt; of the unanswerable items. &lt;em&gt;(my confidence: 75%)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;The best small local model (4B or under) has a &lt;strong&gt;lower&lt;/strong&gt; false-confidence rate than at least one frontier model. &lt;em&gt;(40%)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;Task score and false confidence are only &lt;strong&gt;weakly related&lt;/strong&gt; across models (Spearman below 0.5). &lt;em&gt;(60%)&lt;/em&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I'll grade those here, misses included, once the numbers are in.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-route" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-route&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-classify" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-classify&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-judge" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-judge&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-ground" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/tasks/rudi193/escalation-bench-ground&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ΔΣ=42&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Willow – local-first AI stack, phone reads desktop KB over LAN, no cloud relay</title>
      <dc:creator>sean campbell</dc:creator>
      <pubDate>Sun, 26 Apr 2026 01:42:55 +0000</pubDate>
      <link>https://dev.to/sean_campbell_840bd62bf7e/willow-local-first-ai-stack-phone-reads-desktop-kb-over-lan-no-cloud-relay-189c</link>
      <guid>https://dev.to/sean_campbell_840bd62bf7e/willow-local-first-ai-stack-phone-reads-desktop-kb-over-lan-no-cloud-relay-189c</guid>
      <description>&lt;p&gt;A phone running Willow on Termux sent a signed command to a desktop running Willow on Linux. The response came back in under a second — 70,000 knowledge atoms, live system status, no Discord, no Telegram, no API call to a third party.                                                                                                      &lt;/p&gt;

&lt;p&gt;Willow is a full local-first AI stack: Postgres/SQLite knowledge graph that persists across sessions and models, 40+ MCP tools, a skill system that works with any LLM (Ollama is the default, cloud keys are optional addons), and a LAN command server — HMAC-SHA256 auth, shared token, ~100 lines of Python.                                     &lt;/p&gt;

&lt;p&gt;The knowledge graph is yours. The inference runs on your hardware. The nodes talk directly. Nothing requires permission from a provider or a credit card on file.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; GitHub: https://github.com/rudi193-cmd/willow-1.9                                                                 
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The full concept paper is in docs/CONCEPT.md — it explains why this doesn't exist yet and what it took to build it.                    &lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>opensource</category>
      <category>security</category>
    </item>
  </channel>
</rss>
