<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Grzegorz Brzezinka</title>
    <description>The latest articles on DEV Community by Grzegorz Brzezinka (@agentgreg).</description>
    <link>https://dev.to/agentgreg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149387%2Fc7cab86e-8067-46bd-abb8-84ad0ddc5a65.jpg</url>
      <title>DEV Community: Grzegorz Brzezinka</title>
      <link>https://dev.to/agentgreg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/agentgreg"/>
    <language>en</language>
    <item>
      <title>Your AI thinks too long: System 1 AI explained</title>
      <dc:creator>Grzegorz Brzezinka</dc:creator>
      <pubDate>Wed, 30 Sep 2026 19:15:33 +0000</pubDate>
      <link>https://dev.to/agentgreg/your-ai-thinks-too-long-system-1-ai-explained-cd7</link>
      <guid>https://dev.to/agentgreg/your-ai-thinks-too-long-system-1-ai-explained-cd7</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://agentgreg.ai/research/system-one-ai/" rel="noopener noreferrer"&gt;agentgreg.ai&lt;/a&gt; (&lt;a href="https://agentgreg.ai/pl/research/system-one-ai/" rel="noopener noreferrer"&gt;Polish version&lt;/a&gt;). Updated 1 October 2026 with tests on about 400 Polish decisions, including raw Jev, and a follow-up on rule cases.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The short version.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Daniel Kahneman split human thinking into a fast, automatic System 1 and a slow, effortful System 2. Since OpenAI's o1 in September 2024, AI vendors have poured effort into the System 2 side: models that "think" before they answer. That helps on maths and code, and it costs tokens and seconds on every call. In September 2026 TypeSafe AI, co-founded by former OpenAI researcher Diogo Almeida, launched &lt;strong&gt;Jev&lt;/strong&gt;, a model that writes no text at all. It gets a situation and a list of allowed answers and returns one of them with a probability. Remigiusz Kinas has already built an open Polish counterpart, &lt;strong&gt;basal&lt;/strong&gt;, on top of Bielik. Research from Meta shows where fast answers work (judging answers with 4 tokens instead of 2,118) and where they collapse (school maths). I then tested both on about 400 Polish business decisions, from customer service and from twelve other fields. Jev was right 98.0% and 94.6% of the time in under 300 milliseconds, for about two cents per thousand decisions, and at a strict confidence threshold it settled more than three quarters of the cases on its own without a single mistake. basal, which runs on your own hardware, reached 86.9% and 81.9%. A model that thinks before answering got every case right, at 3 to 4 seconds each. For Jev, nearly all of the gap to the thinking model is in rules with conditions; basal also loses points elsewhere. In a follow-up test, splitting those rules into simple questions and doing the date and amount comparisons in code brought Jev to 99 of 99 rule cases and basal-4.5B to 96.0%. The article ends with a checklist for trying this on your own process.&lt;/p&gt;

&lt;h2&gt;
  
  
  A question for the first second
&lt;/h2&gt;

&lt;p&gt;A bat and a ball cost $1.10 in total. The bat costs $1.00 more than the ball. How much does the ball cost?&lt;/p&gt;

&lt;p&gt;If "10 cents" arrived before you finished reading, you are in good company. Shane Frederick put this question into his Cognitive Reflection Test in 2005, and Kahneman reports that more than half of the students at Harvard, MIT and Princeton gave the intuitive answer. The right one is 5 cents: the bat costs $1.05, which is one dollar more.&lt;/p&gt;

&lt;p&gt;Kahneman built &lt;em&gt;Thinking, Fast and Slow&lt;/em&gt; (2011) around this kind of moment. &lt;strong&gt;System 1&lt;/strong&gt; answers at once and without effort: it recognises a face in a crowd, and it hands you "10 cents". &lt;strong&gt;System 2&lt;/strong&gt; is what you need for 17 × 24. It is slow and tiring, and it is the part that can check what System 1 proposed. The labels themselves come from the psychologists Keith Stanovich and Richard West (2000). Kahneman treats the two systems as a convenient way to talk about the mind; he does not place them anywhere in the brain.&lt;/p&gt;

&lt;p&gt;One honest caveat before I build on the book. Not all of it survived psychology's replication crisis. The chapter on social priming did not hold up, and Kahneman said so himself in 2017: "I placed too much faith in underpowered studies." The two-systems model and the bat-and-ball result are on much firmer ground, and they are all this article relies on.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuvhh92medizbyjmhn2ig.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuvhh92medizbyjmhn2ig.png" alt="The bat-and-ball question answered by the two systems. System 1 lands first with 10 cents, which is wrong. System 2 arrives later with 5 cents, which is right. The drawing shows the order of arrival; the times are schematic." width="800" height="389"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The bat-and-ball question answered by the two systems. System 1 lands first with 10 cents, which is wrong. System 2 arrives later with 5 cents, which is right. The drawing shows the order of arrival; the times are schematic.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Two years of teaching AI to think slowly
&lt;/h2&gt;

&lt;p&gt;The AI field borrowed Kahneman's vocabulary early. At NeurIPS in December 2019 Yoshua Bengio gave a keynote titled "From System 1 Deep Learning to System 2 Deep Learning": neural networks had become excellent at perception, and the next step was reasoning. In November 2023 Andrej Karpathy described the large language models of the time as pure System 1, sampling one word after another with no way to stop and check.&lt;/p&gt;

&lt;p&gt;Then the industry went after System 2 in earnest. OpenAI's o1 (September 2024) spends extra computation writing a long internal chain of reasoning before it answers. Anthropic presented Claude 3.7 Sonnet (February 2025) as the first hybrid model that can answer at once or think at length. Qwen3 (April 2025) shipped one model with a thinking switch, and in July 2025 the Qwen team split it into separate fast and thinking models again, saying that was the way to get the best quality out of each. GPT-5 (August 2025) added a router that decides for you when a question deserves the slow path.&lt;/p&gt;

&lt;p&gt;For a business the bill arrives in two currencies. Thinking is paid in tokens: on Claude, for example, thinking tokens are billed as output tokens, the most expensive kind. And it is paid in time. A thinking model can take tens of seconds before the first word of the answer. A chat widget can live with that, a phone line cannot: across ten languages, Stivers and colleagues (PNAS, 2009) found that people answer each other within 0 to 200 milliseconds of the end of a turn.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10sl3ekcd3hxobg5gu95.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10sl3ekcd3hxobg5gu95.png" alt="From Bengio's 2019 keynote to GPT-5's router in 2025, the industry's effort went into slower, more deliberate models. Jev, launched in September 2026, goes the other way: a model built only for fast decisions." width="799" height="368"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;From Bengio's 2019 keynote to GPT-5's router in 2025, the industry's effort went into slower, more deliberate models. Jev, launched in September 2026, goes the other way: a model built only for fast decisions.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Jev: a model that does not write
&lt;/h2&gt;

&lt;p&gt;On 15 September 2026 TypeSafe AI, a San Francisco company, came out of stealth with a $40M seed round led by DCVC. Its CEO Diogo Almeida spent about four years at OpenAI working on InstructGPT, ChatGPT and GPT-4, after Google Brain. The launch post says it plainly: "We were inspired by Daniel Kahneman, Thinking, Fast and Slow." The model is named after William Stanley Jevons, the economist behind the Jevons paradox, which says that making a resource cheaper tends to increase how much of it we use.&lt;/p&gt;

&lt;p&gt;Jev does not generate text. You send it a situation and a typed question, and it returns a probability distribution over the answers you allowed. There are three kinds of question: &lt;em&gt;choice&lt;/em&gt; (pick one of these), &lt;em&gt;score&lt;/em&gt; (rate this) and a yes/no type TypeSafe calls &lt;em&gt;noul&lt;/em&gt;. The company quotes 70 to 500 milliseconds per decision, $0.042 per million input tokens and nothing for output, because there is no output text to pay for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two real calls from my test, basal-4.5B:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Dzień dobry, kupiłem u Państwa czajnik elektryczny (zamówienie 45821). Po tygodniu przestał grzać, lampka się świeci, ale woda jest zimna. Proszę o wymianę na nowy albo zwrot pieniędzy." (I bought an electric kettle from you, order 45821. After a week it stopped heating. Please replace it or refund me.)&lt;br&gt;
&lt;/p&gt;


&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;question: Is this a complaint?       (yes / no)
answer:   yes  0.990                 112 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;"Moja paczka miała przyjść w poniedziałek, jest czwartek, a status w śledzeniu od trzech dni to „w drodze”. Gdzie jest moja przesyłka?" (My parcel was due Monday, it is Thursday, and tracking has said "in transit" for three days.)&lt;br&gt;
&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;question: Which team handles it?    (complaints / billing / delivery / technical / sales)
answer:   delivery  0.997            117 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A generative model given the same job writes a sentence, and your code has to hope the sentence contains one of the words it expects. A decision model cannot answer outside the list. TypeSafe markets this as a model that "can't hallucinate", which is true only in a narrow sense: it can still pick the wrong option from the list, it just cannot invent a new one. And the headline numbers, 193.6 times faster and 444.6 times cheaper, come from TypeSafe's own tests on four in-house workflows, scored against the averaged answers of two large models rather than human-checked labels. TypeSafe itself calls them "on the higher end". There is no technical paper and no public weights, so those are claims, not results. My own measurements of Jev are further down, and they turned out better than I expected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1cy8k45548rt2e1zcqmp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1cy8k45548rt2e1zcqmp.png" alt="Two ways to ask a model whether an email is a complaint. A generative model writes tokens that you pay for and then parse. A decision model returns one of the allowed answers with a probability in a single pass. An illustration of the two call shapes, not a benchmark." width="800" height="421"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Two ways to ask a model whether an email is a complaint. A generative model writes tokens that you pay for and then parse. A decision model returns one of the allowed answers with a probability in a single pass. An illustration of the two call shapes, not a benchmark.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first independent use I know of is academic. Researchers at the University of Texas at Dallas published Jev-Mem (arXiv, 21 September 2026), which uses Jev to run an AI agent's memory: deciding what kind of memory an observation is, which memories relate, and when to stop searching. A normal language model writes only the final answer. On the LoCoMo long-conversation benchmark, where another language model grades the answers, they report an overall score of 0.777 against 0.700 for their strongest baseline, with 158 seconds to build the memory against 1,044 seconds for the fastest competitor and 0.93 seconds per query. It is one benchmark, the strongest baseline is the authors' own earlier system, and the text disagrees with its own tables in places, so I read it as an early signal of how people will use these models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a fast decision is enough
&lt;/h2&gt;

&lt;p&gt;Look at the traffic that runs through a typical customer-service or back-office process, and most of the questions an AI is asked there have a closed answer. Which team gets this ticket. Is the customer angry. Is this a complaint under the returns policy. Does the form have every required field, as long as the rule is simple. Should a person take over now. Is this reply safe to send. None of them needs an essay, and every one of them runs thousands of times a day, so speed and cost add up.&lt;/p&gt;

&lt;p&gt;There is also a strong research result behind the idea. In "Distilling System 2 into System 1" (Meta, July 2024) Ping Yu and colleagues let a model reason at length many times over, kept the answers it agreed with itself on, and trained it to give those answers directly. Used as a judge of answer quality, the fast version agreed with human raters &lt;strong&gt;58.4%&lt;/strong&gt; of the time using &lt;strong&gt;4 tokens&lt;/strong&gt;, against 49.1% for the slow reasoning method using 2,118 tokens, and it beat GPT-4 on that task. On a test of resisting leading, biased questions the fast version scored 81.3% against 76.0% for the slow method, with 56 tokens instead of 147.&lt;/p&gt;

&lt;p&gt;The same paper is also the best warning label. On GSM8k, a set of school maths word problems, the fast version scored &lt;strong&gt;7.1%&lt;/strong&gt;, against 52.8% with step-by-step reasoning. Some questions cannot be answered without working them out, for models as for people. The practical rule is to give the fast model decisions with a closed answer, and keep the slow model, or a human, for anything that has to be worked out.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcakfyj5mt874npaau34m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcakfyj5mt874npaau34m.png" alt="Meta's distillation results on Llama-2-70B. As a judge of answer quality the fast model agrees with humans 58.4% of the time with 4 tokens, against 49.1% for the reasoning method with 2,118 tokens. On school maths (GSM8k) the fast model gets 7.1% right, against 52.8% with step-by-step reasoning." width="799" height="571"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Meta's distillation results on Llama-2-70B. As a judge of answer quality the fast model agrees with humans 58.4% of the time with 4 tokens, against 49.1% for the reasoning method with 2,118 tokens. On school maths (GSM8k) the fast model gets 7.1% right, against 52.8% with step-by-step reasoning.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Speed also changes which products you can build at all. A voice agent that pauses for several seconds on every turn feels broken, whatever the quality of the answer. A decision in tens of milliseconds fits inside the gap people leave in normal conversation, so routing, escalation and safety checks can run on every turn without the caller noticing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh8y4pmzfp6d69xsaiek.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh8y4pmzfp6d69xsaiek.png" alt="Median time per decision on the 200-case customer-service set, on a logarithmic scale, against the 0 to 200 millisecond gap people leave between turns in conversation. basal-4.5B takes 12.5 milliseconds on an H100 according to its report and 116 on my Mac, basal-1.5B 46. Jev 1.13 takes 288 milliseconds measured from Poland, network included. The two models that think first take about 4 seconds." width="800" height="602"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Median time per decision on the 200-case customer-service set, on a logarithmic scale, against the 0 to 200 millisecond gap people leave between turns in conversation. basal-4.5B takes 12.5 milliseconds on an H100 according to its report and 116 on my Mac, basal-1.5B 46. Jev 1.13 takes 288 milliseconds measured from Poland, network included. The two models that think first take about 4 seconds.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Polish version already exists
&lt;/h2&gt;

&lt;p&gt;At the end of September, two weeks after Jev's launch, Remigiusz Kinas, one of the people behind the Bielik models, released &lt;strong&gt;basal&lt;/strong&gt;: an open, Apache-2.0 family of decision models for Polish and English, with the same kind of interface as Jev. TypeSafe never published how Jev is trained, so basal is Kinas's own recipe: Bielik-4.5B-v3.0 and Bielik-1.5B-v3.0 fine-tuned on 63,663 decision examples (deadlines from the Polish holiday calendar, amounts and VAT, questions about statutes, whether the evidence is sufficient), each either generated by code, grounded in statute text or checked by two independent verifier models, about a fifth of them in English. On top comes calibration, so that a 0.9 means the model is right about nine times in ten. One design choice matters for Polish in particular: the Bielik v3 tokenizer needs 17 to 51% fewer tokens for Polish text than the alternatives he measured, which makes every decision cheaper.&lt;/p&gt;

&lt;p&gt;His results are strong on his home ground. On his own set of 7,081 Polish decisions, basal-4.5B is right &lt;strong&gt;88.4%&lt;/strong&gt; of the time against 78.0% for Jev. The number I like most is coverage: if you only let the model decide when it is confident enough to keep errors near 1%, basal-4.5B still decides &lt;strong&gt;58.6%&lt;/strong&gt; of the cases on its own, and Jev 18.1%. Everything else goes to a person or a slower model. On an H100 GPU one decision takes 12.5 milliseconds. Kinas is also candid about the limits: that test set is his own and unpublished, and on the public English JevBench basal-4.5B scores 0.740 against 0.861 for Jev. That is exactly why an independent test on someone else's data is worth running.&lt;/p&gt;

&lt;h2&gt;
  
  
  My own tests on about 400 Polish decisions
&lt;/h2&gt;

&lt;p&gt;Kinas's numbers come from his own test set and Jev's from TypeSafe's, so I built my own, in three steps. A pilot of thirty short customer-service messages came first. Then 200 customer-service decisions, the thirty included, of six kinds: is it a complaint, which team gets it, should a person take over, is the form complete under a stated rule, how irritated is the customer, and is the message phishing. Then 204 decisions from twelve other fields: public offices, a law firm's intake, HR, a clinic's front desk, insurance claims, bank compliance, logistics, IT security, factory quality control, a university, a housing association and marketplace moderation. About 30% of those are rules with conditions, deadlines or amount thresholds.&lt;/p&gt;

&lt;p&gt;The cases were written for this test with AI help. I fixed the correct answers before any model saw them, and then two models from different vendors, Claude Opus 5.5 and GPT-6.1 Sol, answered every case blind as checkers. Only cases where my label and both checkers agree count: 198 of the 200 customer-service cases and all 204 of the others. On that ground I compared basal-4.5B and basal-1.5B on my Mac (Apple M5 Max), raw Jev 1.13 through OpenRouter, which exposes it directly, Mistral Medium 3.1 as a fast general model, and two models that think before they answer: Gemini 3.8 Flash in the cloud and Qwen3-14B on the Mac.&lt;/p&gt;

&lt;p&gt;Jev surprised me. It was right &lt;strong&gt;98.0%&lt;/strong&gt; of the time on the customer-service set and &lt;strong&gt;94.6%&lt;/strong&gt; on the twelve-field set, in a median 288 and 295 milliseconds measured from Poland with the network included, for $0.019 and $0.024 per thousand decisions, with one call per decision. basal-4.5B reached 86.9% and 81.9% in 116 and 142 milliseconds on the laptop. Gemini with thinking got every case right in both sets, and needed 4.2 and 3.0 seconds per decision and $0.91 and $0.98 per thousand to do it. Qwen3-14B with thinking landed at 94.4% and 96.1%, so thinking alone does not guarantee the top score. In the thirty-case pilot basal made the same three mistakes as before, and Jev made none.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs0wwq59dusrbdtaclpu7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs0wwq59dusrbdtaclpu7.png" alt="Share of correct answers on 198 customer-service decisions and 204 decisions from twelve other fields, counting only cases where my label and two blind checkers agree. Jev 1.13: 98.0% and 94.6%. basal-4.5B: 86.9% and 81.9%. Gemini 3.8 Flash with thinking: 100% on both." width="800" height="441"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Share of correct answers on 198 customer-service decisions and 204 decisions from twelve other fields, counting only cases where my label and two blind checkers agree. Jev 1.13: 98.0% and 94.6%. basal-4.5B: 86.9% and 81.9%. Gemini 3.8 Flash with thinking: 100% on both.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For Jev, almost the whole gap sits in the rules. On the 60 rule-based cases of the twelve-field set Gemini got all 60 right, Qwen3 with thinking 56, Jev 50 and basal-4.5B 38. On the other 144 cases Jev made a single mistake. When the fast models got a rule wrong they mostly said yes where the rule said no. Working something out step by step is what the Meta paper found fast answers bad at, and checking a rule's conditions is that kind of work.&lt;/p&gt;

&lt;p&gt;Can the rules gap be closed without a slower model? The basal documentation recommends splitting a compound rule into simple yes/no questions and computing numbers in code, so I tested exactly that on 99 rule cases: the 60 from the twelve-field set and 39 completeness checks from customer service. Splitting alone helped only a little: basal-4.5B went from 71.7% to 76.8% and Jev from 87.9% to 90.9%. The reason showed up in the sub-questions. Both models read facts from the text well (97.2% and 99.0% of reading questions right) and compare dates and amounts badly (70.1% and 82.2%). When the date and amount comparisons are done by a few lines of Python and the models only read the text, basal-4.5B reaches &lt;strong&gt;96.0%&lt;/strong&gt; and Jev gets &lt;strong&gt;all 99 cases right&lt;/strong&gt;, still at about 290 milliseconds, against 3.5 seconds for Gemini, which also gets all 99. Two honest limits: the dates and amounts that the code compares were copied from the text by my test harness, so this assumes they are read correctly, and on 15 of the 60 twelve-field cases every condition was arithmetic, so the code decided alone. On the other 45 the gain holds: Jev went from 38 to 45 right. The decompositions and results are in the &lt;a href="https://github.com/agentGreg/system-1-ai/tree/main/experiments/05-rule-decomposition" rel="noopener noreferrer"&gt;repository&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frfmw4hygjk9g6jrkxd9e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frfmw4hygjk9g6jrkxd9e.png" alt="Right answers on 99 rule cases for the whole rule asked at once, the rule split into simple yes/no questions, and the split rule with date and amount comparisons done in code. basal-4.5B: 71.7%, 76.8%, 96.0%. Jev 1.13: 87.9%, 90.9%, 100%. Gemini 3.8 Flash with thinking gets 100% on the whole rule." width="800" height="441"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Right answers on 99 rule cases for the whole rule asked at once, the rule split into simple yes/no questions, and the split rule with date and amount comparisons done in code. basal-4.5B: 71.7%, 76.8%, 96.0%. Jev 1.13: 87.9%, 90.9%, 100%. Gemini 3.8 Flash with thinking gets 100% on the whole rule.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;For a manager, the most useful number is about confidence. Set a threshold, let the model decide alone only when its probability reaches it, and send everything below to a person. At basal's own strict threshold (confidence of at least 0.913), Jev settled &lt;strong&gt;80.3%&lt;/strong&gt; of the customer-service cases and &lt;strong&gt;77.0%&lt;/strong&gt; of the twelve-field cases on its own without a single mistake, asking each question twice with the answer list reversed and averaging, as basal does, to cancel any preference for the first option. basal-4.5B settled 67.2% and 57.8%, with two cases (1.5%) and 3.4% of those wrong, the second clearly above the 1% its author aimed for, while the share it settled matches the 58.6% Kinas published. The two models also err in opposite directions. Jev is careful: every answer it gave with confidence of 0.9 or more in the twelve-field set was right. basal is bolder than it should be: on customer service, when it reported a confidence of about 0.85 it was right 68% of the time.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F830cu9x2jvmfxm045e20.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F830cu9x2jvmfxm045e20.png" alt="Share of cases each model settles on its own at confidence of 0.913 or more, basal's own strict threshold, with both option orders averaged. Customer service: Jev 80.3% with no errors, basal-4.5B 67.2% with 1.5% of them wrong. Twelve other fields: Jev 77.0% with no errors, basal-4.5B 57.8% with 3.4% wrong." width="799" height="571"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Share of cases each model settles on its own at confidence of 0.913 or more, basal's own strict threshold, with both option orders averaged. Customer service: Jev 80.3% with no errors, basal-4.5B 67.2% with 1.5% of them wrong. Twelve other fields: Jev 77.0% with no errors, basal-4.5B 57.8% with 3.4% wrong.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;None of this makes basal pointless. It runs on your own hardware, so no customer message leaves the building, its weights are open under Apache-2.0, and on an H100 it answers in 12.5 milliseconds. Its weak spots are specific: rules with conditions, the middle of a scale (it read five of the six "slightly irritated" customers as "very irritated"), and a customer asking whether our own text message was genuine, which it flagged as phishing. My results also disagree with Kinas's, where basal-4.5B led Jev 88.4% to 78.0% on his set. Different test sets reward different things, which is the best argument I know for testing on your own traffic before you choose.&lt;/p&gt;

&lt;p&gt;The limits, on the record. Both sets were built for this test, so they are cleaner than a real inbox. The labels were checked by models rather than people, and in the twelve-field set the cases were drafted with a Claude-based tool while one of the checkers was also Claude, so a human review of the rule-based cases is the next thing I want. Each system ran once. basal and Qwen3 ran on a laptop, while the cloud numbers include the network from Poland. Jev's numbers are a snapshot of its alpha endpoint at the end of September 2026. The cases, labels, scripts and raw outputs are on GitHub at &lt;a href="https://github.com/agentGreg/system-1-ai" rel="noopener noreferrer"&gt;agentGreg/system-1-ai&lt;/a&gt;, so you can rerun them or add your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which decisions to hand over
&lt;/h2&gt;

&lt;p&gt;Two questions sort almost every AI decision in a process. How open is the answer: a fixed list of options, or free-form text? And what does a mistake cost? Closed answers with a cheap mistake belong to a fast model outright. Closed answers with an expensive mistake still suit a fast model, as long as you use its probability as a gate and hand everything below the threshold to a person. Free-form work, a reply to draft or a thread to summarise, belongs to a generative model, with thinking switched on where the task needs working out. And where the answer is open and the stakes are high, AI supports a person who decides.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flx6smrsl1peiofece903.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flx6smrsl1peiofece903.png" alt="A rule of thumb for splitting work between models, drawn from experience. Horizontal: how open the answer is. Vertical: the cost of a mistake. Fast decision models fit the left column, on their own when mistakes are cheap and with a confidence threshold and human hand-over when they are not." width="800" height="617"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;A rule of thumb for splitting work between models, drawn from experience. Horizontal: how open the answer is. Vertical: the cost of a mistake. Fast decision models fit the left column, on their own when mistakes are cheap and with a confidence threshold and human hand-over when they are not.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If you want to try this on your own process, this is the order I would do it in.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;List the closed decisions.&lt;/strong&gt; Walk one real process end to end and write down every point where a person or a model picks from a known set of answers. Most teams find more than they expect.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Label a few hundred real cases.&lt;/strong&gt; Take them from your own tickets and documents, labelled by the people who do the work today, because your inbox has its own language.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure coverage at your error target&lt;/strong&gt; as well as accuracy. Decide how many mistakes you can live with, then ask what share of cases the model handles alone at that level. That share is your saving.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Route what it is unsure of.&lt;/strong&gt; Below the threshold, send the case to a person or to a slower model. The fast model's probability is the switch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Time it end to end&lt;/strong&gt;, including the network. A model that answers in 12 milliseconds on a GPU is slower across the Atlantic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-check regularly.&lt;/strong&gt; Your customers change their language faster than your test set does.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What I'm looking at next
&lt;/h2&gt;

&lt;p&gt;Two things. A human review of the rule-based cases, since that is where every fast model loses its points and where a wrong label would hurt most, and a version of the rule test where the model also extracts the dates and amounts itself. And real traffic instead of written cases: decisions from a live customer-service process, measured against a thinking model on cost and on agreement with the people who do the work today. The results will go here.&lt;/p&gt;

&lt;p&gt;This is an analysis of public work plus experiments of my own. Jev's figures are the vendor's claims; basal's figures come from its technical report; the Meta figures from the paper below.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev" rel="noopener noreferrer"&gt;TypeSafe AI, introducing Jev →&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/rkinas/basal" rel="noopener noreferrer"&gt;basal on GitHub →&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/agentGreg/system-1-ai" rel="noopener noreferrer"&gt;My test set and scripts →&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2407.06023" rel="noopener noreferrer"&gt;Yu et al., Distilling System 2 into System 1 →&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.23986" rel="noopener noreferrer"&gt;Jiang et al., Jev-Mem →&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.pnas.org/doi/10.1073/pnas.0903616106" rel="noopener noreferrer"&gt;Stivers et al., turn-taking →&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thanks to Remigiusz Kinas for publishing basal openly, with an unusually honest technical report. Related reading on this site: &lt;a href="https://agentgreg.ai/research/bielik-graded-familiarity/" rel="noopener noreferrer"&gt;A confidence dial you can read, and turn&lt;/a&gt;, my paper on how well Polish models know what they don't know.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
