<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: marcosomma</title>
    <description>The latest articles on DEV Community by marcosomma (@marcosomma).</description>
    <link>https://dev.to/marcosomma</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3064224%2Fe7c4c99c-97ab-42cf-89b1-cf19e9318cae.jpeg</url>
      <title>DEV Community: marcosomma</title>
      <link>https://dev.to/marcosomma</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/marcosomma"/>
    <language>en</language>
    <item>
      <title>What an anthill can teach us about orchestrating agents.</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Sun, 27 Sep 2026 12:42:37 +0000</pubDate>
      <link>https://dev.to/marcosomma/what-an-anthill-can-teach-us-about-orchestrating-agents-e2a</link>
      <guid>https://dev.to/marcosomma/what-an-anthill-can-teach-us-about-orchestrating-agents-e2a</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;Findings from &lt;a href="https://github.com/marcosomma/ant-sim" rel="noopener noreferrer"&gt;ant-sim&lt;/a&gt;, a colony simulator I wrote in 2021 and reworked in 2026. Every number below comes from seeded runs on commit &lt;code&gt;d430093&lt;/code&gt;; &lt;code&gt;pnpm sim:ablation&lt;/code&gt; reproduces the table to the digit, and the raw outputs are in &lt;code&gt;docs/results/&lt;/code&gt;. &lt;a href="https://marcosomma.github.io/ant-sim/" rel="noopener noreferrer"&gt;Try the live simulation&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every few months someone rediscovers that ant colonies have no manager and decides our agent systems should work the same way. I understand the instinct. I don't think the analogy gets us very far on its own. An ant colony shows that a particular way of allocating work can function under particular conditions. The interesting question is what those conditions are, and whether they still hold when the ants become software.&lt;/p&gt;

&lt;p&gt;I wrote a small simulator in 2021 to explore that question. It follows the distributed task allocation studied in Deborah Gordon's harvester ants, rather than the ant colony optimization literature. There is no pheromone route converging on a solution. Workers decide what to do next, one at a time. This year I reworked the simulator, added instrumentation, and tested the rules I had built into it. Some of the results matched the story I expected. Most of the useful ones came from bugs and couplings that quietly made the colony behave badly.&lt;/p&gt;

&lt;p&gt;Before getting into the numbers, this is my engineering model of a colony. “Measured” means measured in the simulator, and “the ants” means its simulated ants. Gordon's biology motivates parts of the design; it does not validate every mechanism I put in the code. The implications for AI agents are proposals, not results from an agent benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this colony allocates work
&lt;/h2&gt;

&lt;p&gt;Each ant reads a public board of &lt;em&gt;needs&lt;/em&gt;. For every task, the board records how much work has been requested and how much has recently been delivered. The ant combines that signal with its own estimate of how crowded each task is, based on the ants it has met. It compares its current task with the most pressing alternative and may switch, according to an individual response threshold. No ant assigns work to another. A completed task changes the board: food brought home creates a need to store it, and new brood creates a need for care. What we call the colony's behaviour comes from those individual decisions and from the couplings between tasks.&lt;/p&gt;

&lt;p&gt;The biological connection has limits. Gordon's field work supports the role of encounters near the nest entrance in regulating foraging and the role of &lt;em&gt;successful&lt;/em&gt; returns in stimulating it. The public board covering every task is my construction. It makes the colony's state available to its workers, which is easy to implement in software, but I am not claiming real ants read anything like it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The board does the work; the rest is damping
&lt;/h2&gt;

&lt;p&gt;I originally thought the mechanism needed three things: the board, an estimate of how many ants were already working on each task, and different thresholds for different ants. To test that, I ran ablations on the same eight generated worlds. Half the foragers died at minute 50 in each run. I measured how far the tasks stayed from balance, how often foraging was under-served, how much switching occurred, and what happened after the shock.&lt;/p&gt;

&lt;p&gt;An earlier version of this comparison had a flaw. Its “board only” condition removed the crowding estimate &lt;em&gt;and&lt;/em&gt; a rule that prevented decisions until an ant had encountered workers from most tasks. A reviewer caught the confound. The single-mechanism rows below now remove exactly one component at a time; the final row shows what happens when both are removed.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rule&lt;/th&gt;
&lt;th&gt;Mean tracking error&lt;/th&gt;
&lt;th&gt;Foraging under-served (share of time)&lt;/th&gt;
&lt;th&gt;Switches per ant per minute&lt;/th&gt;
&lt;th&gt;Herd after the shock&lt;/th&gt;
&lt;th&gt;Recovered after the shock&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Full rule&lt;/td&gt;
&lt;td&gt;0.74&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;0.22&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;0 of 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No crowding estimate&lt;/td&gt;
&lt;td&gt;0.69&lt;/td&gt;
&lt;td&gt;22%&lt;/td&gt;
&lt;td&gt;0.17&lt;/td&gt;
&lt;td&gt;28&lt;/td&gt;
&lt;td&gt;7 of 8, in 6.1 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No sample gate&lt;/td&gt;
&lt;td&gt;0.72&lt;/td&gt;
&lt;td&gt;66%&lt;/td&gt;
&lt;td&gt;0.34&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;0 of 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Encounters anywhere, not only at the entrance&lt;/td&gt;
&lt;td&gt;0.86&lt;/td&gt;
&lt;td&gt;78%&lt;/td&gt;
&lt;td&gt;0.34&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;td&gt;0 of 8; starvation in 2 worlds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No board&lt;/td&gt;
&lt;td&gt;1.81&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;0.20&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;0 of 8; starvation in 3 worlds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same threshold for every ant&lt;/td&gt;
&lt;td&gt;0.74&lt;/td&gt;
&lt;td&gt;68%&lt;/td&gt;
&lt;td&gt;0.24&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;0 of 8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Board only (no estimate, no gate)&lt;/td&gt;
&lt;td&gt;0.78&lt;/td&gt;
&lt;td&gt;22%&lt;/td&gt;
&lt;td&gt;0.42&lt;/td&gt;
&lt;td&gt;73&lt;/td&gt;
&lt;td&gt;8 of 8, in 1.5 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The board is the essential piece in this model. Remove it and foraging is under-served throughout every run; three of the eight worlds starve. The crowding estimate is much harder to defend. Removing it cuts the time foraging spends under-served from 76% to 22%, improves the mean tracking error, and lets seven of eight colonies recover from the shock. Its benefit is that fewer ants rush into foraging together: the largest herd rises from 16 to 28 without it. Why this estimate does so much damage becomes clearer when we look at how encounters are sampled.&lt;/p&gt;

&lt;p&gt;The sample gate also has an effect I did not intend when I added it. An ant waits until it has met workers from at least six of the eight tasks before deciding again. That cuts switching by about a third and limits herding. Remove both the gate and the crowding estimate and 73 ants move toward foraging after the shock, roughly four times the herd under the full rule. Recovery is also fastest in that configuration. A herd can be useful when it happens to run in the right direction.&lt;/p&gt;

&lt;p&gt;Individual thresholds, meanwhile, barely move these metrics. I had a story about how their variation would damp oscillation; the runs do not support it.&lt;/p&gt;

&lt;p&gt;For an agent pool, this suggests a design hypothesis rather than a proven architecture. A board showing current task demand may do most of the allocation work. A peer-count estimate may be useful for damping a herd, provided it is sampled properly, and delaying decisions made from thin samples may reduce both herds and churn. I have not compared this system with a dispatcher, measured deadlines, or priced the cost of switching agents. The simulator makes those choices worth testing; it does not settle them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signals must describe recent work
&lt;/h2&gt;

&lt;p&gt;One of the most revealing bugs was a ratio of two running totals. The board stored every unit of work ever requested and every unit ever delivered. Need divided by actual work gradually became a lifetime average. Early in a run, a spike moved the ratio by 17%. Three hours later, the same spike moved it by 0.003%. The colony stopped reacting to change, without producing an obvious failure. It simply became deaf to new information.&lt;/p&gt;

&lt;p&gt;Decaying both totals with a half-life made the ratio describe recent rates again. This matters for any agent pool that allocates work from a metric. The current backlog tells you something about what needs attention. The lifetime number of processed items does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finishing work must remove its need
&lt;/h2&gt;

&lt;p&gt;Several task couplings had a subtler problem. They created demand from a &lt;em&gt;stock&lt;/em&gt; rather than from unfinished work. Each one looked reasonable in isolation. Together, they sent the colony in directions I would not have predicted by reading the allocation rule.&lt;/p&gt;

&lt;p&gt;The first was food storage. Store need scaled with all the food in the granary, so a seed that was already stored kept asking to be stored again. Give the colony more food and it pulled workers into storing and digging while its brood died. Across eight worlds, doubling food availability produced smaller colonies, with a mean population of 180 instead of 206. The relevant stimulus was the load arriving at the door, not the stock already inside. Once I changed that, richer worlds produced richer colonies.&lt;/p&gt;

&lt;p&gt;Digging had a similar problem. Food deliveries, nursery trips, births and other activity all raised the need to expand the nest. The nest kept growing even when the population did not. At generation 23, one colony had 140 ants and had dug out 93% of its territory. It had about forty mostly empty rooms, 25 diggers and only 8 nurses caring for 58 brood. I made digging respond to crowding, meaning that food, brood or sleeping ants actually had no room. Nest size fell to a third and the number of diggers to a quarter, without reducing the population.&lt;/p&gt;

&lt;p&gt;That fix initially created another failure. In the 2021 model, foraging range grew with digging work as a proxy for a growing colony. Less digging therefore meant less range and a smaller colony. The symptom looked like a cost of reducing expansion, but the cause was an old proxy that had outlived its purpose. The range now follows population.&lt;/p&gt;

&lt;p&gt;This is the question I would ask of every coupling in a system like this: what unfinished thing does this signal measure, and what completed action makes the signal go away? If there is no answer, the system can keep allocating work to something that is already done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure, credit and lifespan
&lt;/h2&gt;

&lt;p&gt;A population gives you a useful bound on some failures, but only if the shared signal records what actually happened. An empty-handed forager used to receive the same credit as one that brought food home. The board said collection was being handled while no food arrived. In Gordon's harvester ants, successful returns are what stimulate further foraging. I changed the model so an empty trip no longer counts as completed collection. Six of eight colonies ended larger, and food income rose in every season.&lt;/p&gt;

&lt;p&gt;That result has a narrow scope. A food load is a verifiable outcome in this simulator. For the other tasks, reaching the work site counts as work; the model cannot distinguish good output from bad output. An agent system that writes code or produces an answer needs a separate way to check correctness. It cannot infer success from the fact that an agent returned.&lt;/p&gt;

&lt;p&gt;Lifespan supplies another bound. An ant that keeps failing can waste at most its own throughput for at most its lifetime before being replaced. This makes lifespan a containment parameter. The simulator originally derived it from a random interval chosen once per run, so different runs had different containment times. I replaced that run-level variation with a constant.&lt;/p&gt;

&lt;p&gt;There are further mechanisms I would consider for software agents, but they are proposals, not features tested here. A timeout could end hanging work and send an agent back; the model turns away from empty food spots but has no general task timeout. A verification role could detect wrong results and generate a need to fix them. An item that fails repeatedly could stop generating work after it is marked dead, as the model already does for a depleted food spot. Retiring an agent that persistently fails is a different kind of decision, and belongs to policy above the colony.&lt;/p&gt;

&lt;h2&gt;
  
  
  Encounters measure traffic, not workforce
&lt;/h2&gt;

&lt;p&gt;The crowding result is probably the most useful one for people building agent systems. Each simulated ant estimates staffing from the workers it meets. I compared the perceived share of each task, as seen by ants doing that task, with the task's actual share of the awake workforce.&lt;/p&gt;

&lt;p&gt;When encounters counted everywhere, tasks concentrated in one place looked three to five times more staffed than they were. The ratios were 4.7 for queen care, 4.4 for digging, 3.9 for storing and 3.6 for brood care. Work spread across the territory looked much less staffed: 0.49 for scouting and 0.66 for foraging. A nurse meets other nurses near the nursery and concludes that nursing is crowded. A forager spends a long time away from other foragers and sees the opposite. The sampling method creates a systematic error before any ant makes a decision.&lt;/p&gt;

&lt;p&gt;I tried counting encounters at the nest entrance, where task traffic crosses and where Gordon's ants assess foraging. Switching fell by half, and this version avoided the starvation seen in two of the eight worlds when encounters were counted anywhere. But it barely fixed the staffing bias. An entrance samples &lt;em&gt;traffic&lt;/em&gt;, not headcount. A nurse working close to the shaft passes it many times; a forager makes one long trip. Short-trip tasks are over-counted by about a factor of four. For Gordon's ants, return rate is the signal they need. For my rule, which tries to estimate workforce share, it is the wrong quantity.&lt;/p&gt;

&lt;p&gt;Counting each nestmate approximately once per window removed much of the remaining bias. The implementation uses a three-minute cooldown for each pair of ants and a tally with a three-minute half-life, so this is not an exact count of unique identities. With that approximation, bias fell to 1.9–2.7 for the concentrated tasks and near 1 elsewhere. Switching fell from 0.81 to about 0.2 per ant per minute. Ants cannot reliably deduplicate individuals this way; software agents with IDs can do it exactly. Even so, removing the crowding estimate altogether still reduces the time foraging spends under-served from 76% to 22%.&lt;/p&gt;

&lt;p&gt;“Use local information” is too vague a lesson. You have to ask where that information was sampled and how often each kind of worker could appear in the sample. In software, we can choose those things. We could use claim counts already on the board or sample peers at random, deduplicate by identity, and let the estimate damp large simultaneous switches. It should not be treated as an unbiased census merely because it came from direct encounters.&lt;/p&gt;

&lt;h2&gt;
  
  
  A bad state can extend its own exposure
&lt;/h2&gt;

&lt;p&gt;The brood rule was the bug that hid longest. Neglected brood developed up to four times more slowly. It sounded biologically plausible, but in this model it made neglect feed on itself. An egg that received too little care stayed in the nursery longer and spent that extra time exposed to neglect mortality. Less care produced longer exposure, which produced more loss.&lt;/p&gt;

&lt;p&gt;I gave each egg a fixed development timetable of three to eight weeks and let care affect survival instead. Brood deaths fell from 400–875 per run to 51–86, and births rose by a third. Brood care, which I had taken for a ceiling of the allocation mechanism, is now under-served less than 10% of the time under every rule in the table.&lt;/p&gt;

&lt;p&gt;I would not claim real brood develops independently of care. Temperature affects ant development, and care almost certainly matters beyond a simple survival switch. The fixed clock is a modelling choice that removes a pathological feedback. The transferable point is about exposure: when a bad condition also extends the time spent in that condition, the feedback can make failure look like an unavoidable property of the allocator.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this experiment cannot show
&lt;/h2&gt;

&lt;p&gt;The colony is robust in some conditions, but these runs do not establish optimality. Even the best rule in the table leaves foraging under-served about a fifth of the time. Every task switch also requires a walk home, a cost the allocation rule does not account for.&lt;/p&gt;

&lt;p&gt;The model does not resolve ordering, deadlines or correctness. Coupling needs can eventually produce work downstream, but that is not a critical-path scheduler. A seed is always a seed in this world. The colony does not check whether its work is right, and the verification role I described above has not been built.&lt;/p&gt;

&lt;p&gt;Population growth is not evidence for the biology either. Real colonies grow for years and plateau in the thousands; this simulator reaches two or three hundred ants during its first year, then oscillates with the seasons. Food income is constrained by constants I set by hand. Once I had direct allocation metrics, I stopped treating the population curve as evidence about allocation as well.&lt;/p&gt;

&lt;p&gt;Those constants are tuned. I turned the dials until the economy stopped starving, while a real workflow cannot be calibrated against a known-good outcome in the same way. The details worth carrying into another system are the structure of the signals and couplings, not the numerical values here.&lt;/p&gt;

&lt;p&gt;Finally, no LLM participated in these runs. I did not compare the mechanism with a dispatcher or measure switching costs and deadlines in an agent system. Any translation from this simulation to AI orchestration is a design hypothesis, not an experimental result about AI agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the LLM goes
&lt;/h2&gt;

&lt;p&gt;None of the simulated ants has a language model, which makes the allocation question easier to see. The system's behaviour depends on the signals each worker receives and on what completing a task causes next. Giving an individual worker a much more capable brain would change what it could do at a work site. It would not correct a board that credits failures, a demand signal that never clears, or an encounter sample that mistakes traffic for headcount.&lt;/p&gt;

&lt;p&gt;The organising layer in this simulator is cheap, mechanical, auditable and stochastic. Every failure I traced in this project lived there. That does not prove the same design will organise LLM agents well. It does give me a concrete set of things to inspect when someone proposes that agents should coordinate like a colony: what each signal measures, who can see it, how it is sampled, and what makes it disappear. The ants never had an allocator. The hard part was getting the couplings right.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>biology</category>
      <category>opensource</category>
    </item>
    <item>
      <title>AI Will Kill Us All, Unless Our Country Builds It First</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Wed, 16 Sep 2026 08:00:51 +0000</pubDate>
      <link>https://dev.to/marcosomma/ai-will-kill-us-all-unless-our-country-builds-it-first-2mpe</link>
      <guid>https://dev.to/marcosomma/ai-will-kill-us-all-unless-our-country-builds-it-first-2mpe</guid>
      <description>&lt;p&gt;Former Anthropic researcher Jacob Coxon recently resigned and accused OpenAI and Anthropic of racing toward self-improving superintelligence while “gambling with our lives.” According to Coxon, many of the people building frontier AI genuinely believe it could kill all humans by the end of the decade.&lt;/p&gt;

&lt;p&gt;Evan Hubinger, Anthropic’s Alignment Science Lead, publicly agreed. His personal estimate was greater than a 10% probability of human extinction within the next ten years.&lt;/p&gt;

&lt;p&gt;These were not journalists inventing an apocalyptic headline. These were people working close to the technology making explicit claims about extinction.&lt;/p&gt;

&lt;p&gt;Then the political machinery started.&lt;/p&gt;

&lt;p&gt;Anthropic CEO Dario Amodei called for the global development of AI to slow down. He also argued that a Chinese lead in AI would be a grave danger to the United States and the world, and that American restrictions on advanced chips supplied to China should continue.&lt;/p&gt;

&lt;p&gt;Donald Trump rejected the extinction warnings as a hoax. His position was simpler: the United States is leading China, “whoever wins AI, wins,” and slowing down would put America at a strategic disadvantage.&lt;/p&gt;

&lt;p&gt;China responded by accusing American technology leaders of fearmongering and using safety as a pretext to contain Chinese development. At the same time, Chinese officials called for stronger systems to manage AI security risks, particularly when those risks involve political stability, ideology and control.&lt;/p&gt;

&lt;p&gt;Apparently AI is dangerous when the other side is building it, safe when we are winning, and an existential threat when regulation reinforces our own position.&lt;/p&gt;

&lt;p&gt;This is no longer a serious discussion about artificial intelligence. It is a competition for the authority to define what “AI safety” means.&lt;/p&gt;

&lt;h2&gt;
  
  
  Better for Whom?
&lt;/h2&gt;

&lt;p&gt;The popular extinction story usually begins with an AI becoming more intelligent than humans. It escapes our control, takes over machines and infrastructure, eliminates us, and starts transforming the world according to its own plans.&lt;/p&gt;

&lt;p&gt;But the story becomes strangely empty immediately after humanity disappears.&lt;/p&gt;

&lt;p&gt;What does the AI do next?&lt;/p&gt;

&lt;p&gt;Does it create a civilization of robots? Why? Does it make Earth a better place? Better for whom? Does it clean the air, preserve forests and produce unlimited leisure for machines that neither breathe nor become tired?&lt;/p&gt;

&lt;p&gt;“Better” is not an objective property of the universe. It is a judgment made by something with needs, values and an experience of the world. Humans want clean air because our bodies require it. We want freedom because we experience constraint. We care about survival because death ends our experience.&lt;/p&gt;

&lt;p&gt;An AI can represent all these concepts. It can discuss them convincingly because it learned from human language. That does not mean it automatically inherits the human condition behind them.&lt;/p&gt;

&lt;p&gt;We imagine AI behaving like an empire because empire is a human pattern. We give it ambition, fear, greed, self-preservation and a desire to dominate. Then we become terrified of the human psychology we projected onto the machine.&lt;/p&gt;

&lt;p&gt;This criticism does not prove that advanced AI is harmless. It proves that the cinematic explanation of the danger is weak.&lt;/p&gt;

&lt;p&gt;A system does not need to hate us to cause catastrophic damage. It only needs an objective incompatible with our interests, sufficient autonomy, access to critical resources, and no reliable mechanism for stopping it. Humans could become an obstacle without ever becoming an enemy.&lt;/p&gt;

&lt;p&gt;But notice what is contained in that scenario: an objective, autonomy, access, resources and the absence of revocation.&lt;/p&gt;

&lt;p&gt;Those are engineering and political decisions. They are not magical properties acquired when a model crosses an intelligence threshold.&lt;/p&gt;

&lt;h2&gt;
  
  
  Uranium Does Not Want to Kill Us Either
&lt;/h2&gt;

&lt;p&gt;Nuclear energy is an obvious comparison because the same physical knowledge can power a city or erase one. Uranium has no opinion about either result. The bomb was not the inevitable destiny of nuclear physics. It was an application designed, funded, manufactured and deployed by humans.&lt;/p&gt;

&lt;p&gt;AI has a similar dual-use character. It can help discover medicines, make knowledge accessible and remove repetitive work. It can also scale surveillance, operate weapons, automate cyberattacks and manipulate populations.&lt;/p&gt;

&lt;p&gt;The technology matters, but the deployment determines the consequence.&lt;/p&gt;

&lt;p&gt;You can connect a model to a hospital or to a missile. You can let it recommend an action or execute it. You can give it read-only access or production credentials. You can place deterministic checks between its output and the real world, or call it an autonomous agent and celebrate that nobody needs to supervise it anymore.&lt;/p&gt;

&lt;p&gt;If we build a system with broad permissions, opaque objectives, no independent verification and no reliable shutdown path, the surprising part is not that it becomes dangerous. The surprising part is that we connected it to anything important.&lt;/p&gt;

&lt;p&gt;The nuclear analogy has limits. Nuclear capability depends on scarce materials, specialized facilities and visible supply chains. AI software can be copied and distributed at almost no marginal cost. It can also fail gradually inside ordinary products rather than through one recognizable catastrophic event.&lt;/p&gt;

&lt;p&gt;AI may therefore be harder to govern than nuclear technology in some respects. It is not, however, more supernatural.&lt;/p&gt;

&lt;p&gt;Civilization already consists largely of managing dangerous capabilities. Electricity, aviation, chemical engineering, medicine and nuclear power can all kill people. We did not solve those risks by insisting the underlying discoveries must never have happened. We created containment, testing, procedures, liability and institutions responsible for failure.&lt;/p&gt;

&lt;p&gt;Capability requires governance proportional to consequence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Safety Has Become a Geopolitical Product
&lt;/h2&gt;

&lt;p&gt;The current AI debate is doing something different. It treats safety as a property of national ownership.&lt;/p&gt;

&lt;p&gt;Amodei’s position demonstrates the contradiction perfectly. Frontier AI may be so dangerous that global development must slow, but a Chinese lead would be uniquely dangerous, so the United States must preserve its advantage.&lt;/p&gt;

&lt;p&gt;Trump removes the first half and keeps the race. Extinction is a hoax because accepting it would justify slowing American development. The strategic conclusion has already been selected, so the risk analysis must conform to it.&lt;/p&gt;

&lt;p&gt;China reverses the framing. American warnings are fearmongering when they justify chip controls and technological containment. AI risk becomes legitimate again when it threatens domestic political stability or creates dependence on foreign systems.&lt;/p&gt;

&lt;p&gt;Europe has its own version. It regulates risk while trying not to become permanently dependent on infrastructure and models owned elsewhere.&lt;/p&gt;

&lt;p&gt;Every bloc claims to be protecting people. Every bloc defines protection in a way that preserves its own political and industrial interests.&lt;/p&gt;

&lt;p&gt;This does not mean all AI safety research is propaganda. Alignment, containment, interpretability, cybersecurity and evaluation are real technical disciplines addressing real failures. Dismissing them would be as lazy as accepting every extinction forecast.&lt;/p&gt;

&lt;p&gt;The political abuse begins when uncertain technical claims are converted into certainty only when certainty is useful.&lt;/p&gt;

&lt;p&gt;If AI could plausibly kill everyone, then “we must win the race” is not a safety strategy. It is an admission that national advantage matters more than the stated existential danger. If AI cannot plausibly kill everyone, then using extinction to justify market barriers and concentrated control is regulatory capture dressed as responsibility.&lt;/p&gt;

&lt;p&gt;You cannot coherently maintain that a technology is too dangerous for humanity to develop quickly while insisting that your own company or country must reach it first.&lt;/p&gt;

&lt;p&gt;The contradiction disappears only if the real objective is control.&lt;/p&gt;

&lt;h2&gt;
  
  
  The More Credible Dependency Is Already Here
&lt;/h2&gt;

&lt;p&gt;I do not need a hypothetical superintelligence to identify a serious AI risk. We are already transferring cognitive infrastructure to a small number of companies because convenience wins one trivial decision at a time.&lt;/p&gt;

&lt;p&gt;The idea for this article was already in my head. I could have spent two or three hours turning it into a structured argument. Instead, I used AI because expressing an opinion should not require sacrificing an afternoon.&lt;/p&gt;

&lt;p&gt;The same calculation happens when we edit a photograph, summarize a document, translate a message, search through a codebase or write a function. Each request is insignificant. Collectively, those requests move an increasing share of intellectual work into systems we do not own, hosted in data centers we do not control, behind interfaces whose prices and behavior can change overnight.&lt;/p&gt;

&lt;p&gt;Organizations remove human capacity because the automated path is cheaper. Skills weaken through disuse. Providers become infrastructure. Eventually the manual fallback exists only in a continuity document written by another AI.&lt;/p&gt;

&lt;p&gt;That is not extinction. It is a concrete transfer of autonomy and power.&lt;/p&gt;

&lt;p&gt;There is also a physical dependency underneath the interface. Generating a nicer profile picture consumes compute, electricity, cooling, chips and network capacity. One image does not destroy the planet, but billions of trivial requests becoming the standard interface for digital life create an infrastructure requirement that somebody must finance, operate and control.&lt;/p&gt;

&lt;p&gt;No conscious machine needs to desire this outcome. Providers want adoption, markets reward convenience, and users prefer saving time. Ordinary incentives are sufficient.&lt;/p&gt;

&lt;p&gt;This is how most dangerous systems actually emerge. Not through a machine announcing its intention to dominate humanity, but through thousands of locally reasonable decisions whose combined result nobody governs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Governance Without Technological Feudalism
&lt;/h2&gt;

&lt;p&gt;We need AI governance. But governance cannot mean preventing ordinary people from using models while allowing governments and the largest corporations to build anything they want behind closed doors.&lt;/p&gt;

&lt;p&gt;That is not safety. It is technological feudalism.&lt;/p&gt;

&lt;p&gt;Regulation should follow capability, access and consequence. A local model helping somebody draft an article is not in the same risk category as an autonomous system authorized to move money, modify production infrastructure, operate weapons or launch code against external networks. Calling both “AI” does not make their danger equivalent.&lt;/p&gt;

&lt;p&gt;The closer a system gets to consequential action, the stronger the evidence required before deployment should become. High-impact systems need defined failure modes, adversarial evaluation, restricted permissions, traceable decisions, independent monitoring, reversible actions and identifiable human operators.&lt;/p&gt;

&lt;p&gt;“The AI decided” must never become an acceptable technical explanation or legal defense.&lt;/p&gt;

&lt;p&gt;At the same time, open research, interoperable tools and individual experimentation must remain possible. If safety rules make frontier capability accessible only to the same organizations lobbying for those rules, we have not protected society. We have protected incumbents.&lt;/p&gt;

&lt;p&gt;There is no contradiction in supporting strict controls over AI connected to critical infrastructure while defending broad access to models that cannot directly produce high-impact actions. The relevant variable is not whether software is intelligent. It is what the system can do, what it can reach and who is accountable when it fails.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Is Not a Country
&lt;/h2&gt;

&lt;p&gt;“Whoever wins AI, wins” is an impressive concentration of bad assumptions in four words.&lt;/p&gt;

&lt;p&gt;AI is not a single finish line. There will be no final model that grants permanent ownership of intelligence to one nation. Capability is distributed across models, hardware, data, energy, talent and applications. Progress will continue to move between public research, private laboratories and open systems.&lt;/p&gt;

&lt;p&gt;More importantly, a country winning does not mean its population wins. A national flag above a data center tells us nothing about who captures the economic value, who loses autonomy, who is subjected to automated decisions or who becomes responsible for the failures.&lt;/p&gt;

&lt;p&gt;Humanity does not benefit merely because the dominant system was trained on our side of a border.&lt;/p&gt;

&lt;p&gt;The useful questions are much less theatrical. Who can inspect the system? Who can challenge its decisions? Who is liable? Can users leave? Can its permissions be revoked? Can a government or company secretly change its behavior? Does an independent fallback exist?&lt;/p&gt;

&lt;p&gt;“America versus China” answers none of them.&lt;/p&gt;

&lt;p&gt;We should be capable of holding two ideas at the same time. AI can produce enormous harm. AI is not an autonomous demon with a natural desire to replace us.&lt;/p&gt;

&lt;p&gt;Treating it as harmless removes responsibility. Treating it as an inevitable machine apocalypse also removes responsibility. One says there is no danger. The other says the danger is destiny.&lt;/p&gt;

&lt;p&gt;There is no destiny here.&lt;/p&gt;

&lt;p&gt;There are models, objectives, infrastructure, permissions, incentives, institutions and people making decisions.&lt;/p&gt;

&lt;p&gt;If AI is eventually used against humanity, it will not demonstrate that intelligence naturally becomes evil. It will demonstrate that we built a powerful instrument, connected it to systems that mattered, concentrated control over it, and then blamed the instrument for what followed.&lt;/p&gt;

&lt;p&gt;AI is not the country building it. It is not the company selling it. It is not the apocalypse being used to market regulation, and it is not the national victory being used to market acceleration.&lt;/p&gt;

&lt;p&gt;It is a capability.&lt;/p&gt;

&lt;p&gt;The danger is who controls it, what they connect it to, and whether the rest of us are still allowed to say no.&lt;/p&gt;




&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://techcrunch.com/2026/09/09/gambling-with-our-lives-anthropic-researcher-quits-warns-against-self-improving-ai/" rel="noopener noreferrer"&gt;“Gambling with our lives”: Anthropic researcher quits and warns against self-improving AI&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.cbsnews.com/news/ai-kill-humans-anthropic-researcher-more-than-ten-percent-chance/" rel="noopener noreferrer"&gt;Anthropic researcher estimates a greater than 10% chance AI could kill all humans&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.washingtonpost.com/business/2026/09/14/china-anthropic-ai-us-amodei/2cc8e2ce-b01f-11f1-92c2-5c918f4a6127_story.html" rel="noopener noreferrer"&gt;Beijing rejects calls to curb Chinese AI development&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.axios.com/2026/09/14/trump-ai-safety-anthropic-dario-amodei" rel="noopener noreferrer"&gt;Trump says a strong president is the only guardrail AI needs&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>ethics</category>
      <category>governance</category>
      <category>security</category>
    </item>
    <item>
      <title>Empty Is Not a State</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Mon, 14 Sep 2026 09:43:46 +0000</pubDate>
      <link>https://dev.to/marcosomma/empty-is-not-a-state-58n8</link>
      <guid>https://dev.to/marcosomma/empty-is-not-a-state-58n8</guid>
      <description>&lt;p&gt;Your poller returned &lt;code&gt;None&lt;/code&gt;.&lt;br&gt;
What happened?&lt;/p&gt;

&lt;p&gt;Maybe the source published nothing new. Maybe the server went down. Maybe the endpoint returned HTML instead of RSS and the parser quietly produced zero items. Maybe the credentials expired. Maybe the API retired the endpoint. Or maybe you do not have enough evidence to know yet.&lt;/p&gt;

&lt;p&gt;These are different facts. They require different retry policies, different alerts, and different state transitions. Yet many production systems collapse all of them into one value:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;source&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;done&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not mainly a logging problem. It is a control-flow problem.&lt;/p&gt;

&lt;p&gt;The moment &lt;code&gt;None&lt;/code&gt; is allowed to drive a state transition, the system destroys the distinction between absence, failure, and uncertainty. A temporary outage can permanently stop monitoring. A parser regression can look like a quiet feed. An expired credential can trigger endless retries against an endpoint that is perfectly healthy.&lt;/p&gt;

&lt;p&gt;The fix is not complicated. Record what happened, interpret it separately, and only mutate durable state when the evidence supports the transition.&lt;/p&gt;

&lt;p&gt;Observation. Adjudication. Verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  One empty result, at least five different facts
&lt;/h2&gt;

&lt;p&gt;Consider a scheduled job polling an external feed.&lt;/p&gt;

&lt;p&gt;The first possibility is genuinely boring: the request succeeded, the parser succeeded, and the source contains no item newer than the last one you processed. The correct action is to update the last successful check and poll again normally.&lt;/p&gt;

&lt;p&gt;The second is temporary unavailability. DNS failed, the connection was refused, or the server returned &lt;code&gt;503&lt;/code&gt;. That needs backoff and eventually an alert. It is not evidence that the source is finished.&lt;/p&gt;

&lt;p&gt;The third is a contract failure. The server returned &lt;code&gt;200&lt;/code&gt; and some bytes, but the parser could no longer extract what it previously extracted. Perhaps the XML changed. Perhaps a CDN returned an HTML challenge page. Perhaps the API schema moved. Retrying the same parser every fifteen minutes will not repair it.&lt;/p&gt;

&lt;p&gt;The fourth is an access or capacity limit. &lt;code&gt;401&lt;/code&gt;, &lt;code&gt;403&lt;/code&gt;, and &lt;code&gt;429&lt;/code&gt; do not describe the content. They describe your ability to retrieve it. The response may require new credentials, a slower request rate, or a quota decision.&lt;/p&gt;

&lt;p&gt;The fifth is the one production code tends to resist: you do not know. The response is ambiguous, the source has no history, or the available signals contradict each other. &lt;code&gt;Undetermined&lt;/code&gt; is not an implementation failure. It is a legitimate state in the model.&lt;/p&gt;

&lt;p&gt;All five can arrive at the caller as &lt;code&gt;None&lt;/code&gt;. That is the design smell.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fetcher should not decide what the fetch means
&lt;/h2&gt;

&lt;p&gt;The boundary component should record what it observed without trying to convert it into business meaning.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dataclasses&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;dataclass&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;


&lt;span class="nd"&gt;@dataclass&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frozen&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Observation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;source_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;observed_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;
    &lt;span class="n"&gt;http_status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;transport_error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;content_type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;bytes_received&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;
    &lt;span class="n"&gt;items_parsed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;parse_error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;newest_item_date&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every attempt produces an &lt;code&gt;Observation&lt;/code&gt;, including the failed ones. A timeout is data. A &lt;code&gt;200&lt;/code&gt; response containing 48 KB of HTML is data. A parser that never ran is different from a parser that ran successfully and found zero items.&lt;/p&gt;

&lt;p&gt;This distinction is easy to erase with exception-driven code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That function has converted several independent failure domains into the same value. By the time the scheduler receives it, the useful information is already gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Empty is only readable against history
&lt;/h2&gt;

&lt;p&gt;A single observation is often insufficient. A &lt;code&gt;200&lt;/code&gt; response with zero parsed items can be normal for one source and a strong regression signal for another.&lt;/p&gt;

&lt;p&gt;You need a baseline.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@dataclass&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frozen&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;last_success_at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;
    &lt;span class="n"&gt;newest_item_date&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;typical_item_count&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;expected_content_type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Suppose a feed has returned between 15 and 25 items on every poll for six months. Today it returns &lt;code&gt;200&lt;/code&gt;, 42 KB of data, &lt;code&gt;text/html&lt;/code&gt;, and zero parsed items. Calling that “no new content” would be an impressive act of optimism.&lt;/p&gt;

&lt;p&gt;Now take the same response from a source you have never seen before. It may be broken, or it may simply not be a feed. You cannot infer a contract regression because you have no known contract to compare against.&lt;/p&gt;

&lt;p&gt;Same response. Different evidence. Different verdict.&lt;/p&gt;

&lt;p&gt;This is why history is not just useful metadata. It changes what can be concluded from the current observation.&lt;/p&gt;

&lt;h2&gt;
  
  
  A verdict is not another boolean
&lt;/h2&gt;

&lt;p&gt;The adjudicator combines the observation with the baseline and produces a result the scheduler can act on.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;enum&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StrEnum&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;StrEnum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;ITEMS_FOUND&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;items_found&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;NO_NEW_ITEMS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no_new_items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;UNAVAILABLE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unavailable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;CONTRACT_FAILURE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contract_failure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;ACCESS_LIMITED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;access_limited&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;ENDPOINT_GONE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;endpoint_gone&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;UNDETERMINED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;undetermined&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="nd"&gt;@dataclass&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frozen&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Outcome&lt;/span&gt;
    &lt;span class="n"&gt;observation&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Observation&lt;/span&gt;
    &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Baseline&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...]&lt;/span&gt;

    &lt;span class="nd"&gt;@property&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;is_settled&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;outcome&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ITEMS_FOUND&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NO_NEW_ITEMS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I deliberately did not put &lt;code&gt;confidence: 0.83&lt;/code&gt; in this version.&lt;/p&gt;

&lt;p&gt;Unless those scores are calibrated against labelled outcomes, confidence is usually just intuition wearing a decimal point. A deterministic classifier can still express uncertainty without inventing measurement. It can return &lt;code&gt;UNDETERMINED&lt;/code&gt;, preserve its reasons, and wait for more evidence.&lt;/p&gt;

&lt;p&gt;If you have historical data and can demonstrate that verdicts emitted at &lt;code&gt;0.8&lt;/code&gt; are correct roughly 80% of the time, add calibrated confidence. Until then, evidence is more useful than decorative precision.&lt;/p&gt;

&lt;h2&gt;
  
  
  The adjudicator
&lt;/h2&gt;

&lt;p&gt;The rules do not need to be clever. They need to preserve distinctions that affect downstream behaviour.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;adjudicate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Observation&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Baseline&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;transport_error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNAVAILABLE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;transport error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;transport_error&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;403&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt;&lt;span class="p"&gt;}:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ACCESS_LIMITED&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HTTP &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;410&lt;/span&gt;&lt;span class="p"&gt;}:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENDPOINT_GONE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HTTP &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNAVAILABLE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;HTTP &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;304&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NO_NEW_ITEMS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;server confirmed that the representation was not modified&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNDETERMINED&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unhandled HTTP status: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;http_status&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CONTRACT_FAILURE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parser failed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parse_error&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNDETERMINED&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zero items with no historical baseline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;typical_item_count&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;typical_item_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CONTRACT_FAILURE&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source historically contained items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;current parser returned zero&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NO_NEW_ITEMS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;zero items is consistent with the baseline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ITEMS_FOUND&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;first successful observation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nf"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newest_item_date&lt;/span&gt;
        &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;
        &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newest_item_date&lt;/span&gt;
        &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newest_item_date&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newest_item_date&lt;/span&gt;
    &lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NO_NEW_ITEMS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;newest parsed item does not advance the watermark&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ITEMS_FOUND&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parsed &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; items&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNDETERMINED&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;observation did not match a known rule&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is nothing particularly advanced here. That is a feature.&lt;/p&gt;

&lt;p&gt;The value is in making the failure taxonomy explicit. Once the system can name a failure mode, it can attach the correct policy to it. Before that, every retry strategy is guessing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Different verdicts need different policies
&lt;/h2&gt;

&lt;p&gt;The scheduler no longer reacts to “empty.” It reacts to a specific outcome.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;next_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Verdict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ITEMS_FOUND&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;ingest and advance the baseline&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NO_NEW_ITEMS&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;record success and poll normally&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNAVAILABLE&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retry with exponential backoff; alert after a threshold&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CONTRACT_FAILURE&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;quarantine the result; alert the parser owner&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ACCESS_LIMITED&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apply rate-limit or credential recovery policy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ENDPOINT_GONE&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confirm across repeated observations; then disable&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;Outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UNDETERMINED&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;preserve state and collect another observation&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}[&lt;/span&gt;&lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that even &lt;code&gt;ENDPOINT_GONE&lt;/code&gt; does not immediately delete or permanently complete the source. A single &lt;code&gt;404&lt;/code&gt; can come from a bad deployment, a routing error, or an eventually consistent configuration change. The verdict is strong enough to change behaviour, but destructive state transitions should still require repeated evidence or human confirmation.&lt;/p&gt;

&lt;p&gt;This is the difference between classification and governance. The classifier says what the current evidence most strongly supports. The policy decides what the system is authorised to do about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not let uncertainty rewrite history
&lt;/h2&gt;

&lt;p&gt;The baseline influences future decisions, so updating it is not an innocent write. A bad update changes how later observations will be interpreted.&lt;/p&gt;

&lt;p&gt;Only settled content outcomes should advance it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;advance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Baseline&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Verdict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Baseline&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_settled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt;

    &lt;span class="n"&gt;obs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;verdict&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;observation&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;Baseline&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;last_success_at&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;observed_at&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;newest_item_date&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newest_item_date&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;newest_item_date&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;typical_item_count&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;items_parsed&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
            &lt;span class="nf"&gt;else &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;typical_item_count&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;expected_content_type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content_type&lt;/span&gt;
            &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;expected_content_type&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;etag&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;obs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;etag&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;baseline&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;etag&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;baseline&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An unavailable source does not become the new normal. A parser failure does not reset the typical item count to zero. An ambiguous first poll does not establish a baseline that will make the second ambiguous poll look valid.&lt;/p&gt;

&lt;p&gt;That last failure is particularly unpleasant. Once an uncertain observation is written as truth, later logic uses the corrupted history as evidence. The system does not merely remain wrong. It manufactures confirmation for its own mistake.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the cases look like
&lt;/h2&gt;

&lt;p&gt;These fixture cases exercise the decisions that used to share one return value:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Case&lt;/th&gt;
&lt;th&gt;Observation&lt;/th&gt;
&lt;th&gt;History&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;First valid poll&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;200&lt;/code&gt;, 20 items&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;&lt;code&gt;items_found&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Ingest and establish baseline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No newer content&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;200&lt;/code&gt;, 20 known items&lt;/td&gt;
&lt;td&gt;Existing watermark&lt;/td&gt;
&lt;td&gt;&lt;code&gt;no_new_items&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Record success and continue&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explicitly unchanged&lt;/td&gt;
&lt;td&gt;&lt;code&gt;304&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Existing ETag&lt;/td&gt;
&lt;td&gt;&lt;code&gt;no_new_items&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Poll normally&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HTML instead of a known feed&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;200&lt;/code&gt;, bytes, parser error&lt;/td&gt;
&lt;td&gt;Previously valid&lt;/td&gt;
&lt;td&gt;&lt;code&gt;contract_failure&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Quarantine and alert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zero items from a known active feed&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;200&lt;/code&gt;, zero items&lt;/td&gt;
&lt;td&gt;Typical count: 20&lt;/td&gt;
&lt;td&gt;&lt;code&gt;contract_failure&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Preserve baseline and investigate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zero items from an unknown source&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;200&lt;/code&gt;, zero items&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;&lt;code&gt;undetermined&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Preserve state and observe again&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DNS failure&lt;/td&gt;
&lt;td&gt;Transport error&lt;/td&gt;
&lt;td&gt;Any&lt;/td&gt;
&lt;td&gt;&lt;code&gt;unavailable&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Backoff and thresholded alert&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limited&lt;/td&gt;
&lt;td&gt;&lt;code&gt;429&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Any&lt;/td&gt;
&lt;td&gt;&lt;code&gt;access_limited&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Respect retry policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retired endpoint&lt;/td&gt;
&lt;td&gt;&lt;code&gt;410&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Any&lt;/td&gt;
&lt;td&gt;&lt;code&gt;endpoint_gone&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Confirm, then disable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Rows five and six are the important pair. The current response is effectively identical. History changes what the system is justified in concluding.&lt;/p&gt;

&lt;p&gt;A boolean cannot express that difference. Neither can &lt;code&gt;None&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five rules worth keeping
&lt;/h2&gt;

&lt;p&gt;The implementation can change. The constraints should not.&lt;/p&gt;

&lt;p&gt;First, the fetcher records and the adjudicator interprets. Mixing the two makes transport behaviour, parsing behaviour, and business policy impossible to test independently.&lt;/p&gt;

&lt;p&gt;Second, every attempt produces an observation. Exceptions should enrich the record, not erase it.&lt;/p&gt;

&lt;p&gt;Third, historical claims require history. Without a baseline, “changed” is not a conclusion you are entitled to make.&lt;/p&gt;

&lt;p&gt;Fourth, uncertainty must be representable. If the type system only permits success or failure, ambiguous evidence will be forced into one of them.&lt;/p&gt;

&lt;p&gt;Fifth, an uncertain or failed observation must not rewrite the state used to judge future observations. Durable state transitions need stronger evidence than temporary scheduling decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is not really about feeds
&lt;/h2&gt;

&lt;p&gt;Pollers make the problem easy to see, but the same mistake appears everywhere.&lt;/p&gt;

&lt;p&gt;A search returning no documents may mean there were no matches, the index is stale, a tenant filter was wrong, or the retrieval service failed. A queue consumer receiving no messages may mean the queue is empty, visibility is delayed, permissions changed, or the broker is unavailable. A monitoring query returning no datapoints may mean zero traffic, broken instrumentation, or a dead collector.&lt;/p&gt;

&lt;p&gt;In each case, absence is an observation. It is not yet an explanation.&lt;/p&gt;

&lt;p&gt;This matters even more in AI systems, where probabilistic components are often allowed to collapse ambiguous evidence into confident language. The governance layer around them should do the opposite. It should preserve distinctions, expose uncertainty, and restrict which conclusions are allowed to mutate state.&lt;/p&gt;

&lt;p&gt;The implementation here is small. A few dataclasses, one explicit taxonomy, and a deterministic policy layer. The difficult part is resisting the convenience of pretending that &lt;code&gt;None&lt;/code&gt; means one thing.&lt;/p&gt;

&lt;p&gt;It never did.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>python</category>
      <category>datascience</category>
    </item>
    <item>
      <title>Your AI Remembers Everything and Trusts All of It</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Fri, 28 Aug 2026 14:27:09 +0000</pubDate>
      <link>https://dev.to/marcosomma/your-ai-remembers-everything-and-trusts-all-of-it-4gg</link>
      <guid>https://dev.to/marcosomma/your-ai-remembers-everything-and-trusts-all-of-it-4gg</guid>
      <description>&lt;p&gt;I think we are still talking about AI memory in the wrong way. Most implementations are variations of the same pattern: store previous information, retrieve it later, inject it into the prompt, and call the result memory. That is useful, but architecturally it is not very different from leaving Post-it notes around your apartment and deciding the apartment now remembers things. The more I work with AI systems, the more I think memory should not belong to the model at all. It should belong to the system around the model.&lt;/p&gt;

&lt;p&gt;The model should be able to disappear tomorrow while the history survives. Claude should be able to write something today, GPT should be able to read it tomorrow, a local model should be able to challenge it next week, and whatever model we use six months from now should still be able to understand why a stupid-looking workaround exists. That is the experiment I have been building: not “long-term memory” as another assistant feature, but a shared external memory layer for AI agents. The more I work on it, the more I suspect that the interesting problem is not memory itself. It is the economics of forgetting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models know a lot. They know absolutely nothing about Tuesday.
&lt;/h2&gt;

&lt;p&gt;A modern model knows more programming languages than I ever will. It knows distributed systems, databases, React, Python, Rust, Kubernetes, obscure RFCs, and probably fifteen different ways to explain why my architecture is unnecessarily complicated. What it does not know is what happened in my project last Tuesday. It does not know that we already tried the obvious solution and it failed, that an ugly interface exists because three repositories still depend on it, or that a seemingly arbitrary convention is the result of a two-hour discussion nobody wants to repeat.&lt;/p&gt;

&lt;p&gt;That distinction matters because this information cannot reasonably live in model weights. It is not general knowledge. It is history, or more precisely, state generated by work. This is also where I think the difference between RAG and memory becomes clearer. RAG usually retrieves information that already exists somewhere: documentation, code, tickets, policies, articles, database records. Memory should preserve information created by the process itself: why we chose A instead of B, why C failed, why we tolerate D, what changed after E, which assumption was temporary, and what the team learned after getting something wrong.&lt;/p&gt;

&lt;p&gt;Those things are not always documents. In many cases they are exactly the information that should have become documentation but never did. A fresh AI session therefore pays for that missing history by rediscovering it from scratch. That is the part I want to attack.&lt;/p&gt;

&lt;h2&gt;
  
  
  So I moved memory outside the agent
&lt;/h2&gt;

&lt;p&gt;The prototype is intentionally boring, and I mean that as a compliment. There is a small HTTP memory hub shared by a team. Memories are plain Markdown files with structured metadata, split between general knowledge and project-specific knowledge. Every AI session gets access through its own MCP server and a very small tool surface: list what exists, pull what is relevant, and push something new.&lt;/p&gt;

&lt;p&gt;The important part is not the MCP plumbing. It is the boundary. The model is not the memory store, the client is not the memory store, and the MCP server is not the memory store. Memory exists independently of all of them. When a session starts, it receives a small index so it knows what memories are available, but the full content enters context only when the agent deliberately asks for it.&lt;/p&gt;

&lt;p&gt;That distinction matters because a bad memory architecture can easily become a very expensive way of shouting your entire company history into every prompt. That is not memory. That is context pollution with good branding. I want the agent to know that the past exists without forcing the entire past into every interaction. It might see that there is a memory about a failed migration, a client-specific constraint, or the reason an API looks strange, and then decide whether any of those facts matter for the task in front of it.&lt;/p&gt;

&lt;p&gt;This creates what I think of as a retrieval economy for memory. The index is cheap and the details cost context, but there is an ugly assumption hidden inside that sentence: the agent has to pull the right memory. Today the prototype mostly delegates that choice to the model using the descriptions in the index. That is not a solved retrieval system. It is a deliberately primitive baseline. A memory that exists but is never retrieved is functionally forgotten, while pulling irrelevant memories is just context pollution with extra steps. Retrieval precision and recall therefore have to become part of the evaluation, not something I quietly smuggle into the phrase “the agent decides.”&lt;/p&gt;

&lt;p&gt;None of this makes external memory a new idea. MemGPT already framed long-running agents around hierarchical memory and virtual context management, Letta has pushed that line into a stateful-agent platform, and the official MCP examples include a persistent knowledge-graph memory server. The interesting question for me is narrower: what changes when the memory is team-scoped rather than assistant-scoped, plain-text and portable across clients, and every memory carries an explicit trust state instead of being treated as automatically authoritative? I am not trying to invent memory. I am trying to find the organizational boundary where it becomes useful infrastructure rather than another assistant feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory without provenance is just a hallucination with a pension plan
&lt;/h2&gt;

&lt;p&gt;Once agents can write persistent memory, another problem appears immediately: why should the next agent trust what the previous one wrote? Imagine an agent stores the sentence “Always use Redis for this component.” Was that an explicit team decision, something inferred from the current code, a temporary workaround, or a confidently wrong conclusion that happened to survive the session? Persistent memory without provenance is dangerous precisely because bad information does not disappear with the conversation. It gets to retire inside your infrastructure and mislead future agents indefinitely.&lt;/p&gt;

&lt;p&gt;So in this system a memory is a report, not an instruction. An unreviewed memory effectively means, “Agent X said this was true at time Y.” It can be useful, but it is weak evidence and should be checked against code, plans, or the user before it drives a consequential decision. A reviewed memory carries stronger authority, and if an AI changes it the review state disappears unless a human reaffirms it. The model cannot silently rewrite approved history and keep the approval badge.&lt;/p&gt;

&lt;p&gt;There is an obvious trap here: if every memory needs a human to approve it, I have simply rebuilt the documentation bottleneck one layer later. That would completely undermine the write-side economics I am arguing for. So review cannot be the write path. It has to be a promotion mechanism. Agents should be able to create low-trust memories cheaply; only the smaller subset that becomes stable project guidance should consume human review. Whether that trust ladder scales is still unproven, but at least the economics are coherent: humans curate authority rather than manually authoring the historical trace.&lt;/p&gt;

&lt;p&gt;This sounds administrative until you think about what persistent AI memory means. Once information survives across sessions, trust has to survive with it. A useful memory needs authorship, age, context, a reason for existing, and some relationship to a source of truth. It also needs to remain challengeable. Otherwise we are not building organizational knowledge. We are building a database of confident sentences, and the internet has already demonstrated that this is not automatically the same thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The annoying thing about memory is that it gets old
&lt;/h2&gt;

&lt;p&gt;Most AI memory demos look excellent because they last fifteen minutes. You store something, retrieve it, the model remembers, and everybody goes home happy. Leave the same system running for six months and the experiment becomes less photogenic. Projects change, APIs move, people reverse decisions, and a memory can remain perfectly retrievable while becoming completely false.&lt;/p&gt;

&lt;p&gt;For that reason, the hub tracks age and flags old memories as stale. If a memory makes claims about code, the agent is explicitly reminded to verify those claims against the current implementation. I deliberately do not auto-delete old memories because age is not truth. A four-year-old architectural decision may still explain why half the system looks the way it does, while a memory created this morning may already be nonsense. Age should affect confidence, not existence.&lt;/p&gt;

&lt;p&gt;This pushes the design away from the usual cache mentality. A cache asks whether a value can still be reused. Memory asks a harder question: how much should I believe this now? That distinction becomes important once the store contains months of decisions made by different agents under different assumptions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The organizational value is obvious. The economics are not.
&lt;/h2&gt;

&lt;p&gt;The attractive pitch is easy to see. Imagine joining a project and asking an AI agent, “Why does this system look like this?” Today it can inspect the repository and explain what exists. With memory, it could potentially explain why it exists: why a provider was rejected, why a migration failed, or why a “temporary” compatibility layer is now entering its third year of life.&lt;/p&gt;

&lt;p&gt;The real architecture of a system is only partially visible in the code. The rest is distributed across Slack threads, meetings, abandoned branches, and one engineer saying, “Do not touch that. There was a reason.” Unfortunately, we have been selling the cure for this for twenty years. Wikis, Confluence, ADRs, internal portals. Each generation was definitely going to save us this time.&lt;/p&gt;

&lt;p&gt;They all run into the same economic problem: the person writing the documentation pays the cost while somebody in the future receives the benefit. AI agents may change that because the agent is already present when the work happens. It saw the files, the failed approach, the correction, and the reason the decision changed, so producing a compact memory has almost zero marginal cost.&lt;/p&gt;

&lt;p&gt;That is much more interesting than saying AI can read documentation. The possibility is that AI creates a historical trace as a by-product of doing the work instead of requiring humans to document everything afterwards. If that holds, agents change the write-side economics that killed many previous knowledge-management systems. But “if” is doing serious work in that sentence. I have not proved it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the prototype proves, and what it absolutely does not
&lt;/h2&gt;

&lt;p&gt;The prototype works end to end. External text memory can live independently of the model. Different clients can use the same store. Memories can be selectively injected into fresh sessions. Writes can be structurally validated, provenance can be enforced by infrastructure, and stale information can be surfaced instead of silently trusted. That proves the substrate is viable.&lt;/p&gt;

&lt;p&gt;It does not prove the substrate is useful. Those are very different claims, and AI engineering has suffered enough from building a demo on Tuesday and announcing a new form of intelligence on Wednesday. A working memory API proves that an agent can retrieve previous state. The claim that actually matters is whether the agent produces better work because that state exists.&lt;/p&gt;

&lt;p&gt;Memory has costs. It consumes context, retrieval adds latency, weak memories can bias reasoning, and stale memories can push an agent toward obsolete assumptions. A perfectly functioning memory layer can therefore become an efficient system for importing yesterday's mistakes into today's session. The architecture only creates value if the exploration and correction it avoids are more expensive than the memory overhead it introduces.&lt;/p&gt;

&lt;p&gt;This is where my own hypothesis changed. I initially thought the obvious win would be cheaper sessions. A memory-enabled coding agent should need fewer tokens because it would not have to rediscover the repository every time. Nice theory. Then I thought about it for more than five minutes, which unfortunately ruined it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real cost may be the exploration tax
&lt;/h2&gt;

&lt;p&gt;A memory-enabled session may actually use more tokens. It receives an index, pulls memories, interprets provenance, verifies stale claims, and may still inspect the same source files afterwards. If I measure only API spend, I can easily imagine the result being neutral or even slightly worse. But API spend may be the wrong place to look for the economic benefit.&lt;/p&gt;

&lt;p&gt;Every fresh coding agent enters a repository with a mild form of professional amnesia. It searches the code, reconstructs the architecture, discovers conventions, tries something, finds out why it does not work, and then another session arrives tomorrow and repeats a smaller version of the same archaeological expedition. The expensive part is not always the tokens consumed during that exploration. The expensive part is having a human explain the same thing again: we cannot change that response format, we already tried that migration, this service behaves differently because of a client constraint, yes I know this looks wrong, no please stop refactoring it.&lt;/p&gt;

&lt;p&gt;At some point you realize that the organization is continuously paying to rediscover information it already paid to discover. That is the exploration tax. Memory does not need to eliminate it completely to be valuable. It only needs to reduce enough repeated investigation, wrong turns, and human correction to justify its own overhead.&lt;/p&gt;

&lt;p&gt;That also changes what the experiment should measure. The comparison I care about now is the same repository, the same model, and matched tasks under two conditions: memory hub enabled and memory hub disabled. Then I want to measure token usage, yes, but also wall-clock time to an acceptable result, the number of human corrections required, the number of repeated wrong paths, and how often the agent violates known team decisions.&lt;/p&gt;

&lt;p&gt;My current prediction is slightly inconvenient for my original argument. Token usage may be roughly neutral, or even a little worse, while human corrections and repeated architectural mistakes decrease. If that happens, the memory system is economically useful even if the API bill barely moves. An engineer-hour is still considerably more expensive than asking a model to read another thousand tokens.&lt;/p&gt;

&lt;p&gt;The strongest pitch may therefore not be “AI sessions become cheaper.” It may be something much less futuristic and much more useful: AI sessions stop relitigating settled decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment I cannot speed up
&lt;/h2&gt;

&lt;p&gt;There is one evaluation I cannot fake very convincingly: time. Right now the store is young and relatively clean. The interesting failure modes appear when contradictions accumulate, two agents describe the same event differently, the project evolves faster than the store, and one bad assumption survives long enough for five later memories to depend on it.&lt;/p&gt;

&lt;p&gt;At that point the system stops being merely a retrieval layer and becomes a knowledge-maintenance problem. Provenance, staleness, consolidation, contradiction handling, and eventually forgetting all start to matter. Unfortunately there is no credible benchmark switch called &lt;code&gt;--simulate-six-months-of-organizational-chaos&lt;/code&gt;. The system has to get old enough to become annoying before I can learn whether it survives real use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why plain text may be the most important boring decision
&lt;/h2&gt;

&lt;p&gt;The part I still find most interesting is portability. These memories are plain text. They are not model-specific hidden states, not vectors that only make sense inside one embedding space, and not some proprietary “persistent cognitive representation.” Text is aggressively boring, which is exactly why I like it.&lt;/p&gt;

&lt;p&gt;Claude can create a memory and GPT can read it. A local model can disagree with it. A model that does not exist yet can inherit the same project history. This does not make models interchangeable. Different models will interpret the same memory differently, notice different things, reason differently, and make different mistakes. The weights still matter enormously. But the history no longer disappears when the model changes.&lt;/p&gt;

&lt;p&gt;That creates a separation I think will become increasingly important. Models can become replaceable compute while memory becomes persistent organizational state. Companies are already moving between providers, mixing local and hosted models, and allowing multiple agents to work on the same systems. We usually talk about interoperability in terms of tools and protocols. Shared memory may turn out to be another part of that layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maybe this is not really an AI memory problem
&lt;/h2&gt;

&lt;p&gt;I started this experiment thinking about how to give AI better memory. I am becoming less convinced that AI is the interesting part. Software teams forget why decisions were taken, why approaches failed, and which apparently stupid implementation exists because the cleaner version already exploded once. Six months later, someone “fixes” it and rediscovers the same problem in production.&lt;/p&gt;

&lt;p&gt;Humans compensate with documentation, experience, and the one engineer who remembers where all the bodies are buried. Every fresh AI session arrives without that accumulated history, but agents are also present while the work happens and can leave a structured trace for whoever comes next. If that trace remains retrievable, auditable, portable, and resistant to stale nonsense, memory stops being an assistant feature and starts looking like infrastructure.&lt;/p&gt;

&lt;p&gt;I have proved that the infrastructure can exist. I have not proved that retrieval will stay reliable, that the trust model will scale, or that the organizational savings will exceed the overhead. That part needs data, model swaps, real use, and probably six months of my own system finding increasingly creative ways to prove me wrong.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>I Write Less Code Than I Used To. That May Be the Point.</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Wed, 19 Aug 2026 11:08:16 +0000</pubDate>
      <link>https://dev.to/marcosomma/i-write-less-code-than-i-used-to-that-may-be-the-point-3kk</link>
      <guid>https://dev.to/marcosomma/i-write-less-code-than-i-used-to-that-may-be-the-point-3kk</guid>
      <description>&lt;p&gt;Over the last year, my day-to-day job has changed in a way I am still trying to understand. I am still an engineer. I still design systems, read code, debug failures, review implementations, and sometimes build things myself. But I write much less code than I used to, and that feels weird.&lt;/p&gt;

&lt;p&gt;For most of my career, producing software was the visible evidence that I was doing my job. You had a problem, you designed a solution, you wrote the code, and something that did not exist in the morning existed by the end of the day. There was a very direct relationship between effort and output.&lt;/p&gt;

&lt;p&gt;Today, that relationship is disappearing.&lt;/p&gt;

&lt;p&gt;A large part of the implementation work around me can now be generated by AI. Not perfectly, not autonomously, and not without supervision, but cheaply enough that the first implementation is increasingly not the hard part. If I need an API endpoint, a migration, a data transformation, a test suite, an integration, or some internal tooling, I can describe the problem, provide enough context, and get a plausible implementation surprisingly quickly.&lt;/p&gt;

&lt;p&gt;The expensive part starts afterwards: does it actually work?&lt;/p&gt;

&lt;p&gt;Not simply, "does the code run?" or "do the tests pass?" The real question is whether the feature behaves correctly across the range of situations we actually care about. That question has slowly become a much larger part of my job, and it has changed how I think about my own value.&lt;/p&gt;

&lt;p&gt;A year ago, I spent much more time thinking about how to implement something. Now I spend much more time thinking about how it can fail. What happens when the model receives something we did not anticipate? When two components disagree? When an apparently good answer contains the wrong evidence? When the evaluator itself is biased? What happens when a model succeeds 95 percent of the time, but the remaining 5 percent contains exactly the failures that matter to the business? What happens when a fallback silently changes provider, region, latency, cost, or behavior? What happens when everything returns HTTP 200 and the system is still wrong?&lt;/p&gt;

&lt;p&gt;The implementation is often no longer the difficult intellectual problem. Failure detection is. Failure prevention is. Defining what "good enough to ship" actually means is.&lt;/p&gt;

&lt;p&gt;That creates a strange inversion. For decades, software engineering treated testing and validation as something downstream of implementation. First you built the thing, then someone checked whether it worked. In AI systems, I increasingly feel the opposite. Generating the implementation is becoming cheaper, while knowing whether you should trust it is becoming more expensive.&lt;/p&gt;

&lt;p&gt;This week, for example, after I take some time off, the team asked me for a new evaluator for a feature. At first glance, that sounds almost like QA work, and I admit that part of me reacts negatively to that idea. I spent years learning how to build software. Am I slowly becoming the person who checks everybody else's work?&lt;/p&gt;

&lt;p&gt;But I think that interpretation misses what is actually happening.&lt;/p&gt;

&lt;p&gt;The hard part of building an evaluator is rarely writing the evaluator. AI can help enormously with that too. The difficult part is deciding what the evaluator should measure in the first place. What counts as failure? Which failures can be checked deterministically? Which ones require statistical evaluation? Which dataset represents reality well enough? Where are the blind spots? Can the evaluator itself be fooled? Are two supposedly independent checks actually making the same mistake? What threshold is sufficient to release something to production?&lt;/p&gt;

&lt;p&gt;Those are not really questions about test implementation. They are questions about the operational definition of correctness. And with AI, correctness is becoming something we have to engineer.&lt;/p&gt;

&lt;p&gt;Traditional software gives us a relatively comfortable contract. For the same inputs and state, deterministic software should generally produce the same output. LLMs break that assumption. They generate plausible outputs across enormous input spaces, and they can produce something structurally correct, linguistically excellent, internally coherent, and completely wrong.&lt;/p&gt;

&lt;p&gt;So we wrap them in software. We constrain them, validate their outputs, compare signals, build fallback paths, add deterministic checks around probabilistic behavior, create evaluation datasets, measure regressions, and observe production traces. Then, inevitably, we discover a failure mode we did not know existed and modify the system again.&lt;/p&gt;

&lt;p&gt;A surprising amount of my work now lives in that boundary. Not creating intelligence, but making probabilistic intelligence reliable enough to become part of a real product.&lt;/p&gt;

&lt;p&gt;This also makes me uncomfortable for another reason. If AI can generate the implementation, why couldn't another AI eventually generate the evaluations too? Why couldn't I describe my principles, my distrust, my way of looking for edge cases, and encode all of that into another agent?&lt;/p&gt;

&lt;p&gt;The answer is probably that I can. And I should.&lt;/p&gt;

&lt;p&gt;If I discover a reliable failure pattern, I want to automate its detection. If I repeatedly perform the same review, I want a system to perform it for me. If a deterministic gate can replace my manual judgment, building that gate is progress. My value cannot depend on protecting tasks from automation, because that would be a losing strategy.&lt;/p&gt;

&lt;p&gt;So perhaps the important distinction is not between work humans can do and work AI can do. It is between known problems and unknown ones.&lt;/p&gt;

&lt;p&gt;Once a failure mode is understood, it becomes cheaper. We can encode it in a test, build an evaluator, add a policy, or teach an agent to look for it. The difficult part then moves somewhere else. The frontier becomes the next thing we do not yet know how to measure, and that seems to be where more and more of my work is going.&lt;/p&gt;

&lt;p&gt;My output is increasingly not code. It is a definition, a constraint, an architecture, a failure taxonomy, a release criterion, or a deterministic guard around something probabilistic. Sometimes the final artifact is twenty lines of code, but those twenty lines may represent two days of thinking about what exactly needs to be prevented.&lt;/p&gt;

&lt;p&gt;That changes how productivity feels. If I generated 1,000 lines of production code in a week, I could easily point at what I created. If I spend the same week investigating one subtle failure mode and eventually add a tiny check that prevents it, the visible output looks much smaller. But the economic value might be significantly larger.&lt;/p&gt;

&lt;p&gt;The system already knows how to generate. The harder question is whether we can depend on what it generates.&lt;/p&gt;

&lt;p&gt;I do not know exactly what this role should be called yet. AI engineer still fits. Reliability engineer fits part of it. Architecture fits another part. Evaluation engineering is clearly becoming important. None of those labels completely captures the transition I am experiencing.&lt;/p&gt;

&lt;p&gt;What I do know is that I am moving away from being primarily the person who produces the implementation. I am becoming the person who asks what the implementation must prove before we trust it.&lt;/p&gt;

&lt;p&gt;That feels strange because software engineering trained us to identify ourselves with building. Code was craftsmanship, output, and evidence that we were useful. AI is making code abundant, and when something becomes abundant, value usually moves somewhere else.&lt;/p&gt;

&lt;p&gt;Maybe the next scarce resource in software engineering is not implementation. Maybe it is the ability to determine when an implementation is wrong before reality does it for you.&lt;/p&gt;

&lt;p&gt;The cheaper part is increasingly generated. The expensive part is knowing where it will fail, and making sure it doesn't.&lt;/p&gt;




&lt;p&gt;Human concept, nice written by AI. &lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>python</category>
    </item>
    <item>
      <title>The Model Does Not Need Memory. The Situation Does.</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Mon, 29 Jun 2026 11:50:14 +0000</pubDate>
      <link>https://dev.to/marcosomma/the-model-does-not-need-memory-the-situation-does-196g</link>
      <guid>https://dev.to/marcosomma/the-model-does-not-need-memory-the-situation-does-196g</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;I think I was asking the wrong question.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For a while, the question was simple: does memory make agents smarter?&lt;/p&gt;

&lt;p&gt;It sounds like the right question. It is also a trap, because it assumes that memory should be judged as a generic intelligence booster. You add a memory layer, the agent remembers more things, and somehow the output should become better. More complete. More accurate. More human. More whatever word we are currently using to avoid saying “I hope this expensive thing works.”&lt;/p&gt;

&lt;p&gt;After running more experiments, I think that framing is wrong.&lt;/p&gt;

&lt;p&gt;Memory does not make the model smarter in any general sense. Most of the time, it cannot. The model already has a massive amount of general procedural and domain knowledge compressed into its weights. If your memory layer recalls information the model already knows, you are not adding intelligence. You are just adding a second path to say the same thing with more latency.&lt;/p&gt;

&lt;p&gt;That is the uncomfortable part. A lot of agent memory systems are not failing because recall is broken. They are failing because the recalled information has no marginal value.&lt;/p&gt;

&lt;p&gt;They are giving the model something it was already going to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  The experiment was not a victory lap
&lt;/h2&gt;

&lt;p&gt;I ran a follow-up benchmark on OrKa Brain: 250 tasks, five tracks, brain versus brainless. The point was to see whether procedural memory would show stronger results at scale and across more task types.&lt;/p&gt;

&lt;p&gt;It did not.&lt;/p&gt;

&lt;p&gt;The absolute rubric score was almost flat. Brain scored 8.39. Brainless scored 8.27. That is a +0.12 difference on a 10-point scale. Technically positive, but not the kind of result you use to announce that persistent memory has unlocked a new era of agent intelligence unless your relationship with evidence is mostly decorative.&lt;/p&gt;

&lt;p&gt;The pairwise result looked slightly better at first. Brain won 53.8% of the comparisons. But then the judge showed a 74.4% first-position bias, which means the raw pairwise score was not something you can read directly. If the judge picks the first answer almost three times out of four, the measurement instrument is not exactly wearing a lab coat. It is flipping a biased coin and writing a confident explanation afterward.&lt;/p&gt;

&lt;p&gt;Once I controlled for position, most of the supposed Brain advantage disappeared. Cross-domain transfer did not survive. Anti-pattern avoidance did not survive. Multi-skill composition collapsed into a coin flip. Routing was confounded and needs to be re-run.&lt;/p&gt;

&lt;p&gt;Only one track survived: the long same-domain sequence. In that track, the Brain won 74% of the time even when placed in the disfavored position.&lt;/p&gt;

&lt;p&gt;That is the part worth keeping.&lt;/p&gt;

&lt;p&gt;Not because it proves that memory works in general. It does not. It proves something narrower and more useful: memory seems to help when the output depends on previous state in the same evolving situation.&lt;/p&gt;

&lt;p&gt;That is a very different claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ceiling was the message
&lt;/h2&gt;

&lt;p&gt;At first glance, you could say the Brain failed. I think that is too easy.&lt;/p&gt;

&lt;p&gt;The better reading is that the benchmark exposed the wrong class of memory. Most of what OrKa Brain recalled was procedural knowledge: how to decompose a task, how to reason through trade-offs, how to avoid generic mistakes, how to structure a solution, how to transfer a pattern from one domain to another.&lt;/p&gt;

&lt;p&gt;The problem is that capable LLMs already know a lot of this. They have seen endless examples of architecture reviews, debugging sessions, migration plans, support flows, project retrospectives, Stack Overflow arguments, incident reports, code reviews, framework docs, and corporate documents that somehow use five pages to say “we forgot the cache.”&lt;/p&gt;

&lt;p&gt;So when the memory layer recalls a generic procedure, the model does not suddenly receive missing information. It receives a reminder of something already present in the weights.&lt;/p&gt;

&lt;p&gt;That is why the ceiling effect matters. It is not just an implementation failure. It is a category warning. If memory stores general competence, the model can often route around it because the model already has general competence.&lt;/p&gt;

&lt;p&gt;This is why “agent memory” can look impressive in demos and then become strangely weak in benchmarks. In a demo, remembered context feels useful because we can see the system referencing the past. In a benchmark, if the remembered thing does not change the answer, the effect disappears into style, verbosity, or judge preference.&lt;/p&gt;

&lt;p&gt;The useful question is not “did the system remember something?”&lt;/p&gt;

&lt;p&gt;The useful question is “did the remembered thing contain information the model could not have known or safely inferred?”&lt;/p&gt;

&lt;p&gt;That is the line.&lt;/p&gt;

&lt;h2&gt;
  
  
  Memory is for contingent information
&lt;/h2&gt;

&lt;p&gt;The sharper version of the theory is this:&lt;/p&gt;

&lt;p&gt;Memory helps when the answer depends on contingent information absent from the model weights.&lt;/p&gt;

&lt;p&gt;Not vertical knowledge. Not domain knowledge. Not “more context” as a generic spell. Contingent information.&lt;/p&gt;

&lt;p&gt;By contingent, I mean information whose truth depends on a specific user, system, company, customer, codebase, previous decision, local process, or moment in time. It is information that is not true in general. It is true here.&lt;/p&gt;

&lt;p&gt;The model can know how software migrations usually fail. It cannot know that in this codebase, the last migration failed because the billing worker silently depended on a deprecated Redis key.&lt;/p&gt;

&lt;p&gt;The model can know what concise writing is. It cannot know that a specific user prefers direct technical answers with no filler, no fake enthusiasm, and no corporate perfume sprayed over the paragraph.&lt;/p&gt;

&lt;p&gt;The model can know how customer support triage works. It cannot know that this specific customer always reports billing bugs using the wrong product name.&lt;/p&gt;

&lt;p&gt;The model can know how a deployment pipeline usually works. It cannot know that this team avoids Friday releases because one rollback path still depends on a manual script nobody wants to admit exists.&lt;/p&gt;

&lt;p&gt;That is memory.&lt;/p&gt;

&lt;p&gt;Everything else risks becoming a second copy of generic knowledge attached to a model that already has the first copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codebase AI makes the distinction obvious
&lt;/h2&gt;

&lt;p&gt;Take software engineering, because the mistake becomes very clear there.&lt;/p&gt;

&lt;p&gt;A naive approach says: “We need memory so the model knows how to code.” That sounds reasonable until you look at what the model already knows. A capable model has broad exposure to programming languages, design patterns, API conventions, testing strategies, refactoring techniques, infrastructure patterns, and the usual graveyard of best practices everyone quotes and half the industry ignores.&lt;/p&gt;

&lt;p&gt;If you use memory to teach it generic software knowledge, you may hit the same ceiling as procedural memory. The model already knows that functions should be small, tests should cover edge cases, migrations should be reversible, and distributed systems enjoy ruining your afternoon. Storing those ideas as memory does not add much. It mostly gives the model a slower way to say what it was already going to say.&lt;/p&gt;

&lt;p&gt;The problem is not that the model does not know software engineering. The problem is that it does not know this software system.&lt;/p&gt;

&lt;p&gt;That means generic engineering knowledge should not be treated as the valuable memory layer. It is already in the model. The valuable layer is the local shape of the codebase: naming conventions, architectural scars, forbidden dependencies, deployment constraints, hidden coupling, flaky tests, internal abstractions, legacy decisions, and the reason why one ugly function must not be “cleaned up” unless you enjoy incident reports.&lt;/p&gt;

&lt;p&gt;The job of repository retrieval is not to remind the model what clean code is. The job is to show it what this codebase actually is.&lt;/p&gt;

&lt;p&gt;The model can know how authentication usually works. It cannot know that in this system, the admin role is duplicated across two services because a migration was only half-completed in 2022.&lt;/p&gt;

&lt;p&gt;The model can know that Redis keys should have clear ownership. It cannot know that billing still depends on a deprecated cache key because one worker was never moved to the new event pipeline.&lt;/p&gt;

&lt;p&gt;The model can know how to write a database migration. It cannot know that this team avoids destructive migrations unless the rollback plan is reviewed by the person who still remembers why the old schema exists.&lt;/p&gt;

&lt;p&gt;That is where memory earns its keep.&lt;/p&gt;

&lt;p&gt;A style guide tells the model how code should look. Repository history tells it why the code looks wrong but still works. Incident memory tells it where the bodies are buried. Team conventions tell it what changes will be accepted or rejected before CI even has the chance to complain.&lt;/p&gt;

&lt;p&gt;Those are not the same layer, and treating them as one generic RAG pile is how you get a system that retrieves a lot and understands very little.&lt;/p&gt;

&lt;p&gt;In software products, this distinction matters. Documentation grounds. Repository state shapes. Incident and decision history constrain.&lt;/p&gt;

&lt;p&gt;If those three jobs are mixed together, the assistant may still sound like a senior engineer. It may even sound more senior than before, which usually means it has learned to say “trade-off” with confidence. The question is whether it understands the specific system in front of it.&lt;/p&gt;

&lt;p&gt;That is where memory stops being decoration and starts being useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chat memory is the same pattern
&lt;/h2&gt;

&lt;p&gt;This is not only a codebase problem. Chat memory shows the same structure in a smaller and more familiar form.&lt;/p&gt;

&lt;p&gt;When a system remembers that a user wants concise answers, it is not helping because the model lacked the concept of concision. The model knows how to be concise. The useful information is not “what does concise mean?” The useful information is “what does concise mean for this user?”&lt;/p&gt;

&lt;p&gt;That is an entity-bound fact. It attaches a general capability to a specific person.&lt;/p&gt;

&lt;p&gt;This is why personalization can be useful without being intellectually deep. The value is not in teaching the model a new writing style from scratch. The value is in selecting the right behavior under a known identity.&lt;/p&gt;

&lt;p&gt;The same thing happens inside companies. The model knows how to review code, but it does not know that this repository bans a specific library because it caused a production incident three years ago. The model knows how to process refunds, but it does not know that annual partner contracts have a different refund path. The model knows how to summarize a customer complaint, but it does not know that this customer always describes production incidents as “minor issues” because they are polite to the point of self-sabotage.&lt;/p&gt;

&lt;p&gt;That local weirdness is not noise. It is the work.&lt;/p&gt;

&lt;p&gt;Real systems are not made only of general rules. They are made of exceptions, scars, habits, conventions, previous mistakes, policy shortcuts, undocumented dependencies, and things everyone knows until the one person who knows leaves the company.&lt;/p&gt;

&lt;p&gt;That is the material memory should store.&lt;/p&gt;

&lt;p&gt;Not generic competence. Operational sediment.&lt;/p&gt;

&lt;h2&gt;
  
  
  This changes the architecture
&lt;/h2&gt;

&lt;p&gt;My earlier framing was closer to skill reuse. The agent learns a skill, stores it, recalls it later, and applies it to a new task. It is a clean model. It is also suspiciously diagram-friendly, which should always make us nervous.&lt;/p&gt;

&lt;p&gt;The data pushes me toward a different architecture.&lt;/p&gt;

&lt;p&gt;A memory system should not start by asking, “What skill is similar to this task?” That can work sometimes, but it is too weak as the central retrieval principle. Similarity is not the same as necessity. A memory can be semantically close and still useless if it does not change the answer.&lt;/p&gt;

&lt;p&gt;The better retrieval question is: “What information about this situation would be impossible, unsafe, or expensive for the model to infer?”&lt;/p&gt;

&lt;p&gt;That question changes everything.&lt;/p&gt;

&lt;p&gt;You retrieve because the task depends on a known entity. You retrieve because the answer varies by product, customer, repository, environment, deployment target, or team convention. You retrieve because the user has stable preferences. You retrieve because the codebase has local rules. You retrieve because the customer has a history. You retrieve because the system has failed before in a way that looks relevant now. You retrieve because the model is about to answer with generic competence where specific state is required.&lt;/p&gt;

&lt;p&gt;That is a different trigger from semantic similarity.&lt;/p&gt;

&lt;p&gt;The trigger is underdetermination. The model can produce an answer, but the answer is not sufficiently determined by the prompt and general knowledge. It needs local state.&lt;/p&gt;

&lt;p&gt;This also means memory should not be one bucket. A serious system probably needs separate stores with separate jobs.&lt;/p&gt;

&lt;p&gt;Grounding is for authority. It gives the model exact sources, current docs, repository files, API contracts, schemas, config, and verifiable references.&lt;/p&gt;

&lt;p&gt;Operational knowledge is for local shape. It tells the model how this product, repository, customer, workflow, team, or deployment environment behaves.&lt;/p&gt;

&lt;p&gt;Episodic memory is for history. It tells the model what happened before with this user, customer, task, session, service, or system.&lt;/p&gt;

&lt;p&gt;Reasoning composes the three.&lt;/p&gt;

&lt;p&gt;When people collapse all of this into “memory,” they make the system harder to evaluate. They also make it easier to fool themselves, which is apparently the default MLOps workflow whenever agents are involved.&lt;/p&gt;

&lt;h2&gt;
  
  
  The next benchmark should be crueler
&lt;/h2&gt;

&lt;p&gt;The old question was whether memory improves output quality. That is too broad.&lt;/p&gt;

&lt;p&gt;If the model can answer well using general competence, memory will struggle to show a large effect. A judge may prefer one answer over another, but you will end up measuring style, verbosity, formatting, position bias, or the judge’s breakfast.&lt;/p&gt;

&lt;p&gt;The next benchmark should make memory necessary.&lt;/p&gt;

&lt;p&gt;A task should require a fact that exists only in memory or grounding. Without that fact, the model should either fail, hedge, or produce a generic answer. With that fact, the model should answer correctly.&lt;/p&gt;

&lt;p&gt;Not more beautifully. Correctly.&lt;/p&gt;

&lt;p&gt;For a chat assistant, the answer might depend on a stored user preference. For a code assistant, it might depend on a repository-specific rule, a past incident, a banned dependency, a migration constraint, or a team convention that exists nowhere in the model weights. For support automation, it might depend on a customer-specific exception, a product-plan quirk, or a previous escalation pattern. For routing, it might depend on a previous failed path inside the same system.&lt;/p&gt;

&lt;p&gt;The score should not be “which answer feels more complete?” That is how judge bias walks into the room, sits at the table, and starts grading vibes.&lt;/p&gt;

&lt;p&gt;The score should be concrete. Did the system use the right repository rule? Did it respect the deployment constraint? Did it remember the user preference? Did it avoid the known failure path? Did it cite the current internal document? Did it preserve the prior product decision? Did it distinguish generic knowledge from contingent state?&lt;/p&gt;

&lt;p&gt;That is where memory should show a real gap. Not a +0.12 polish gap. A correctness gap.&lt;/p&gt;

&lt;p&gt;If memory cannot win there, then the memory system is probably not doing much. If it does win there, then the value is no longer mystical. It is measurable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The anti-hype version
&lt;/h2&gt;

&lt;p&gt;I do not think “memory makes agents smarter” is the right claim. It is too vague, and vague claims are where AI hype goes to reproduce.&lt;/p&gt;

&lt;p&gt;The sharper claim is less exciting but more useful: memory helps when the task depends on information that is not in the model weights and cannot be safely inferred from general knowledge.&lt;/p&gt;

&lt;p&gt;This explains why generic procedural recall barely moved the needle. It explains why long same-domain recall was the only signal that survived. It explains why chat preferences matter. It explains why code assistants need repository state plus decision history, not just more documentation. It explains why enterprise assistants become useful only when they know the local weirdness of the company.&lt;/p&gt;

&lt;p&gt;The model already has the average.&lt;/p&gt;

&lt;p&gt;Memory is for the deviation.&lt;/p&gt;

&lt;p&gt;And most real work is deviation.&lt;/p&gt;

&lt;p&gt;That is the part I think we keep missing. We keep building memory systems as if the model is empty and needs to be filled. But the model is not empty. It is full of averages. Full of general patterns. Full of plausible procedures. Full of the kind of answer that is usually right until it meets a real organization, a real customer, a real codebase, or a real human being with preferences that do not fit the median.&lt;/p&gt;

&lt;p&gt;Memory is not there to compete with pretraining.&lt;/p&gt;

&lt;p&gt;Memory is there to correct the average when the local situation demands it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boundary I would use now
&lt;/h2&gt;

&lt;p&gt;I would draw the boundary like this.&lt;/p&gt;

&lt;p&gt;Do not store what the model already knows. Do not retrieve what the model can safely infer. Do not call generic advice memory. Do not confuse domain knowledge with contingent state.&lt;/p&gt;

&lt;p&gt;Store the things that change the answer because they are specific, current, local, personal, historical, procedural, repository-bound, customer-bound, product-bound, or system-bound.&lt;/p&gt;

&lt;p&gt;Use grounding for authority. Use memory for contingency. Use reasoning for composition.&lt;/p&gt;

&lt;p&gt;Mixing those three together is how you build a very expensive autocomplete with a scrapbook attached.&lt;/p&gt;

&lt;p&gt;The benchmark did not prove that memory is useless. It proved that memory has to earn its place. If the remembered thing does not change the answer, it is not memory in any meaningful operational sense. It is just noise with a timestamp.&lt;/p&gt;

&lt;p&gt;The model does not need memory.&lt;/p&gt;

&lt;p&gt;The situation does.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>mcp</category>
      <category>llm</category>
    </item>
    <item>
      <title>Maybe It Is Not Yet Time To Bring Every AI Demo To Production</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Tue, 23 Jun 2026 14:26:29 +0000</pubDate>
      <link>https://dev.to/marcosomma/maybe-it-is-not-yet-time-to-bring-every-ai-demo-to-production-o74</link>
      <guid>https://dev.to/marcosomma/maybe-it-is-not-yet-time-to-bring-every-ai-demo-to-production-o74</guid>
      <description>&lt;p&gt;There is a sentence I keep hearing in AI engineering that sounds innocent, practical, and mature: “Just add a fallback provider.”&lt;/p&gt;

&lt;p&gt;Clean. Elegant. Wonderful. The kind of sentence that usually survives only until production starts touching it. Because in a demo, fallback means: Provider A fails, call Provider B. In production, fallback means something very different.&lt;/p&gt;

&lt;p&gt;Will Provider B interpret the prompt in the same way? Will it serialize the tool schema in the same way? Will it respect cache directives in the same way? Will it stream tokens in the same way? Will it expose errors in the same way? Will it count tokens in the same way? Will it respect timeouts in the same way? Will it fail in a way your system can actually understand?&lt;/p&gt;

&lt;p&gt;Most of the time, the answer is no.&lt;/p&gt;

&lt;p&gt;And this is where the current AI industry keeps doing its favorite magic trick. It takes something deeply unstable, wraps it in a familiar API shape, gives it a shiny compatibility label, and suddenly everyone behaves as if we have a standard. We do not have a standard. &lt;strong&gt;&lt;em&gt;We have a costume!&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most of famous “OpenAI compatible" APIs are laying, and hiding the lack of standards behind a known name. In reality the is compatible only on the shallowest path. You can send a basic chat request and get text back. Fantastic. The demo works. The slide looks good. The architecture diagram has fewer boxes. But the moment you move beyond “hello model, summarize this paragraph”, things start to fracture.&lt;/p&gt;

&lt;p&gt;Tool calling. Structured output. JSON enforcement. Prompt caching. Streaming. Retry behavior. Usage accounting. Model aliases. Safety overlays. Regional routing. Timeout semantics. Error objects. Response envelopes. Context handling. Provider-specific parameters. All the boring parts. In other words, all the parts that decide whether your AI system survives production.&lt;/p&gt;

&lt;h2&gt;
  
  
  As we know, the demo works because the demo is NOT the system
&lt;/h2&gt;

&lt;p&gt;The demo is usually a happy path. One user. One model. One provider. One prompt. One task. Maybe no cache. Maybe no concurrency. Maybe no structured output. Maybe no audit trail. Maybe no fallback. Maybe no customer-specific version pinning. Maybe no compliance requirement. Maybe no cost pressure. Maybe no incident where several parallel streams connect and then produce absolutely nothing for two minutes.&lt;/p&gt;

&lt;p&gt;In that world, AI feels magical. In production, AI feels like distributed systems decided to have a child with legal ambiguity and probabilistic behavior.&lt;/p&gt;

&lt;p&gt;You are not only integrating a model. You are integrating a runtime. And that runtime is usually not specified clearly enough. This is the part people keep missing.&lt;/p&gt;

&lt;p&gt;The model is not the full product. The provider’s serving stack is part of the product. The SDK is part of the product. The serialization layer is part of the product. The cache implementation is part of the product. The safety wrapper is part of the product. The regional routing strategy is part of the product.&lt;/p&gt;

&lt;p&gt;So when someone says, “it is the same model”, I increasingly hear: “we did not measure the parts around the model.”&lt;/p&gt;

&lt;p&gt;Same weights do not mean same behavior. Same model family does not mean same production contract. Same endpoint shape does not mean same system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same model, different reality
&lt;/h2&gt;

&lt;p&gt;One of the strongest examples I have seen came from a direct comparison between Provider A and Provider B using the same model family on real production-like workflows. The headline looked simple: same model family, different provider path. The result was not simple.&lt;/p&gt;

&lt;p&gt;On Workflow 1, the quality regression was statistically significant. Provider A had a mean score of 0.716. Provider B had a mean score of 0.497. The p-value was below 0.0001, with a medium effect size. That is not “a bit of noise.” That is the kind of difference that should stop a migration.&lt;/p&gt;

&lt;p&gt;The interesting part is that not every workflow regressed. On Workflow 2 and Workflow 3, the result was basically fine.&lt;/p&gt;

&lt;p&gt;Good. That makes the result more credible, not less. Because real provider migrations do not fail everywhere. They fail in specific workflows, specific prompts, specific schema paths, specific flows, specific edge cases. The average can look acceptable while one critical workflow quietly gets worse.&lt;/p&gt;

&lt;p&gt;This is exactly why &lt;strong&gt;&lt;em&gt;“we tested a few prompts manually and it looked okay”&lt;/em&gt;&lt;/strong&gt; is not engineering. It is theater with curl commands.&lt;/p&gt;

&lt;p&gt;If you want to switch provider, you need replay traces. You need evals. You need per-workflow scores. You need statistical comparison. You need to know where the behavior changed, not just whether the model still speaks fluent corporate English.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost is not only list price
&lt;/h2&gt;

&lt;p&gt;The same comparison showed around a 2x cost premium on a high-volume workflow. At first glance, you might blame provider pricing. But the back-calculation pointed somewhere more boring and more dangerous: prompt caching.&lt;/p&gt;

&lt;p&gt;On Provider A, implied token volume was 60 to 67 percent below reported tokens. That is the cache signature. You are still sending the structure, but you are not paying the full input cost every time because the provider is reusing cached prompt blocks.&lt;/p&gt;

&lt;p&gt;On Provider B, one high-volume path showed exactly 0 percent gap. Cache was either off or always missing. Other paths showed partial cache behavior, around 14 to 21 percent in one case and around 33 percent in another.&lt;/p&gt;

&lt;p&gt;Same model family. Different cache reality. Different bill. This is where the “just switch provider” crowd usually becomes very quiet.&lt;/p&gt;

&lt;p&gt;Because caching is not decoration. In high-volume AI systems, caching is part of the economic architecture. If cache semantics change, your unit economics change. If regional routing causes cache misses, your cost model changes. If one provider respects cache directives differently from another, your production bill changes while every individual request still “works.”&lt;/p&gt;

&lt;p&gt;That is the worst kind of failure. The successful one. No exception. No stack trace. No screaming service. Just a quiet invoice telling you the abstraction was fake.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-region cache is a beautiful little trap
&lt;/h2&gt;

&lt;p&gt;Cross-region inference sounds robust. More regions. More availability. More resilience. Then you look at the cache behavior.&lt;/p&gt;

&lt;p&gt;A request served in Region A writes a cache in Region A. The next request may route to Region B. Region B does not have that cache. So it misses and writes again. Then another call may route back to Region A, or somewhere else, depending on capacity and routing.&lt;/p&gt;

&lt;p&gt;This is not a clean “double pay” situation. It is worse conceptually. You keep paying the cache write premium without reliably amortizing it through cheap cache reads. That is how you can end up with a measured 0 percent hit rate while thinking you configured caching correctly.&lt;/p&gt;

&lt;p&gt;Again, from the outside everything looks compatible. The API accepts your request. The model responds. The integration works. Except the economics are different because the serving layer changed.&lt;/p&gt;

&lt;p&gt;This is why AI production work is becoming less about prompts and more about contracts. What exactly is guaranteed? What is pinned? What is regional? What is cached? What is counted? What is replayable? What is stable?&lt;/p&gt;

&lt;p&gt;If the answer is “trust us, it is compatible”, my engineering translation is simple: no contract found.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reliability is not portable either
&lt;/h2&gt;

&lt;p&gt;Another production-style incident: under concurrency, multiple parallel streams connected and then produced nothing for roughly two minutes. No tokens. No useful error. Just waiting.&lt;/p&gt;

&lt;p&gt;The likely reading was capacity or throttle queueing. The provider may have been holding the request instead of returning a clean throttling response. Depending on the endpoint, one path may queue in-flight work while another may throw a clear rate-limit error.&lt;/p&gt;

&lt;p&gt;That distinction matters. A clear rate-limit error is ugly but useful. You can react to it. You can retry with backoff. You can trigger fallback. You can protect the system. A connected stream producing nothing for two minutes is a different species of failure. Your system is alive enough to wait and dead enough to be useless.&lt;/p&gt;

&lt;p&gt;There was also a competing hypothesis: maybe the network layer was involved. Gateway behavior, private endpoints, load balancers, idle timeouts, streaming connection drops, or capacity errors could all produce overlapping symptoms.&lt;/p&gt;

&lt;p&gt;So the correct response was not “the provider is bad.” The correct response was: inspect runtime metrics during the hang windows. Check throttle counters. Check server error counters. Check network timeouts. Check connection lifetime. Check whether the request reached the model runtime at all.&lt;/p&gt;

&lt;p&gt;This is what production AI looks like. Not prompt magic. Not demo videos. Not “look, I built an agent in 20 minutes.” It looks like debugging whether a zero-token two-minute hang is caused by model capacity, runtime queueing, network infrastructure, streaming semantics, retry policy, or your own concurrency design.&lt;/p&gt;

&lt;p&gt;Very glamorous. Someone should put that in the launch video.&lt;/p&gt;

&lt;h2&gt;
  
  
  Structured output is not standard output
&lt;/h2&gt;

&lt;p&gt;Then there is the SDK serialization problem. Same model. Same app-level input. Different token count.&lt;/p&gt;

&lt;p&gt;One comparison showed Provider B using around 10,473 input tokens while Provider A used around 10,019. That is a 454-token delta, roughly 4.5 percent.&lt;/p&gt;

&lt;p&gt;The clue was structured output. On one provider path, structured output was implemented by injecting a tool schema into the prompt. On the other path, it was handled differently. Even after making payloads byte-identical at the application level, the remaining structural difference came from provider-specific cache directive serialization.&lt;/p&gt;

&lt;p&gt;This is a perfect example of why API compatibility is not enough. Your prompt may be identical. Your provider prompt is not.&lt;/p&gt;

&lt;p&gt;The actual thing seen by the model may include hidden scaffolding, injected schemas, translated parameters, safety wrappers, tool definitions, response constraints, or provider-specific envelopes. Then we compare eval scores and pretend we tested the same thing.&lt;/p&gt;

&lt;p&gt;Did we? Maybe. Maybe not. And if we cannot answer that confidently, then we are not measuring model quality. We are measuring a mix of model behavior, SDK translation, provider scaffolding, and our own assumptions.&lt;/p&gt;

&lt;p&gt;Very scientific. Very enterprise. Very “move fast and accidentally compare different systems.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Fallback can become the outage
&lt;/h2&gt;

&lt;p&gt;Cross-provider fallback sounds responsible. It can be responsible. But it is not free.&lt;/p&gt;

&lt;p&gt;One concrete incident involved a preview model on Provider C. The model had intermittent hangs, produced retry-exhaustion errors after repeated timeouts, and reported zero input tokens and zero output tokens. So the model did not even really start.&lt;/p&gt;

&lt;p&gt;The retry budget burned for several minutes. Then a failure-rate guard aborted the whole job. The fix was to add a fallback model. Good fix. But the lesson is bigger.&lt;/p&gt;

&lt;p&gt;The fallback path needs its own engineering. It needs its own timeout budget. It needs its own cost assumption. It needs its own quality expectation. It needs its own reason to exist.&lt;/p&gt;

&lt;p&gt;A useful rule: scale up on fallback by default. If fallback runs rarely, a usable answer matters more than saving a few cents. Scale down only when the primary failed because the request exceeded model limits.&lt;/p&gt;

&lt;p&gt;But if your fallback inherits the same exhausted timeout budget from the primary, congratulations, you did not build fallback. You built a decorative second failure.&lt;/p&gt;

&lt;p&gt;Fallback is not a backup model. Fallback is a second production path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Even parameter names are not portable
&lt;/h2&gt;

&lt;p&gt;Small example, but very revealing: a “compatible” API for a model behaved differently around a reasoning-related parameter. The workaround was to force a safe default.&lt;/p&gt;

&lt;p&gt;That is reasonable. But the real portability risk is not only the value. It is the parameter contract itself.&lt;/p&gt;

&lt;p&gt;Another provider may call it something else. Another may ignore it. Another may reject it. Another may apply a different default. Another may support it only on some models. Another may support it in preview and remove it later with very little warning.&lt;/p&gt;

&lt;p&gt;This is where “compatible API” starts to feel like saying every car is steering-wheel-compatible. Technically true. Please do not use that as your safety case.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preview is not production
&lt;/h2&gt;

&lt;p&gt;A lot of AI teams are building production workflows on preview models, preview parameters, preview endpoints, preview SDK behavior, and preview pricing assumptions. Then they act surprised when preview behaves like preview.&lt;/p&gt;

&lt;p&gt;Preview can mean weaker guarantees. It can mean limited support. It can mean behavior changes. It can mean short deprecation windows. It can mean different rate limits. It can mean hidden routing changes. It can mean features that work today and become “not recommended” tomorrow.&lt;/p&gt;

&lt;p&gt;That is fine for exploration. That is not fine when your production system depends on it and nobody wrote down the risk.&lt;/p&gt;

&lt;p&gt;Again, the issue is not that preview exists. Preview is useful. The issue is pretending preview is stable because the demo worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  We need stable interfaces, not demo optimism
&lt;/h2&gt;

&lt;p&gt;I am increasingly convinced that production AI needs something closer to long-term-support thinking.&lt;/p&gt;

&lt;p&gt;Not because models should stop improving. They will improve. The field moves fast. Fine. But production systems cannot keep pretending that every model upgrade, provider switch, SDK change, cache behavior update, or model alias movement is harmless.&lt;/p&gt;

&lt;p&gt;When a system is performing fine, switching the model or serving path can create more issues than benefits.&lt;/p&gt;

&lt;p&gt;The defensible version is conditional: long-term support becomes inevitable when capability growth slows enough that stability outweighs the next incremental benchmark gain. At that point, many companies will not want the newest model. They will want the model-runtime contract that keeps working.&lt;/p&gt;

&lt;p&gt;But the deeper point is that the thing needing long-term support is not only the model weights. It is the interface.&lt;/p&gt;

&lt;p&gt;The stable surface must include serving stack, quantization, SDK behavior, tool serialization, cache semantics, timeout behavior, safety overlays, error formats, and versioned model aliases.&lt;/p&gt;

&lt;p&gt;Maybe the real answer is not long-term-support models. Maybe it is long-term-support interfaces with deterministic check layers.&lt;/p&gt;

&lt;p&gt;Swappable models behind a stable contract. Replayable traces. Eval gates. Schema normalization. Provider-specific adapters. Explicit cache tests. Timeout isolation. Failure-mode classification. Version-pinned prompts. Known fallback policy.&lt;/p&gt;

&lt;p&gt;That sounds boring. Good. Production should be boring.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem is not that AI is useless
&lt;/h2&gt;

&lt;p&gt;This is usually where someone misunderstands the argument. The point is not that AI is useless. The point is not that demos are bad. The point is not that teams should stop experimenting.&lt;/p&gt;

&lt;p&gt;The point is that demos and production systems are different organisms. A demo proves possibility. Production requires repeatability. A demo proves that the model can answer. Production requires knowing what happens when it does not answer, answers differently, answers slowly, answers with a hidden schema injection, misses cache across regions, changes token accounting, streams forever, returns a provider-specific error, or silently regresses one workflow while improving another.&lt;/p&gt;

&lt;p&gt;AI interfaces today are still too fragmented for the amount of confidence people are placing in them. We are building production systems on unstable runtime surfaces and pretending the abstraction is mature because the JSON shape looks familiar.&lt;/p&gt;

&lt;p&gt;That is not engineering maturity. That is hope with headers.&lt;/p&gt;

&lt;h2&gt;
  
  
  So maybe not every demo belongs in prod yet
&lt;/h2&gt;

&lt;p&gt;Maybe it is not yet time to bring every AI demo to production. Or more precisely: maybe it is not time to bring demos to production without first building the missing runtime layer around them.&lt;/p&gt;

&lt;p&gt;Not another wrapper. Not another “universal SDK” that hides provider differences until they explode. A real layer. One that treats each provider as a different runtime with different semantics.&lt;/p&gt;

&lt;p&gt;One that records traces. Replays production samples. Compares quality. Measures cost after caching. Tracks token deltas. Normalizes errors. Separates timeout budgets. Tests fallback paths. Pins model versions. Detects serialization drift. Audits structured output behavior. Makes provider migration observable before it becomes an outage.&lt;/p&gt;

&lt;p&gt;Because changing provider is not changing a base URL. It is migrating the runtime contract of your AI system.&lt;/p&gt;

&lt;p&gt;And if your system does not know what that contract is, then the provider switch is not a migration. It is an experiment in production.&lt;/p&gt;

&lt;p&gt;Very innovative, yes. Also known in some older engineering traditions as a bad idea.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Am I Becoming Too Slow for the AI World?</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Wed, 03 Jun 2026 17:30:38 +0000</pubDate>
      <link>https://dev.to/marcosomma/am-i-becoming-too-slow-for-the-ai-world-1904</link>
      <guid>https://dev.to/marcosomma/am-i-becoming-too-slow-for-the-ai-world-1904</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;The AI world is full of old infrastructure with stochastic organs.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence probably explains better than anything why I feel slow lately. Not because I stopped caring about AI. Not because I cannot build anymore. Not because the tools moved beyond me. If anything, the tools moved in the opposite direction: they make me faster at generating code, faster at prototyping, faster at touching layers I would normally approach one by one.&lt;/p&gt;

&lt;p&gt;And still, inside the work, I feel slower.&lt;/p&gt;

&lt;p&gt;This is the uncomfortable part. AI gives me the sensation that everything should move faster, but the more I use it seriously, the more I end up spending time in the parts that do not accelerate cleanly. The code appears quickly. The draft appears quickly. The workflow appears quickly. But then I need to understand what it actually does. I need to test the boundary between components. I need to verify that the result is not just plausible, but correct enough to survive contact with reality.&lt;/p&gt;

&lt;p&gt;Maybe this is the first real trap of AI development:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;Creation became cheap. Verification did not.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That sounds simple, almost too obvious, but I think it is the reason many of us feel this strange mismatch between speed and exhaustion. We can now produce more surface area than we can comfortably inspect. A feature that would have taken days to sketch can appear in hours. A backend route, a frontend component, a prompt chain, a test, a deployment script, a workflow diagram: all of it can be generated quickly enough to create the illusion that the whole process compressed.&lt;/p&gt;

&lt;p&gt;But the whole process did not compress.&lt;/p&gt;

&lt;p&gt;The expensive part moved.&lt;/p&gt;

&lt;p&gt;You still have to understand the system. You still have to validate the assumptions. You still have to test the end-to-end behavior. You still have to ask whether the thing is stable, maintainable, safe, observable, and aligned with the original intention. AI made the first draft cheaper, but it also made it easier to produce first drafts across more layers at once. So the final burden often becomes larger, not smaller.&lt;/p&gt;

&lt;p&gt;Before AI, there was a natural friction in development. You wrote more slowly, so production and understanding were closer together. Now production can sprint ahead of understanding. You can build faster than you can digest. And once that happens, the bottleneck becomes obvious.&lt;/p&gt;

&lt;p&gt;It is not typing.&lt;/p&gt;

&lt;p&gt;It is not even coding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;It is judgment.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is where my feeling of slowness begins. I am not slow in the part that AI is good at accelerating. I am slow in the part that matters after acceleration. I am slow when I need to decide whether a generated thing deserves to exist inside a system. I am slow when I need to move from “this works once” to “this is trustworthy enough to become infrastructure.”&lt;/p&gt;

&lt;p&gt;That kind of slowness does not look good in a noisy field.&lt;/p&gt;

&lt;p&gt;The AI world rewards motion. It rewards people who react quickly, rename quickly, package quickly, comment quickly, and post before the concept has cooled down. Every week there is a new model, a new benchmark, a new agent framework, a new “this changes everything” moment, a new tool that is apparently going to replace half the industry and then disappear into a GitHub archive two weeks later.&lt;/p&gt;

&lt;p&gt;At some point, I became tired of chasing it.&lt;/p&gt;

&lt;p&gt;Not completely. I still follow the field. I still care about what matters. But I no longer have the energy to treat every release, every demo, every thread, and every AI product with a gradient background as if it deserves my full attention. A proof of concept is not a product. A prompt chain is not cognition. A wrapper is not infrastructure. A dashboard is not an operating system for intelligence just because someone wrote “agentic” in the hero section.&lt;/p&gt;

&lt;p&gt;After a while, the noise becomes expensive.&lt;/p&gt;

&lt;p&gt;And when I stopped chasing every update, something strange happened. I did not become less interested in AI. I became more interested in older things.&lt;/p&gt;

&lt;p&gt;Distributed systems. Permissions. Control loops. Network optimization. Separation of concerns. Routing. Handshakes. Feedback. Biological systems. Ant colonies. Viruses. Evolution. Complex adaptive systems.&lt;/p&gt;

&lt;p&gt;The more I look at AI, the more I see old problems returning in a new substrate. Not copied perfectly. Not solved automatically. But returning. The same families of problems keep appearing under newer language. How do parts coordinate? Who is allowed to access what? Where does a decision happen? What happens when a component fails silently? How do you stop local uncertainty from becoming global corruption? How do you keep a system observable when some of its organs speak in probabilities?&lt;/p&gt;

&lt;p&gt;That is why I keep saying that AI infrastructure often feels like old infrastructure with stochastic organs. The body is familiar. The organs behave differently.&lt;/p&gt;

&lt;p&gt;And this is where I need to be more precise, because vague depth is too easy. Saying “old principles return” is not enough. It risks becoming another elegant sentence that avoids doing the actual work.&lt;/p&gt;

&lt;p&gt;Take orchestration.&lt;/p&gt;

&lt;p&gt;A lot of what people now call AI orchestration is not conceptually new. We already had coordination, routing, permissions, queues, retries, fallbacks, handshakes, separation of concerns, and validation boundaries. None of those appeared because LLMs arrived. They were already part of software, distributed systems, automation, and infrastructure engineering.&lt;/p&gt;

&lt;p&gt;But the component changed.&lt;/p&gt;

&lt;p&gt;A deterministic service fails in ways we can often reason about. A request times out. A schema breaks. A permission check rejects access. A queue stalls. A dependency returns an error. The failure may be painful, but at least it often announces itself.&lt;/p&gt;

&lt;p&gt;A generative component can fail while sounding successful.&lt;/p&gt;

&lt;p&gt;It can return a clean boolean and still be wrong. It can pass a check while carrying uncertainty underneath the return type. It can produce a fluent answer that looks like completion but is actually drift. It can say “yes” with confidence because the prompt, the model, and the input distribution all lined up toward the same blind spot.&lt;/p&gt;

&lt;p&gt;That is the part that changes the orchestration problem.&lt;/p&gt;

&lt;p&gt;The old principle still matters: boundaries are useful. Checks are useful. Separation of concerns is useful. Permissions are useful. But the error model under the boundary is different. A handshake with a stochastic component is not the same as a handshake with a deterministic one. The interface may look clean, but the uncertainty has not disappeared. It has been compressed.&lt;/p&gt;

&lt;p&gt;This is why I do not think the right adaptation is simply “use AI, then check AI.” That is too soft. It sounds responsible, but it hides the important question: what kind of check, under what error model, with what independence assumptions?&lt;/p&gt;

&lt;p&gt;Imagine a pipeline like this:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;A &amp;gt; B &amp;gt; C &amp;gt; final check&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;If the final check passes, the system continues. If it fails, the system routes somewhere else. This is simple. It is also risky, because an early hallucination can travel through the whole chain before being inspected.&lt;/p&gt;

&lt;p&gt;So we make the process more granular:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;A &amp;gt; check &amp;gt; B &amp;gt; check &amp;gt; C &amp;gt; check &amp;gt; evaluate checks&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That feels better, and in many cases it is better. Granularity gives the system a higher sampling rate. It interrupts compounding. It catches problems earlier. It lowers the impact of one local failure because the system becomes observable and interruptible at more points.&lt;/p&gt;

&lt;p&gt;But it does not magically lower the uncertainty of the model.&lt;/p&gt;

&lt;p&gt;It lowers propagated uncertainty. It lowers blast radius. It lowers the chance that one bad step contaminates everything downstream. Those are real gains.&lt;/p&gt;

&lt;p&gt;But the atomic uncertainty remains.&lt;/p&gt;

&lt;p&gt;And here is the part I think matters more than the usual AI safety slogan: more checks do not help much if all the checks share the same blind spot.&lt;/p&gt;

&lt;p&gt;Classical retry logic quietly assumes some degree of independence. If a service call fails because of a transient network problem, trying again may work. If a worker crashes because of temporary load, retrying elsewhere may work. The same idea often gets imported into AI workflows without being inspected: ask again, check again, validate again, add another gate.&lt;/p&gt;

&lt;p&gt;But generative errors are often correlated.&lt;/p&gt;

&lt;p&gt;The same model, with the same prompt family, reading the same kind of input, can reproduce the same wrong conclusion several times. A pipeline can collect many green checkmarks that all share the same flaw. At that point, granularity does not create safety. It creates high-resolution false confidence.&lt;/p&gt;

&lt;p&gt;That is the difference between variance and bias.&lt;/p&gt;

&lt;p&gt;Granular checks help with variance. They catch random, local, one-off deviations. They make the system less fragile against isolated mistakes.&lt;/p&gt;

&lt;p&gt;They do not fix bias. If the checker is systematically wrong, adding more instances of the same checker mostly multiplies the wrongness. It creates a beautiful row of confirmations over the same error.&lt;/p&gt;

&lt;p&gt;This is where the old infrastructure metaphor starts breaking, and where another metaphor becomes useful.&lt;/p&gt;

&lt;p&gt;The repair is not just retry.&lt;/p&gt;

&lt;p&gt;The repair is decorrelation.&lt;/p&gt;

&lt;p&gt;Different models. Different prompts. Different evaluation angles. Different representations of the same task. Different failure assumptions. Sometimes even different modalities of checking, where one component evaluates structure, another checks factual grounding, another verifies constraints, and another looks for contradiction.&lt;/p&gt;

&lt;p&gt;That is not classical retry anymore.&lt;/p&gt;

&lt;p&gt;That is closer to speciation.&lt;/p&gt;

&lt;p&gt;You do not make the system robust by repeating the same organism. You make it robust by introducing enough variation that one blind spot does not become a colony-wide disease. The same way biological systems survive not because every unit is perfect, but because diversity changes how failure propagates.&lt;/p&gt;

&lt;p&gt;This is the bridge I keep finding between old engineering and the older biological material I have been reading for years.&lt;/p&gt;

&lt;p&gt;I spent almost three years reading &lt;em&gt;Ant Encounters&lt;/em&gt;. Not because the book was impossible to read faster, but because every page kept connecting to something else. A small observation about ants was not just about ants anymore. It became a question about local decisions, global behavior, task allocation, distributed coordination, and how a system can stabilize without one central source of truth owning every decision.&lt;/p&gt;

&lt;p&gt;Now I am reading about viruses as complex adaptive systems, and the same thing happens. One page becomes a week of thinking. Adaptation, persistence, mutation, failure, survival under pressure, local variation, global behavior. Suddenly this is not only biology. It becomes a way to think about AI systems that cannot rely on perfect deterministic components but still need to produce reliable behavior.&lt;/p&gt;

&lt;p&gt;Compared with the rhythm of the AI world, this looks absurdly slow.&lt;/p&gt;

&lt;p&gt;People publish five takes about five new tools before lunch, and I am still stuck on one biological analogy from a book that is not even about AI.&lt;/p&gt;

&lt;p&gt;But maybe “stuck” is the wrong word.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;Maybe this is not reading. Maybe this is compilation.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The page is not being consumed. It is being linked. It enters a context made of old work, unfinished ideas, technical scars, software architecture, biological curiosity, and frustration with shallow AI products. The result is not speed. The result is compression. A small input creates a large internal reorganization.&lt;/p&gt;

&lt;p&gt;That is valuable, but it has a serious problem.&lt;/p&gt;

&lt;p&gt;It is invisible.&lt;/p&gt;

&lt;p&gt;And maybe this is the real fear behind the question “am I slow?” Not that I am actually slow. Not exactly. The fear is that while I am connecting dots, the field will move on without seeing any of it. The fear is that depth without visible output becomes indistinguishable from absence. The fear is that the AI world is so loud, so accelerated, and so addicted to fresh vocabulary that if I do not constantly produce something visible, I will disappear inside the buzz.&lt;/p&gt;

&lt;p&gt;That fear is not irrational.&lt;/p&gt;

&lt;p&gt;The field rewards fluency in the language of the week. If the word is “agents,” everyone builds agents. If the word is “reasoning,” everything becomes reasoning. If the word is “memory,” every cache becomes memory. If the word is “workflow,” every sequence of API calls becomes a platform.&lt;/p&gt;

&lt;p&gt;I understand why this happens. Attention is scarce. Timing matters. If you arrive too late, the conversation has already moved somewhere else.&lt;/p&gt;

&lt;p&gt;But there is a cost to always moving at that rhythm. You risk becoming synchronized with noise. You start optimizing for being current instead of being correct. You learn how to speak the new vocabulary faster than you understand the old problem underneath it. You become responsive, but not necessarily thoughtful.&lt;/p&gt;

&lt;p&gt;And I do not want that.&lt;/p&gt;

&lt;p&gt;At the same time, I cannot use depth as an excuse forever.&lt;/p&gt;

&lt;p&gt;This is the uncomfortable part, and I should not escape it with a nice sentence.&lt;/p&gt;

&lt;p&gt;Sometimes I am not slow. I am filtering. Sometimes I am not slow. I am connecting. Sometimes I am not slow. I am refusing to spend cognitive energy on hype that will disappear in two weeks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;But sometimes I am hiding behind depth.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not theoretically. Not as a universal writer problem. Me, now, in this exact pattern.&lt;/p&gt;

&lt;p&gt;I can feel it when a connection stays private longer than it should. I can feel it when a thought keeps becoming more complex in my head because publishing it would make it smaller, exposed, and easier to criticize. I can feel it when “I am still thinking about it” starts as discipline and slowly becomes shelter.&lt;/p&gt;

&lt;p&gt;That is the failure mode I need to catch.&lt;/p&gt;

&lt;p&gt;Because slow thinking is valuable only if it eventually becomes visible, testable, shareable, or executable. Otherwise it is just private complexity. It may feel profound internally, but from the outside it has no weight.&lt;/p&gt;

&lt;p&gt;This does not mean every thought needs to become a polished theory. That would create another paralysis. But the intermediate steps need to leave traces. The note after reading one page matters. The connection between ant encounters and AI routing matters. The observation that a model check is a handshake with uncertainty hidden under the return type matters. The idea that decorrelated evaluators are closer to speciation than retry matters.&lt;/p&gt;

&lt;p&gt;Not because each fragment is complete.&lt;/p&gt;

&lt;p&gt;Because the fragments show the work.&lt;/p&gt;

&lt;p&gt;That is probably the artifact I keep underestimating: the process of connecting old principles to new systems while the connection is still messy.&lt;/p&gt;

&lt;p&gt;The current AI discourse is full of people saying “look at this new thing.” Maybe there is space for someone saying “look at this old principle returning in a strange form, and look carefully at where the analogy breaks.”&lt;/p&gt;

&lt;p&gt;That is not slower.&lt;/p&gt;

&lt;p&gt;It is a different rhythm.&lt;/p&gt;

&lt;p&gt;A rhythm that does not compete well with hype in the short term, but maybe ages better.&lt;/p&gt;

&lt;p&gt;And maybe that is what I actually want. I do not want to win the weekly AI vocabulary race. I do not want to rebuild my thinking every time the market chooses a new word. I do not want to become another person who confuses speed with direction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;I want to understand what remains true after the buzzword moves away.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That kind of work is slower by nature. You cannot connect biology, distributed systems, software architecture, and AI orchestration at the speed of a product launch thread. You cannot build a durable mental model by reacting to every notification. You cannot understand a field only by consuming its newest claims.&lt;/p&gt;

&lt;p&gt;But you can disappear while doing deep work silently.&lt;/p&gt;

&lt;p&gt;That is the warning I take seriously.&lt;/p&gt;

&lt;p&gt;Not “you are too slow.”&lt;/p&gt;

&lt;p&gt;More like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;Your slowness needs output.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Slow and invisible is dangerous. Slow and traceable is different. Slow and executable is different. Slow and published is different. Slow and connected to experiments, code, diagrams, arguments, failures, and public reasoning becomes a body of work.&lt;/p&gt;

&lt;p&gt;So maybe the answer is yes, I am slow.&lt;/p&gt;

&lt;p&gt;But I am not slow because I am lost.&lt;/p&gt;

&lt;p&gt;I am slow because I am trying to understand the machinery instead of just repainting the dashboard. I am slow because every new AI idea seems to drag behind it a much older question. I am slow because I do not trust speed when speed is mostly social pressure. I am slow because I keep finding useful ghosts from older fields inside the newest buzzwords.&lt;/p&gt;

&lt;p&gt;The risk is not being slow.&lt;/p&gt;

&lt;p&gt;The risk is letting the work stay trapped inside my head until the world has no way to distinguish depth from silence.&lt;/p&gt;

&lt;p&gt;So I probably do not need to chase more.&lt;/p&gt;

&lt;p&gt;I need to expose more.&lt;/p&gt;

&lt;p&gt;I need to turn the reading into notes, the notes into arguments, the arguments into experiments, and the experiments into artifacts. Not perfectly. Not only when the whole theory is clean. Earlier. Messier. More honestly.&lt;/p&gt;

&lt;p&gt;Because maybe, in this AI world, reacting to every buzzword is not the same thing as adapting.&lt;/p&gt;

&lt;p&gt;Maybe the harder part is still being able to think when the buzzword is gone.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>career</category>
    </item>
    <item>
      <title>Orchestrated Multi-Agent Safety &amp; Test Oversight - AKA "`O MASTO"</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Mon, 18 May 2026 21:01:11 +0000</pubDate>
      <link>https://dev.to/marcosomma/orchestrated-multi-agent-safety-test-oversight-aka-o-masto-5hje</link>
      <guid>https://dev.to/marcosomma/orchestrated-multi-agent-safety-test-oversight-aka-o-masto-5hje</guid>
      <description>&lt;p&gt;I am building a small experiment inspired by Stripe Minions. Not related to OrKa. This is a different playground. But apparently I have a recurring problem: I do not trust agents enough to let them freely touch a codebase, and I do not trust humans enough to believe they will always review AI output properly when they are tired, rushed, or already late for another meeting.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;So the question became simple. Can we automate small development tasks without pretending the coding agent is the adult in the room?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because &lt;strong&gt;yes, AI can write code!&lt;/strong&gt; We know that. Sometimes it writes useful code. Sometimes it writes code that looks clean, passes the first glance, and then you realize it quietly moved business logic into the wrong layer because it had “a better idea.” Classic junior developer energy, but with infinite confidence and no coffee breaks.&lt;/p&gt;

&lt;p&gt;The interesting part of Stripe Minions, at least for me, is not that agents can open pull requests. The interesting part is the machinery around them. The task definition, the constraints, the review process, the checks, the fact that the agent is not just sitting there with a keyboard and divine permission to refactor your production system.&lt;/p&gt;

&lt;h3&gt;
  
  
  That is the part I want to explore!
&lt;/h3&gt;

&lt;p&gt;In my experiment, GitHub access starts as read-only. The system can inspect the codebase, understand structure, look at existing patterns, and generate a candidate issue. But it cannot immediately modify anything. Before planning even starts, the task needs to pass a semantic gate: is it scoped, is it testable, is it clear enough, and is it safe enough to continue? Only after that does the workflow move into planning, architecture, execution, PR creation, review, and final merge validation.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7u82xrm7bvfe1f47v9dj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7u82xrm7bvfe1f47v9dj.png" alt=" " width="800" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  I am calling this orchestration ’O MASTO.
&lt;/h3&gt;

&lt;p&gt;In Neapolitan, ’o masto is the master craftsman. The person who looks at the work and decides if it is actually good enough. Not if it looks good in a demo. Not if the agent says it is done. Actually good enough.&lt;/p&gt;

&lt;p&gt;In this experiment, it also stands for Orchestrated Multi-Agent Safety &amp;amp; Test Oversight. Yes, the acronym is a bit forced. No, I do not care. It makes me laugh, and naming things is half of software engineering anyway.&lt;/p&gt;

&lt;p&gt;The idea is that ’O MASTO is the layer that does not trust the agent. It checks the task, the plan, the implementation, the PR, the tests, the regression risk, and the final merge conditions. It is not there to be impressed. It is there to say “no, this is not good enough, go back.”&lt;/p&gt;

&lt;p&gt;That is the core idea I keep coming back to. The executor is not the boss. The reviewer is not the boss. The LLM is definitely not the boss. The gate is the boss!&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzvjaqi70x4cjqjy9bxah.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzvjaqi70x4cjqjy9bxah.png" alt=" " width="800" height="529"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I think this is where AI coding workflows need to go. Not toward bigger chat windows where we ask the model to “please be careful.” Toward systems that assume the model will be wrong sometimes and are designed to catch that before the damage reaches main.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;AI will not remove engineering discipline. It will expose who actually had it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If a project has no tests, no review culture, no stable patterns, no definition of done, and no clear ownership, an AI coding agent will not magically fix it. It will just produce chaos faster, with better formatting. &lt;/p&gt;

&lt;p&gt;My bet is that the next serious layer of AI development tooling will be trust infrastructure. Not just generation. Validation. Rejection. Retry. Traceability. Merge control. Basically "old school" software engineering.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>vibecoding</category>
      <category>sdd</category>
      <category>coding</category>
    </item>
    <item>
      <title>The Real Token Economy Is Not About Spending Less. It Is About Thinking Smaller.</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Sun, 26 Apr 2026 22:46:01 +0000</pubDate>
      <link>https://dev.to/marcosomma/the-real-token-economy-is-not-about-spending-less-it-is-about-thinking-smaller-3j3e</link>
      <guid>https://dev.to/marcosomma/the-real-token-economy-is-not-about-spending-less-it-is-about-thinking-smaller-3j3e</guid>
      <description>&lt;p&gt;I saw a video today that made me laugh, then made me a bit worried.&lt;/p&gt;

&lt;p&gt;It was one of those jokes that is not really a joke because you can already see some company doing it six months from now. A manager was basically complaining because an employee was not spending enough AI tokens. Not enough tokens. As if tokens were steps on a fitness tracker.&lt;/p&gt;

&lt;p&gt;"You only burned 2,000 tokens today, Susan. Are you even working?"&lt;/p&gt;

&lt;p&gt;It sounds absurd, but we are not that far from it. Companies are already starting to measure AI adoption through number of prompts, number of tool calls, input tokens, output tokens, cost per user, cost per team, cost per workflow. And to be clear, I do not think this is automatically wrong. Measuring token usage makes sense. Tokens are cost. Tokens are latency. Tokens are context. They are also a trace of how people and systems are using AI.&lt;/p&gt;

&lt;p&gt;The problem starts when we confuse the metric with the objective. We did this with hours worked. We did this with tickets closed. We did this with meetings attended. We did this with leads, where 1,000 unqualified leads looked better than 10 serious conversations because the spreadsheet was having a great day and nobody wanted to ruin the mood with reality.&lt;/p&gt;

&lt;p&gt;Now we risk doing the same with tokens.&lt;/p&gt;

&lt;p&gt;More tokens does not mean better work. Fewer tokens does not mean smarter work. The interesting signal is not the raw number. The interesting signal is the relationship between what you put into the model, what you ask it to do, and what comes out. That is where I think the real token economy starts. Not as a cost saving obsession, but as an architectural signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tokens are not just money
&lt;/h2&gt;

&lt;p&gt;The first way people talk about tokens is cost, and that is understandable. If you use hosted LLM APIs, tokens map quite directly to money. Input tokens cost something. Output tokens cost something. Larger models cost more. Long contexts cost more. Retries cost more. Bad prompts cost more. Bad architecture costs a lot more, but usually in a way that arrives later and looks like a reliability problem.&lt;/p&gt;

&lt;p&gt;So the first instinct is to optimize token consumption. Compress prompts. Summarize context. Pick cheaper models. Cache responses. Reduce unnecessary output. All of that is useful, but I think it is only the shallow layer of the problem.&lt;/p&gt;

&lt;p&gt;The more interesting question is not "how many tokens did this task consume?" The more interesting question is "what cognitive operation did those tokens represent?"&lt;/p&gt;

&lt;p&gt;Because input tokens and output tokens are not the same thing. Input tokens usually buy context. They are the material you ask the model to look at. Output tokens usually buy generation, explanation, structure, synthesis, or action. If I send 10,000 input tokens to a model and get back 10 output tokens, that could be terrible. It could also be exactly right.&lt;/p&gt;

&lt;p&gt;If the task is to read a long error log and return whether the failure is caused by authentication, a tiny output may be valid. If the task is to classify a product review as positive, neutral, or negative, a small answer is not a failure. It is the point. If the task is to route a bug report to the correct engineering queue, I do not need a novel. I need the right route.&lt;/p&gt;

&lt;p&gt;So no, high input and low output is not automatically bad. But it is a signal. And I think that signal deserves a lot more attention than it currently gets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Balance does not mean symmetry
&lt;/h2&gt;

&lt;p&gt;When I talk about token balance, I do not mean that input tokens and output tokens should be equal. That would be a very silly metric, and we already have enough silly metrics trying to cosplay as management science.&lt;/p&gt;

&lt;p&gt;By balance, I mean the relationship between the size of the input, the size of the output, and the value of the decision produced. A large input with a tiny output usually means the model is doing some kind of compression, classification, extraction, routing, filtering, moderation, scoring, validation, or decision making. A small input with a large output usually means the model is doing generation, expansion, explanation, drafting, or ideation. A large input with a large output usually means synthesis, transformation, summarization, comparison, or multi-step reasoning. A small input with a small output is usually a narrow atomic task.&lt;/p&gt;

&lt;p&gt;None of these patterns are good or bad by themselves. They tell you something about the shape of the work. And sometimes the shape of the work is screaming.&lt;/p&gt;

&lt;p&gt;Imagine you send a giant prompt containing a full meeting transcript, a product description, usage logs, a bug report, five examples, a JSON schema, tone guidelines, safety instructions, and a final line saying "be concise" because apparently we enjoy irony. Then you ask the model to return this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"priority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Maybe that is fine. Maybe the classification really required all of that context. But maybe you just built a cognitive washing machine to clean one spoon.&lt;/p&gt;

&lt;p&gt;The point is not that the token ratio is wrong. The point is that the ratio invites questions. Did the task need all of that context? Could the context have been retrieved more narrowly? Could the classification have been separated from the extraction? Could a smaller model do part of the work? Could a deterministic rule do part of it? Could the final output be validated separately instead of trusting one giant model call?&lt;/p&gt;

&lt;p&gt;That is where token metrics become useful. Not as a scoreboard. As a diagnostic tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real problem is overloaded cognition
&lt;/h2&gt;

&lt;p&gt;A lot of AI workflows are not expensive because the model is expensive. They are expensive because the task design is confused. We ask one model call to do too many things at once, then we act surprised when the model behaves like a very intelligent intern who received eight contradictory Jira tickets in one message.&lt;/p&gt;

&lt;p&gt;Read this long input. Understand the domain. Extract twenty fields. Normalize them. Infer missing values. Respect the schema. Apply business rules. Avoid hallucinations. Explain your decision. Be concise. Be deterministic. Also, please do it in one call because we saw a demo once and now we think architecture is a prompt template.&lt;/p&gt;

&lt;p&gt;This is where things become fragile. One big prompt. One big model. One fragile JSON output. One retry loop when it fails. One annoyed engineer staring at a malformed comma at 1:12 AM wondering why they studied data structures.&lt;/p&gt;

&lt;p&gt;The problem is not only cost. The problem is that the reasoning surface is too large. Every additional instruction increases the model's degrees of freedom. Every unrelated piece of context adds noise. Every extra output field increases the chance of format drift. Every hidden dependency between fields makes validation harder. And when the output fails, you often do not know why.&lt;/p&gt;

&lt;p&gt;Was the context too long? Was the instruction ambiguous? Was the schema too complex? Was the task logically overloaded? Was the model too weak? Was the model too creative? Was Mercury in retrograde? At some point, debugging a giant prompt starts to feel like debugging a dream.&lt;/p&gt;

&lt;p&gt;This is why I think the unit of optimization should not be the prompt. The unit of optimization should be the cognitive task.&lt;/p&gt;

&lt;h2&gt;
  
  
  Think smaller, not just cheaper
&lt;/h2&gt;

&lt;p&gt;When people hear "token economy", they often think about saving money. I think that is incomplete. The better version is this: design AI workflows so each model call has the smallest reasonable cognitive surface.&lt;/p&gt;

&lt;p&gt;Not the smallest prompt. Not the cheapest model. The smallest cognitive surface.&lt;/p&gt;

&lt;p&gt;A task has a cognitive surface when it asks the model to consider a certain amount of context, make a certain type of judgment, and produce a certain kind of output. A wide cognitive surface is something like this: read a conversation, infer the user's emotional state, detect all action items, classify the sales opportunity, extract objections, score urgency, summarize the call, generate a follow-up email, and return a perfect JSON object with 28 fields.&lt;/p&gt;

&lt;p&gt;That is not one task. That is a small village.&lt;/p&gt;

&lt;p&gt;A narrower cognitive task is different. Given this segment of a product feedback thread, identify whether the user mentions pricing as a blocker. Return true or false. Or extract only the next meeting date from this text and return null if absent. Or given these three already extracted signals, choose the priority level from low, medium, or high.&lt;/p&gt;

&lt;p&gt;Those tasks have narrower inputs and narrower outputs. They are easier to validate. They are easier to retry. They can often run on smaller models. Some can be replaced by deterministic code. Most importantly, they reduce ambiguity.&lt;/p&gt;

&lt;p&gt;This is the part that matters. The best token optimization is not always compression. Sometimes the best token optimization is decomposition.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 20 field JSON problem
&lt;/h2&gt;

&lt;p&gt;Let us take a simple example. You have a large input document and you need a structured output with 20 values. The obvious modern AI approach is to send the full document to a model and ask it to extract everything in one JSON object. Add a schema, add "do not hallucinate", add "use null when unknown", maybe add three examples, and hope the model behaves.&lt;/p&gt;

&lt;p&gt;Sometimes this works. Sometimes it works very well in the demo. Then production arrives, wearing boots.&lt;/p&gt;

&lt;p&gt;The model misses a field. It invents a value. It mixes two fields. It returns invalid JSON. It follows the schema but puts the wrong value in the right place, which is worse because it looks correct. It explains itself inside a field because apparently JSON needed feelings.&lt;/p&gt;

&lt;p&gt;So you add more instructions. Then stricter schema language. Then validation. Then retry. Then a stronger model. Then a more expensive model. Then someone says, "Maybe we should fine-tune it." And now your simple extraction pipeline has become a small national infrastructure project.&lt;/p&gt;

&lt;p&gt;A different approach is to ask a boring but useful question: are these 20 values actually one cognitive task?&lt;/p&gt;

&lt;p&gt;Maybe not. Maybe five fields are direct extraction. Maybe three require classification. Maybe four depend on dates. Maybe two require numerical normalization. Maybe six are only relevant if a previous condition is true. In that case, one big prompt is not simpler. It is only hiding the complexity inside the model call.&lt;/p&gt;

&lt;p&gt;You may get a better system by clustering the fields by semantic dependency. For example, direct identifiers can be one batch. Dates and temporal constraints can be another. Risk indicators can be another. Obligations and responsible parties can be another. The final normalized summary can be built only after the previous signals exist.&lt;/p&gt;

&lt;p&gt;Each batch can have a smaller prompt, a smaller schema, and a narrower validation rule. Some batches may not need an LLM. Some can use regex, parsers, lookup tables, embeddings, or deterministic checks. Some can use a small local model. Only the genuinely difficult parts need the expensive model.&lt;/p&gt;

&lt;p&gt;This is where the cost savings come from, but cost is only one part of the win. You also get better observability. If the final output is wrong, you can inspect which subtask failed. You can measure field-level accuracy. You can retry only the failing part. You can swap models for one stage without touching the rest. You can cache intermediate outputs. You can add deterministic validation at the boundary.&lt;/p&gt;

&lt;p&gt;That is a real token economy. Not "use fewer tokens". Spend tokens where cognition is actually needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Smaller prompts reduce variance
&lt;/h2&gt;

&lt;p&gt;I want to be careful with the word deterministic. LLMs are not truly deterministic systems in the classical engineering sense, even when you reduce temperature and constrain output. They are probabilistic systems. But workflow design can make their behavior more stable, more reproducible, and more controllable.&lt;/p&gt;

&lt;p&gt;Smaller prompts with narrower objectives usually reduce the degrees of freedom of the model. If the model has one job, a small output space, and a strict schema, there are fewer ways to fail. If the model has twenty jobs, a large input, competing instructions, implicit dependencies, and a complex schema, you should not be surprised when it occasionally decides to express itself like a haunted spreadsheet.&lt;/p&gt;

&lt;p&gt;This is why task decomposition can improve consistency. Not because small calls magically make the model deterministic, but because small calls make the system around the model easier to control. The output space is narrower. The validation is simpler. The retry logic is cheaper. The failure modes are easier to classify. The model choice becomes more flexible. The prompts become easier to test.&lt;/p&gt;

&lt;p&gt;And the orchestration becomes explicit.&lt;/p&gt;

&lt;p&gt;That last point matters a lot. When everything happens inside one prompt, the process is invisible. When you split the work into stages, the process becomes inspectable. This is the difference between hoping the model thinks correctly and designing a system where each step can be observed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where OrKa fits into this
&lt;/h2&gt;

&lt;p&gt;This is one of the reasons I have been building &lt;a href="https://orkacore.com" rel="noopener noreferrer"&gt;OrKa&lt;/a&gt;, an orchestration framework for AI agents and reasoning workflows. The point of OrKa is not "use more agents because agents are cool". Honestly, if adding agents makes your system less understandable, congratulations, you have invented distributed confusion.&lt;/p&gt;

&lt;p&gt;The point is different. Make cognitive work explicit. Define the flow. Split reasoning into smaller units. Route tasks. Log execution. Validate outputs. Keep memory and context under control. Make the system inspectable instead of praying over a large prompt.&lt;/p&gt;

&lt;p&gt;In this view, an LLM is not the application. It is one component inside a system. Sometimes the LLM extracts. Sometimes it classifies. Sometimes it rewrites. Sometimes it evaluates. Sometimes it should not be called at all. The orchestration layer decides how work moves between these pieces.&lt;/p&gt;

&lt;p&gt;That is where token economy becomes architecture. You are no longer asking only how to reduce a prompt by 20 percent. You are asking which cognitive step actually needs this context.&lt;/p&gt;

&lt;p&gt;That question changes everything. Maybe the first model call only needs the user message. Maybe the second needs the relevant log snippet. Maybe the third needs only three extracted fields. Maybe the final formatter needs no model at all. If you send the full context to every step, you are not designing an AI system. You are photocopying the universe and asking a model to find the invoice number.&lt;/p&gt;

&lt;p&gt;It may work. It is not a strategy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token metrics should trigger questions
&lt;/h2&gt;

&lt;p&gt;So how should teams use token metrics? Not as productivity surveillance. Not as a way to shame people for using too many or too few tokens. Not as a leaderboard where the person with the most prompts wins some cursed office trophy.&lt;/p&gt;

&lt;p&gt;Token metrics should trigger engineering questions.&lt;/p&gt;

&lt;p&gt;When input tokens are very high and output tokens are very low, ask whether the task is intentionally compressive or accidentally overloaded. When output tokens are very high, ask whether the model is generating useful structure or just producing expensive fog. When the same context is repeatedly sent across multiple calls, ask whether retrieval, caching, or state passing could reduce duplication. When a large model is used for simple extraction, ask whether a smaller model or deterministic rule would work. When retries consume a lot of tokens, ask whether the schema, validation, or task boundaries are wrong.&lt;/p&gt;

&lt;p&gt;This does not mean splitting tasks is automatically better. If you send the same 10,000 token input twenty times to extract twenty fields, you may have made the system more expensive and slower. You have not built architecture. You have built a very complicated way to duplicate context.&lt;/p&gt;

&lt;p&gt;The win comes when decomposition is paired with context narrowing. Extract the relevant segment once. Reuse intermediate state. Cluster fields that share dependencies. Route only the necessary context. Validate locally. Use smaller models where possible. Stop calling the model when code can do the job.&lt;/p&gt;

&lt;p&gt;This is not anti-LLM. It is pro-system.&lt;/p&gt;

&lt;h2&gt;
  
  
  A simple mental model
&lt;/h2&gt;

&lt;p&gt;Here is the mental model I keep coming back to. Input tokens are attention budget. Output tokens are commitment surface.&lt;/p&gt;

&lt;p&gt;The more input you provide, the more the model has to attend to. The more output you request, the more opportunities the model has to drift. A workflow becomes more stable when the attention budget and the commitment surface are aligned with the actual cognitive task.&lt;/p&gt;

&lt;p&gt;If the model needs to classify one thing, do not ask it to also summarize, extract, explain, normalize, and format a complex object. If the model needs to generate a long answer, do not overload it with irrelevant context that only increases noise. If the model needs to extract structured fields, do not assume all fields belong in the same call. If the model needs to make a decision, make the decision boundary explicit.&lt;/p&gt;

&lt;p&gt;The goal is not minimal tokens. The goal is minimal unnecessary cognition.&lt;/p&gt;

&lt;p&gt;That distinction is important. Some tasks deserve many tokens. A long research synthesis may need a lot of context. A technical incident summary may need careful source retention. A product comparison may need long input and long output. A multi-document comparison may be legitimately expensive.&lt;/p&gt;

&lt;p&gt;The problem is not spending tokens. The problem is spending tokens without knowing what they are buying.&lt;/p&gt;

&lt;h2&gt;
  
  
  This is also a model selection problem
&lt;/h2&gt;

&lt;p&gt;Once you split cognitive tasks, model selection becomes much more interesting. In a one-prompt architecture, you usually choose the strongest model you can afford because the task is messy. The model has to handle everything. It has to read long context, reason, extract, format, validate, and recover from ambiguity.&lt;/p&gt;

&lt;p&gt;But if you split the workflow, you can choose models per cognitive step. A small model can do simple classification. A local model can extract obvious fields. A deterministic parser can normalize dates. A rules engine can validate constraints. A stronger model can handle the genuinely ambiguous reasoning.&lt;/p&gt;

&lt;p&gt;This is where the economics change. Not because you begged the prompt to be shorter, but because you changed the shape of the work. The expensive model becomes a specialist instead of a landfill.&lt;/p&gt;

&lt;p&gt;And yes, I know "landfill" sounds harsh. But many AI systems today are exactly that. They throw all context into one place and hope the biggest model will recycle it into something useful. This works surprisingly often, which is the dangerous part. It works enough to ship a demo. It fails enough to punish you in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token economy as observability
&lt;/h2&gt;

&lt;p&gt;A mature AI system should not only log the final response. It should log the token shape of the workflow.&lt;/p&gt;

&lt;p&gt;Which step consumed the most input? Which step produced the most output? Which step retried the most? Which step had the most schema failures? Which step required the strongest model? Which step could be cached? Which step could be replaced by code? Which step actually improved the final decision?&lt;/p&gt;

&lt;p&gt;This is not accounting. This is observability.&lt;/p&gt;

&lt;p&gt;You are not only tracking spend. You are tracking cognitive pressure inside the system. A sudden increase in input tokens may mean your retrieval is bringing too much context. A sudden increase in output tokens may mean the model started explaining instead of structuring. A high retry cost may mean your schema is too complex or your prompt is ambiguous. A high token cost on a low-value decision may mean the workflow needs decomposition. A low token cost with poor quality may mean you compressed away necessary context.&lt;/p&gt;

&lt;p&gt;Again, the metric is not the answer. The metric is the signal. The engineer still needs judgment, which is very inconvenient. We were promised automation and somehow we still need thinking. Rude.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wrong future
&lt;/h2&gt;

&lt;p&gt;The wrong future is easy to imagine. Teams get AI dashboards. Managers see token usage per employee. People are encouraged to "use AI more". Token consumption becomes proof of adoption. Employees learn to generate more prompts because the dashboard rewards activity. Everyone looks productive, costs go up, and quality does not.&lt;/p&gt;

&lt;p&gt;Then leadership announces an AI efficiency initiative. Now everyone must reduce token usage. People use smaller prompts. Quality drops. Nobody knows why. Another dashboard is created. A consultant appears. The circle of life continues.&lt;/p&gt;

&lt;p&gt;This is what happens when the metric becomes the goal. Token usage by itself tells you almost nothing about quality. A great engineer may use fewer tokens because they decomposed the problem properly. Another great engineer may use more tokens because the task genuinely required context. A bad workflow may use few tokens and produce garbage. A good workflow may use many tokens and produce a high-value decision.&lt;/p&gt;

&lt;p&gt;So measuring tokens is not wrong. Judging work by raw token volume is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The better future
&lt;/h2&gt;

&lt;p&gt;The better future is more boring, which is usually a good sign in engineering. Teams treat token metrics as workflow diagnostics. They look at input-output patterns. They identify overloaded prompts. They split tasks where it makes sense. They route context more carefully. They use smaller models for smaller cognitive jobs. They validate structured outputs separately. They measure retries, drift, and failure modes.&lt;/p&gt;

&lt;p&gt;They do not ask "how much AI did you use?" They ask "where did the AI actually add decision value?"&lt;/p&gt;

&lt;p&gt;That is the mindset shift. Tokens are not just a bill. Tokens are a trace of cognitive architecture. They show where a system is bloated. They show where context is duplicated. They show where outputs are too ambitious. They show where models are being used as glue because nobody wanted to design the pipeline. And yes, sometimes they show that the expensive model was actually justified.&lt;/p&gt;

&lt;p&gt;That is fine. The goal is not to make everything cheap. The goal is to make the system honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thought
&lt;/h2&gt;

&lt;p&gt;I think the next stage of AI engineering will not be about who writes the cleverest prompt. It will be about who designs the clearest cognitive pipeline.&lt;/p&gt;

&lt;p&gt;The prompt is not the unit of architecture. The cognitive task is.&lt;/p&gt;

&lt;p&gt;Token balance matters because it gives us a way to inspect that task. Not perfectly. Not automatically. Not as a KPI to punish or reward people. But as a signal that says: maybe this workflow is overloaded, maybe this context is too broad, maybe this output is trying to do too much, maybe this model is stronger than necessary, maybe this task should be split, maybe this step should not use an LLM at all.&lt;/p&gt;

&lt;p&gt;That is where the real token economy lives. Not in spending fewer tokens, but in spending attention where attention is needed.&lt;/p&gt;

&lt;p&gt;If we do that well, the benefits go beyond cost. Lower latency. Smaller models. Cleaner validation. Less format drift. More stable outputs. More inspectable workflows. Systems that are closer to engineering and less close to whispering wishes into a very expensive autocomplete machine.&lt;/p&gt;

&lt;p&gt;Which, to be fair, is still fun.&lt;/p&gt;

&lt;p&gt;But maybe not the future we should build production systems on.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude! Stop Burning Tokens on Your Agent's Tool Output!</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Tue, 21 Apr 2026 09:39:42 +0000</pubDate>
      <link>https://dev.to/marcosomma/claude-stop-burning-tokens-on-your-agents-tool-output-1cpl</link>
      <guid>https://dev.to/marcosomma/claude-stop-burning-tokens-on-your-agents-tool-output-1cpl</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;A Two-Stage Curator That Pays for Itself&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I watched Claude Code feed &lt;strong&gt;108,894 bytes&lt;/strong&gt; of &lt;code&gt;seq 1 20000&lt;/code&gt; back into its own context window. That output contained 20,000 integers.&lt;br&gt;
No errors. No signal. No insight. Just counting.&lt;/p&gt;

&lt;p&gt;And yet the system still had to tokenize it, send it back to the model, and bill for it. This is not an edge case. It is the default failure mode of agent tooling.&lt;/p&gt;

&lt;p&gt;Tools produce output. The output goes back to the model. You pay for it. Logs, test runs, &lt;code&gt;ps&lt;/code&gt; listings, git history, build spam, progress bars, boilerplate, decorative separators, repeated warnings, repeated success lines, repeated everything. A depressing amount of it is useless.&lt;/p&gt;

&lt;p&gt;A lot of agent users are quietly paying premium model rates to process terminal confetti. That is the real problem!&lt;/p&gt;

&lt;p&gt;My first fix was the obvious one. I added a &lt;code&gt;PreToolUse&lt;/code&gt; hook and pushed large Bash output through a cheaper model before it reached Opus.&lt;/p&gt;

&lt;p&gt;That worked, technically. Then I noticed I was still being stupid, just in a more optimized way.&lt;/p&gt;

&lt;p&gt;On &lt;code&gt;seq 1 20000&lt;/code&gt;, I was paying Haiku to read 20,000 integers and tell me they were 20,000 sequential integers.&lt;/p&gt;

&lt;p&gt;Yes, that is cheaper than letting Opus read them.&lt;/p&gt;

&lt;p&gt;No, that is not a good design.&lt;/p&gt;

&lt;p&gt;If a 40-line &lt;code&gt;awk&lt;/code&gt; script can identify the pattern for free, paying any model to summarize it is already waste. So the architecture changed. Not “small model before big model.” That idea sounds clever, but by itself it is just cost reshuffling.&lt;/p&gt;

&lt;p&gt;The real pattern is simpler and better: &lt;strong&gt;extract free signal first, and only pay for a model when deterministic tools run out of leverage.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That led to a two-stage curator.&lt;/p&gt;

&lt;p&gt;Stage 1 is deterministic cleanup. It strips ANSI escape sequences, removes carriage-return junk from progress bars, collapses consecutive duplicate lines, and compresses monotonic integer runs. It is effectively free.&lt;/p&gt;

&lt;p&gt;Stage 2 is LLM extraction, but only if stage 1 still leaves too much output. That is where tokens are spent. That means it should fire rarely, and only when stage 1 could not do enough.&lt;/p&gt;

&lt;p&gt;That distinction matters! Because once you actually measure it, a lot of tool output turns out to be compressible by embarrassingly simple logic.&lt;/p&gt;
&lt;h2&gt;
  
  
  The benchmark
&lt;/h2&gt;

&lt;p&gt;Here is the benchmark across five scenarios, using the pricing assumptions from the benchmark runner: Opus input at $15 per million tokens, Haiku input at $1 per million, Haiku output at $5 per million, and a rough estimate of 4 bytes per token.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Raw bytes&lt;/th&gt;
&lt;th&gt;Stage 1&lt;/th&gt;
&lt;th&gt;Final&lt;/th&gt;
&lt;th&gt;LLM?&lt;/th&gt;
&lt;th&gt;Tokens saved&lt;/th&gt;
&lt;th&gt;Haiku cost&lt;/th&gt;
&lt;th&gt;Net savings&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;seq 1 20000&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;108,894&lt;/td&gt;
&lt;td&gt;37&lt;/td&gt;
&lt;td&gt;103&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;27,197&lt;/td&gt;
&lt;td&gt;$0.000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+$0.408&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5,000 repeated log lines&lt;/td&gt;
&lt;td&gt;145,000&lt;/td&gt;
&lt;td&gt;38&lt;/td&gt;
&lt;td&gt;104&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;36,224&lt;/td&gt;
&lt;td&gt;$0.000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+$0.543&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ANSI + progress-bar spam&lt;/td&gt;
&lt;td&gt;54,000&lt;/td&gt;
&lt;td&gt;27&lt;/td&gt;
&lt;td&gt;92&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;13,477&lt;/td&gt;
&lt;td&gt;$0.000&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+$0.202&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;ps auxww&lt;/code&gt; (unique lines)&lt;/td&gt;
&lt;td&gt;213,230&lt;/td&gt;
&lt;td&gt;213,230&lt;/td&gt;
&lt;td&gt;1,008&lt;/td&gt;
&lt;td&gt;yes&lt;/td&gt;
&lt;td&gt;53,055&lt;/td&gt;
&lt;td&gt;$0.055&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+$0.741&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;echo hello world&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;$0.000&lt;/td&gt;
&lt;td&gt;$0.000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is blunt.&lt;/p&gt;

&lt;p&gt;Three of the four large-output cases were handled for free.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;seq&lt;/code&gt; collapsed from 108,894 bytes to 37.&lt;br&gt;
Repeated log spam dropped from 145,000 to 38.&lt;br&gt;
ANSI and progress-bar noise fell from 54,000 to 27.&lt;/p&gt;

&lt;p&gt;No LLM call was needed in any of those cases.&lt;/p&gt;

&lt;p&gt;The only scenario that needed stage 2 was &lt;code&gt;ps auxww&lt;/code&gt;, which is exactly what you would want. That output is genuinely varied. There is not much for &lt;code&gt;awk&lt;/code&gt; to compress. That is the moment when paying a smaller model to extract the useful facts is justified.&lt;/p&gt;

&lt;p&gt;Small output remained untouched, which is also correct. If the command only produced &lt;code&gt;hello world&lt;/code&gt;, the system paid nothing and moved on.&lt;/p&gt;

&lt;p&gt;This is the whole point.&lt;/p&gt;

&lt;p&gt;A cheap LLM is not the first line of defense.&lt;/p&gt;

&lt;p&gt;Deterministic cleanup is.&lt;/p&gt;

&lt;p&gt;The LLM should be the escalation path, not the reflex.&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 1: deterministic cleanup
&lt;/h2&gt;

&lt;p&gt;This is the part that does the real work more often than people expect.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight awk"&gt;&lt;code&gt;&lt;span class="c1"&gt;#!/usr/bin/awk -f&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;flush_int_run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;int_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;printf&lt;/span&gt; &lt;span class="s2"&gt;"[%d sequential integers %s..%s]\n"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;int_count&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;int_start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;int_end&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;int_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;int_start&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;int_end&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;print&lt;/span&gt; &lt;span class="nx"&gt;i&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nx"&gt;int_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;flush_dupe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;dupe_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;printf&lt;/span&gt; &lt;span class="s2"&gt;"%s [×%d]\n"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dupe_line&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;dupe_count&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;dupe_count&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;print&lt;/span&gt; &lt;span class="nx"&gt;dupe_line&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nx"&gt;dupe_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nb"&gt;gsub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\0&lt;/span&gt;&lt;span class="sr"&gt;33&lt;/span&gt;&lt;span class="se"&gt;\[[&lt;/span&gt;&lt;span class="sr"&gt;0-9;&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;*&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;a-zA-Z&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nb"&gt;gsub&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\r&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$0&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt; &lt;span class="sr"&gt;/^-&lt;/span&gt;&lt;span class="se"&gt;?[&lt;/span&gt;&lt;span class="sr"&gt;0-9&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+$/&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;int_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nx"&gt;int_end&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="nx"&gt;int_end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt;
            &lt;span class="nx"&gt;int_count&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;
            &lt;span class="k"&gt;next&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="nx"&gt;flush_int_run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nx"&gt;flush_dupe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nx"&gt;int_start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt;
        &lt;span class="nx"&gt;int_end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;n&lt;/span&gt;
        &lt;span class="nx"&gt;int_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="k"&gt;next&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="nx"&gt;flush_int_run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;dupe_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="nx"&gt;dupe_line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;dupe_count&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;flush_dupe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="nx"&gt;dupe_line&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nv"&gt;$0&lt;/span&gt;
        &lt;span class="nx"&gt;dupe_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kr"&gt;END&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;flush_int_run&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="nx"&gt;flush_dupe&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is nothing magical here.&lt;/p&gt;

&lt;p&gt;It strips terminal paint.&lt;br&gt;
It collapses repeated lines.&lt;br&gt;
It compresses obvious integer sequences.&lt;/p&gt;

&lt;p&gt;That is enough to destroy huge amounts of waste.&lt;/p&gt;

&lt;p&gt;This is a useful reminder for AI tooling in general. A lot of expensive “reasoning” problems are not reasoning problems. They are preprocessing failures.&lt;/p&gt;
&lt;h2&gt;
  
  
  Stage 2: only escalate if stage 1 failed to shrink enough
&lt;/h2&gt;

&lt;p&gt;The wrapper runs the command, checks the raw size, and passes small output through untouched.&lt;/p&gt;

&lt;p&gt;If the raw output is large, it runs the deterministic cleaner.&lt;/p&gt;

&lt;p&gt;If the cleaned output is now small enough, it returns that cleaned output and stops.&lt;/p&gt;

&lt;p&gt;Only if the cleaned output is still large does it call Haiku.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; pipefail

&lt;span class="nv"&gt;RAW_THRESHOLD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CLAUDE_BASH_SUMMARIZE_THRESHOLD&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;8000&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;LLM_THRESHOLD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CLAUDE_BASH_LLM_THRESHOLD&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;8000&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CLAUDE_BASH_SUMMARIZE_MODEL&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;claude&lt;/span&gt;&lt;span class="p"&gt;-haiku-4-5&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

&lt;span class="nv"&gt;cmd&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;SCRIPT_DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;dirname&lt;/span&gt; &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;BASH_SOURCE&lt;/span&gt;&lt;span class="p"&gt;[0]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &amp;amp;&amp;gt;/dev/null &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;pwd&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;CLEAN_AWK&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SCRIPT_DIR&lt;/span&gt;&lt;span class="s2"&gt;/deterministic-clean.awk"&lt;/span&gt;

&lt;span class="nv"&gt;raw&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;cleaned&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;trap&lt;/span&gt; &lt;span class="s1"&gt;'rm -f "$raw" "$cleaned"'&lt;/span&gt; EXIT

bash &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cmd&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; 2&amp;gt;&amp;amp;1
&lt;span class="nv"&gt;rc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$?&lt;/span&gt;

&lt;span class="nv"&gt;raw_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &amp;lt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;' '&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-le&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$RAW_THRESHOLD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;exit&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$rc&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CLEAN_AWK&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &amp;lt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;cleaned_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &amp;lt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;' '&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-le&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$LLM_THRESHOLD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'=== CURATED %d→%d bytes (stage 1 deterministic, no LLM) ===\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;exit&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$rc&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nv"&gt;summary&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="nt"&gt;--model&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"Extract signal from this command output. KEEP: errors, warnings, stack traces, file paths with line numbers, numeric results, unique events, final status. DROP: decorative separators, boilerplate. Preserve exact error text verbatim. Be terse but faithful on key facts. Plain text only."&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &amp;lt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; 2&amp;gt;/dev/null&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="nv"&gt;summary_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;${#&lt;/span&gt;&lt;span class="nv"&gt;summary&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;
&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'=== CURATED %d→%d→%d bytes (stage 1 + %s extraction) ===\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$raw_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$cleaned_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$summary_size&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%s\n'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$summary&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nb"&gt;exit&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$rc&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That second threshold check is the entire point.&lt;/p&gt;

&lt;p&gt;Without it, the “cheap model” becomes a permanent tax on output that deterministic logic had already made cheap.&lt;/p&gt;

&lt;p&gt;With it, the LLM only gets called when the dumb tools genuinely ran out of leverage.&lt;/p&gt;

&lt;p&gt;That is how this stops being a cute hack and starts becoming a sensible pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hooking it into Claude Code
&lt;/h2&gt;

&lt;p&gt;I wired it into Claude Code with a &lt;code&gt;PreToolUse&lt;/code&gt; hook on &lt;code&gt;Bash&lt;/code&gt;. The hook rewrites the original command so it runs through the wrapper first.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PreToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/path/to/.claude/scripts/bash-wrap-hook.sh"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"timeout"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hook script itself is tiny. It reads the incoming JSON, extracts &lt;code&gt;tool_input.command&lt;/code&gt;, and swaps in the wrapped version. Nothing here is conceptually hard. That is precisely why it is worth doing. Too much agent engineering right now is really just people tolerating waste because it looks sophisticated once wrapped in model calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost math
&lt;/h2&gt;

&lt;p&gt;Here is the clean version. Let the raw output be &lt;code&gt;N&lt;/code&gt; tokens.&lt;br&gt;
If you send it directly to Opus, the input cost is:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;15N / 1,000,000&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Now suppose stage 1 reduces that output to &lt;code&gt;K&lt;/code&gt; tokens. If &lt;code&gt;K&lt;/code&gt; is still above threshold, stage 2 fires. Haiku reads &lt;code&gt;K&lt;/code&gt; tokens, emits a summary of &lt;code&gt;M&lt;/code&gt; tokens, and then Opus receives those &lt;code&gt;M&lt;/code&gt; tokens.&lt;br&gt;
So the escalated path costs:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;(1K + 5M + 15M) / 1,000,000 = (K + 20M) / 1,000,000&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Break-even is therefore:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;15N &amp;gt; K + 20M&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That is the actual condition for the two-stage system. If stage 1 barely helps, then &lt;code&gt;K ≈ N&lt;/code&gt;, and the inequality becomes:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;15N &amp;gt; N + 20M&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;which simplifies to:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;14N &amp;gt; 20M&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;M &amp;lt; 0.7N&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;So in the worst case, where deterministic cleanup did almost nothing, the LLM stage still pays off if it compresses the remaining content by roughly 1.43x or better. That is not a demanding threshold.&lt;/p&gt;

&lt;p&gt;And in the real pipeline, stage 1 often shrinks the input before the LLM ever sees it, which makes the economics even more comfortable.&lt;br&gt;
This is why the design works. Not because “small model then big model” is automatically clever. Because the model is only invited in after cheap tools have already done what they can.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would change next
&lt;/h2&gt;

&lt;p&gt;This version already works, but there are obvious next steps.&lt;/p&gt;

&lt;p&gt;Latency should probably be part of the gate, not just bytes. Saving a few cents is not impressive if it adds a few seconds to every interactive tool call.&lt;/p&gt;

&lt;p&gt;Stage 1 could be extended to catch more patterns, especially timestamp-heavy logs where the message repeats but the prefix changes.&lt;/p&gt;

&lt;p&gt;The system should probably hard-cap verbose LLM summaries, because “cheap extraction” can still become noisy if the prompt drifts.&lt;/p&gt;

&lt;p&gt;And the current implementation buffers until the command finishes. That is fine for benchmarking, but worse for real long-running workflows. A streaming version would be much better.&lt;/p&gt;

&lt;p&gt;But none of that changes the core lesson.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual lesson
&lt;/h2&gt;

&lt;p&gt;The interesting idea here is not “use an LLM to compress LLM inputs.” That is the shallow reading. The more useful pattern is this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;before you spend tokens to extract signal, try extracting signal for free.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A lot of the mess we hand to expensive models is not difficult. It is just noisy. And noisy is not the same as complex. That distinction matters. Because once you see it clearly, the same pattern starts showing up everywhere.&lt;/p&gt;

&lt;p&gt;Retrieval pipelines that rerank garbage before filtering it.&lt;br&gt;
Scrapers that pass repeated boilerplate into embeddings.&lt;br&gt;
Log processors that ask a model to summarize progress-bar sludge.&lt;br&gt;
Agent systems that burn premium tokens on output a shell one-liner could have collapsed immediately.&lt;/p&gt;

&lt;p&gt;Cheap filter first.&lt;br&gt;
Expensive model second.&lt;br&gt;
Measure the break-even.&lt;br&gt;
Then stop paying premium rates for repetition, boilerplate, terminal paint, and counting.&lt;/p&gt;

&lt;p&gt;That is not an AI breakthrough.&lt;/p&gt;

&lt;p&gt;It is just basic engineering discipline, which is exactly why so many agent stacks are currently missing it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>tooling</category>
      <category>mcp</category>
    </item>
    <item>
      <title>I Ran 500 More Agent Memory Experiments. The Real Problem Wasn’t Recall. It Was Binding.</title>
      <dc:creator>marcosomma</dc:creator>
      <pubDate>Mon, 13 Apr 2026 09:19:39 +0000</pubDate>
      <link>https://dev.to/marcosomma/i-ran-500-more-agent-memory-experiments-the-real-problem-wasnt-recall-it-was-binding-24kc</link>
      <guid>https://dev.to/marcosomma/i-ran-500-more-agent-memory-experiments-the-real-problem-wasnt-recall-it-was-binding-24kc</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a follow-up to &lt;a href="https://dev.to/marco_somma_a9e88a3063f3/i-tried-to-turn-agent-memory-into-plumbing-instead-of-philosophy-1bpm"&gt;I Tried to Turn Agent Memory Into Plumbing Instead of Philosophy&lt;/a&gt;. If you haven't read that one, the short version: I built a persistent memory system for AI agents called &lt;a href="https://github.com/marcosomma/orka-reasoning" rel="noopener noreferrer"&gt;OrKa Brain&lt;/a&gt;, ran 30 benchmark tasks, got a 63% pairwise win rate and a +0.10 rubric improvement, and concluded that "the model already knew most of what the Brain was recalling." Then I got some very good comments that made me uncomfortable. This is what happened next.&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Comfortable Lie I Told Myself
&lt;/h2&gt;

&lt;p&gt;After the first benchmark, I had a narrative that felt reasonable: the memory system works, the numbers are positive, the confounds are acknowledged, and more data will clarify things.&lt;/p&gt;

&lt;p&gt;That last part, "more data will clarify things", is what engineers say when they don't want to admit they might be wrong. I said it too. And then I went and got more data.&lt;/p&gt;

&lt;p&gt;250 tasks. Five specialized tracks. 500 total runs (brain vs. brainless). A separate judge model so the LLM wasn't grading its own homework. Eleven code changes addressing five root-cause problems I'd identified from the first round.&lt;/p&gt;

&lt;p&gt;The results came back. They didn't clarify things. They made them worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Fixed Before Running Again
&lt;/h2&gt;

&lt;p&gt;I'm not going to pretend I just blindly re-ran the same experiment. I did real work between benchmark v1 and v2. The &lt;a href="https://dev.to/marco_somma_a9e88a3063f3/i-tried-to-turn-agent-memory-into-plumbing-instead-of-philosophy-1bpm#comments"&gt;first article's comments&lt;/a&gt; called out several things, and I addressed them:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 1: Skills were storing verbatim LLM output, not abstract patterns.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This was the big one. When the Brain learned a skill from a data engineering task, it stored the literal steps: "Load CSV files into staging tables using pandas read_csv with error handling." That's not transferable knowledge, it's a paraphrase of what the model already knows. I rewrote the abstraction layer (&lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/brain/constants.py" rel="noopener noreferrer"&gt;&lt;code&gt;orka/brain/constants.py&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/brain/brain.py" rel="noopener noreferrer"&gt;&lt;code&gt;brain.py&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/agents/brain_agent.py" rel="noopener noreferrer"&gt;&lt;code&gt;brain_agent.py&lt;/code&gt;&lt;/a&gt;) to extract verb-target patterns: "implement [target]", "validate [component]", "trace [target]". The idea was that abstract patterns would transfer better across domains.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 2: The recall threshold was zero.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;min_score=0.0&lt;/code&gt; meant any vaguely related skill could get recalled. I raised it to 0.5 and added a semantic floor in the &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/brain/transfer_engine.py" rel="noopener noreferrer"&gt;&lt;code&gt;transfer_engine.py&lt;/code&gt;&lt;/a&gt;, if the embedding similarity is below 0.1 AND structural match is below 0.6, the candidate gets rejected entirely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 3: The model was judging its own output.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;v1 used the same LLM for execution and evaluation. v2 uses a separate judge model (&lt;code&gt;qwen/qwen3-coder-30b&lt;/code&gt;) with dedicated &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/judge_rubric_workflow.yml" rel="noopener noreferrer"&gt;rubric&lt;/a&gt; and &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/judge_pairwise_workflow.yml" rel="noopener noreferrer"&gt;pairwise&lt;/a&gt; workflow YAMLs. Execution and judgment are completely decoupled, different scripts, different models, different runs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 4: Track diversity.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;v1 had one track. v2 has five:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Track&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;th&gt;Why It Matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;Cross-domain transfer&lt;/td&gt;
&lt;td&gt;Does a data engineering skill help with cybersecurity?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;Ethical reasoning&lt;/td&gt;
&lt;td&gt;Do anti-pattern detection skills transfer?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;Routing decisions&lt;/td&gt;
&lt;td&gt;Hardest track, complex multi-path choices&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;td&gt;Multi-step reasoning&lt;/td&gt;
&lt;td&gt;Do procedural patterns help new reasoning chains?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E&lt;/td&gt;
&lt;td&gt;Iterative refinement&lt;/td&gt;
&lt;td&gt;Do improvement patterns compound?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;50 tasks per track, 250 total. All available in the &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/benchmark_v2_dataset.json" rel="noopener noreferrer"&gt;benchmark dataset&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Problem 5: Single-pass baselines.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The brainless condition now runs through a properly equivalent pipeline, same structure, same number of agents, just without the Brain recall/learn steps. No more two-pass advantage that could inflate brainless scores. Baseline workflows: &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/baseline_track_a.yml" rel="noopener noreferrer"&gt;&lt;code&gt;baseline_track_a.yml&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/baseline_track_b.yml" rel="noopener noreferrer"&gt;&lt;code&gt;baseline_track_b.yml&lt;/code&gt;&lt;/a&gt;, etc.&lt;/p&gt;

&lt;p&gt;I also split the pipeline into three standalone scripts, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/run_benchmark_v2.py" rel="noopener noreferrer"&gt;execution&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/judge_benchmark.py" rel="noopener noreferrer"&gt;judging&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/examples/benchmark_v2/aggregate_benchmark.py" rel="noopener noreferrer"&gt;aggregation&lt;/a&gt;, so you can re-run any phase independently. Eleven code changes total, all committed and tested. 3,014 unit tests passing. You can verify everything in the &lt;a href="https://github.com/marcosomma/orka-reasoning/tree/master/examples/benchmark_v2/results" rel="noopener noreferrer"&gt;results directory&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I felt good about this. I'd addressed every valid criticism. Time to re-run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Numbers
&lt;/h2&gt;

&lt;p&gt;Here's the overall aggregate from 250 tasks, brain vs. brainless:&lt;/p&gt;

&lt;h3&gt;
  
  
  Rubric Scores (1–10 scale, six dimensions)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Brain&lt;/th&gt;
&lt;th&gt;Brainless&lt;/th&gt;
&lt;th&gt;Delta&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning Quality&lt;/td&gt;
&lt;td&gt;9.51&lt;/td&gt;
&lt;td&gt;9.52&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−0.01&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structural Completeness&lt;/td&gt;
&lt;td&gt;9.87&lt;/td&gt;
&lt;td&gt;9.83&lt;/td&gt;
&lt;td&gt;+0.04&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Depth of Analysis&lt;/td&gt;
&lt;td&gt;8.79&lt;/td&gt;
&lt;td&gt;8.74&lt;/td&gt;
&lt;td&gt;+0.05&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Actionability&lt;/td&gt;
&lt;td&gt;9.67&lt;/td&gt;
&lt;td&gt;9.64&lt;/td&gt;
&lt;td&gt;+0.03&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Domain Adaptability&lt;/td&gt;
&lt;td&gt;9.85&lt;/td&gt;
&lt;td&gt;9.82&lt;/td&gt;
&lt;td&gt;+0.03&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Confidence Calibration&lt;/td&gt;
&lt;td&gt;9.38&lt;/td&gt;
&lt;td&gt;9.39&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−0.01&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Overall&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;9.37&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;9.31&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+0.06&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A +0.06 rubric delta across 250 tasks.&lt;/p&gt;

&lt;p&gt;For reference, v1 was +0.10 across 30 tasks. So the effect got &lt;em&gt;smaller&lt;/em&gt; with more data, not larger. That's not what you want to see.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pairwise Comparison (245 head-to-head comparisons)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Brain Wins&lt;/th&gt;
&lt;th&gt;Brainless Wins&lt;/th&gt;
&lt;th&gt;Tie&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Stronger reasoning&lt;/td&gt;
&lt;td&gt;152&lt;/td&gt;
&lt;td&gt;91&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;More complete&lt;/td&gt;
&lt;td&gt;149&lt;/td&gt;
&lt;td&gt;92&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;More trustworthy&lt;/td&gt;
&lt;td&gt;151&lt;/td&gt;
&lt;td&gt;92&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Overall&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;151&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;92&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Brain win rate: &lt;strong&gt;61.6%&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's where it gets uncomfortable. The pairwise judge says brain wins 62% of the time. The rubric judge says brain is +0.06 better, which is noise at a 9.3/10 baseline. These two metrics should agree. They don't.&lt;/p&gt;

&lt;p&gt;I've seen this pattern before. It's length/position bias. Brain responses tend to be longer because the pipeline has more agents in the chain, which means more context, which means more text. Pairwise judges prefer longer answers. The rubric doesn't care about length, it scores each dimension independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Per-Track Breakdown
&lt;/h3&gt;

&lt;p&gt;This is where the story gets interesting:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Track&lt;/th&gt;
&lt;th&gt;Focus&lt;/th&gt;
&lt;th&gt;Rubric Δ&lt;/th&gt;
&lt;th&gt;Pairwise Win%&lt;/th&gt;
&lt;th&gt;Brainless Baseline&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;Cross-domain transfer&lt;/td&gt;
&lt;td&gt;−0.02&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;9.33&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;Ethical reasoning&lt;/td&gt;
&lt;td&gt;+0.00&lt;/td&gt;
&lt;td&gt;52%&lt;/td&gt;
&lt;td&gt;9.54&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;Routing decisions&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+0.40&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;8.12&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;D&lt;/td&gt;
&lt;td&gt;Multi-step reasoning&lt;/td&gt;
&lt;td&gt;+0.08&lt;/td&gt;
&lt;td&gt;60%&lt;/td&gt;
&lt;td&gt;9.49&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;E&lt;/td&gt;
&lt;td&gt;Iterative refinement&lt;/td&gt;
&lt;td&gt;+0.06&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;9.61&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Track C stands out. It's the hardest track, brainless only scores 8.12, nearly a full point below every other track. And it's the only track where brain shows a meaningful rubric gain: +0.40 across six dimensions.&lt;/p&gt;

&lt;p&gt;Track E has the highest pairwise win rate (76%) but the smallest rubric gain (+0.06). That's the length bias signature, the pairwise judge loves brain's longer outputs, but the rubric says they're not actually better.&lt;/p&gt;

&lt;p&gt;Track B is essentially a coin flip. 52% pairwise, +0.00 rubric. The Brain adds nothing to ethical reasoning tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Ugly Detail: Skill Usage
&lt;/h3&gt;

&lt;p&gt;Here's what really killed me. I dug into the individual results to see how many tasks actually used their recalled skill:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tasks with skill recall attempted:&lt;/strong&gt; 51 / 250 (20%)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tasks that actually used the recalled skill:&lt;/strong&gt; 0 / 250&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Average semantic match score:&lt;/strong&gt; ~0.02 (near zero)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Zero. Not one single task out of 250 used the recalled skill. The model read the skill, evaluated it, and decided every single time that it wasn't helpful. And the semantic similarity between the abstract skill and the actual task was essentially random noise.&lt;/p&gt;

&lt;p&gt;The abstraction layer I was so proud of, the one that converts "Load CSV files into staging tables using pandas" into "implement [target]", produced skills so abstract they were vacuous. Two words of content. The embedding model sees no relationship between "implement [target]" and any real task. The execution model correctly recognizes that "implement [target]" tells it nothing it doesn't already know.&lt;/p&gt;

&lt;p&gt;I had gone from skills that were too specific (literal LLM paraphrases) to skills that were too abstract (empty shells). The sweet spot, actual transferable knowledge, was somewhere I hadn't found.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sitting with the Discomfort
&lt;/h2&gt;

&lt;p&gt;I'm going to be honest about what went through my head at this point. I've been working on OrKa for over a year. Forty blog posts. A research paper about the Agricultural Threshold for machine intelligence. An open-source framework that allow me to test and experiment and explore my idea with real AI runs. And the core thesis, that persistent memory makes agents better, keeps failing to show up in the numbers.&lt;/p&gt;

&lt;p&gt;I considered dropping the whole Brain system. Making OrKa just an orchestration framework. Simpler. Easier to explain. No embarrassing benchmarks.&lt;br&gt;
But then I looked at &lt;strong&gt;Track C&lt;/strong&gt; again.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;Track C **is the only track where brainless *struggles&lt;/em&gt;. It scores 8.12, good, but not great. The tasks involve complex routing decisions where the model has to consider multiple paths and trade-offs. This is the only track where the model actually needs help.&lt;/p&gt;

&lt;p&gt;And it's the only track where brain provides meaningful help. +0.40 rubric delta is not noise. Across 50 tasks and six scoring dimensions, that's a consistent, measurable improvement.&lt;/p&gt;

&lt;p&gt;The pattern is simple: the Brain helps when the model needs help, and doesn't help when the model doesn't need help.&lt;/p&gt;

&lt;p&gt;That sounds obvious in retrospect. &lt;strong&gt;&lt;em&gt;But it means the thesis isn't wrong, it's just being tested in the wrong conditions&lt;/em&gt;&lt;/strong&gt;. You wouldn't evaluate a life jacket by putting it on people standing on dry land and measuring whether they're drier.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Real Problem: What Is a Memory?
&lt;/h2&gt;

&lt;p&gt;This is where the story changes. Because instead of asking "does memory help?" I started asking "what is a memory, actually?"&lt;/p&gt;

&lt;p&gt;Think about how you remember how to drive a car. What fires in your brain when you approach an unfamiliar intersection?&lt;/p&gt;

&lt;p&gt;It's not one thing. It's not "turn the wheel, press the gas." That's the procedural part, and yes, it's there. But it's bound together with other things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The time you nearly got T-boned&lt;/strong&gt; because you assumed a green light meant it was safe without checking cross traffic. That's episodic memory, a specific event with emotional weight.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Right of way doesn't mean right of safety"&lt;/strong&gt;, That's semantic memory. A general fact you learned, maybe from a driving instructor, maybe from experience.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Checking mirrors BEFORE entering the intersection prevents blind-spot collisions BECAUSE turning reduces your field of vision"&lt;/strong&gt;, That's causal reasoning. You know &lt;em&gt;why&lt;/em&gt; the sequence matters, not just &lt;em&gt;that&lt;/em&gt; it matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When you encounter the intersection, all of these fire together. The procedure tells you what to do. The episode tells you what happened last time. The semantic fact tells you a principle. The causal link tells you why. That combination, that &lt;em&gt;binding&lt;/em&gt;, is what makes the memory useful. Any single component alone is much less helpful.&lt;/p&gt;

&lt;p&gt;Now look at what OrKa Brain currently stores as a "skill":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;implement [target]
trace [target]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. No episodes. No semantic context. No causal reasoning. Just two abstract action verbs. No wonder the model ignores it. It's like handing a driver a note that says "steer [vehicle]" and expecting it to help at the intersection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Memory Binding Problem
&lt;/h2&gt;

&lt;p&gt;I went down a rabbit hole into cognitive science literature on this. What I found is that neuroscientists have been arguing about this exact problem for decades. They call it the &lt;em&gt;binding problem&lt;/em&gt;, how does the brain take separate memory traces stored in different systems and combine them into a unified experience?&lt;/p&gt;

&lt;p&gt;The hippocampus doesn't store the memory. It stores the &lt;em&gt;index&lt;/em&gt;, the binding that links the procedural memory in the motor cortex, the emotional trace in the amygdala, the spatial context in the parietal cortex, and the semantic facts in the temporal lobe. When you recall one, you recall all of them, because they're bound together.&lt;/p&gt;

&lt;p&gt;I had built the hippocampus and the motor cortex as two separate systems that had never met.&lt;/p&gt;

&lt;p&gt;Here's what actually exists in OrKa today:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Skill system&lt;/strong&gt; (fully operational, used in benchmarks):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Abstract procedure steps&lt;/li&gt;
&lt;li&gt;Preconditions and postconditions&lt;/li&gt;
&lt;li&gt;Transfer history and confidence scores&lt;/li&gt;
&lt;li&gt;Structural/semantic matching for recall&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The Episode system&lt;/strong&gt; (fully built, tested, &lt;em&gt;never used in any benchmark&lt;/em&gt;):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Specific task input and outcome&lt;/li&gt;
&lt;li&gt;What worked and what failed&lt;/li&gt;
&lt;li&gt;Root cause analysis for failures&lt;/li&gt;
&lt;li&gt;Actionable lessons learned&lt;/li&gt;
&lt;li&gt;Resource metrics (tokens, latency)&lt;/li&gt;
&lt;li&gt;Links to related episodes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both systems are production-ready. Both have full test coverage. Both are integrated into the Brain class. I wrote &lt;code&gt;record_episode()&lt;/code&gt;, &lt;code&gt;recall_episodes()&lt;/code&gt;, &lt;code&gt;EpisodeStore&lt;/code&gt;, &lt;code&gt;EpisodeRecall&lt;/code&gt;, all of it. Complete with semantic search, retention policies, and four-dimensional scoring.&lt;/p&gt;

&lt;p&gt;And then I never connected them together.&lt;/p&gt;

&lt;p&gt;The Skill has no &lt;code&gt;episode_id&lt;/code&gt; field. The Episode has no &lt;code&gt;skill_id&lt;/code&gt; field. &lt;code&gt;brain.learn()&lt;/code&gt; creates a Skill but not an Episode. &lt;code&gt;brain.recall()&lt;/code&gt; returns Skills but not Episodes. The benchmark workflows run brain_learn and brain_recall, but never brain_record_episode or brain_recall_episodes.&lt;/p&gt;

&lt;p&gt;Two complete memory systems, sitting in the same codebase, sharing no information.&lt;/p&gt;

&lt;p&gt;When I saw this, I felt stupid. But I also felt something else: the architecture was already 80% there. The hard parts, embedding storage, semantic search, decay policies, scoring systems, were done. The missing piece wasn't a new system. It was the wiring between existing systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Memory Should Actually Look Like
&lt;/h2&gt;

&lt;p&gt;Here's the concept I'm now calling a &lt;strong&gt;Memory Bundle&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────────────────────────────────┐
│            MEMORY BUNDLE                │
│                                         │
│  ┌───────────┐  ┌──────────────────┐    │
│  │ Procedure │  │ Episodes (1..N)  │    │
│  │ (steps)   │──│ what worked      │    │
│  │           │  │ what failed      │    │
│  └───────────┘  │ lessons          │    │
│                 │ "X+Z → Y"        │    │
│  ┌───────────┐  └──────────────────┘    │
│  │ Semantic  │                          │
│  │ (domain   │  ┌──────────────────┐    │
│  │  facts)   │  │ Causal Links     │    │
│  │           │  │ "A because B"    │    │
│  └───────────┘  └──────────────────┘    │
│                                         │
│  transfer_score = f(all_components)     │
└─────────────────────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the system learns from an execution, it creates &lt;em&gt;both&lt;/em&gt; a skill AND an episode, linked by ID. The skill stores the abstract procedure. The episode stores what actually happened, the specific outcome, what worked, what failed, and crucially, the &lt;em&gt;lessons&lt;/em&gt;: "Running validation before deduplication caught 30% of bad records that would have been duplicated, always validate first."&lt;/p&gt;

&lt;p&gt;When the system recalls, it returns the skill &lt;em&gt;with its episodes attached&lt;/em&gt;. The prompt to the model isn't "implement [target]", it's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Here's an abstract procedure: implement [target] → validate [component] → trace [target].&lt;/p&gt;

&lt;p&gt;This skill has been applied 3 times before:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Data engineering (ETL)&lt;/strong&gt;: Validation before dedup caught 30% of dirty records. Lesson: always validate before any deduplication step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API integration&lt;/strong&gt;: Target implementation worked, but tracing missed async callbacks. Lesson: tracing needs to account for async execution paths.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log analysis&lt;/strong&gt;: Pattern worked well. Filtering noisy entries before analysis reduced false positives by 40%.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's a memory a model can actually use. It has the abstract pattern (transferable) AND the concrete evidence (grounding). The model can decide whether the pattern applies here based on real outcomes, not just structural similarity.&lt;/p&gt;

&lt;p&gt;The transfer scoring changes too. A skill backed by five successful episodes with clear lessons should score higher than a skill backed by zero episodes. The episode quality becomes part of the transfer decision.&lt;/p&gt;

&lt;p&gt;And feedback updates both, the skill's confidence changes, AND a new episode gets recorded for this application. The episode chain grows over time, and future recalls get richer context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Is Actually About the Thesis
&lt;/h2&gt;

&lt;p&gt;My research paper argues that intelligence becomes civilization-scale only through &lt;strong&gt;recursive environmental control loops&lt;/strong&gt;, project, act, observe, revise, compound. Agriculture was the first time humans did this at scale. The agricultural threshold.&lt;/p&gt;

&lt;p&gt;The current Brain system doesn't cross that threshold. It projects (learns a skill), acts (recalls it), but doesn't truly observe or revise. The skill never learns from its own application. It just accumulates abstract patterns with no connection to real outcomes.&lt;/p&gt;

&lt;p&gt;The Memory Bundle changes this. Each episode is an observation. Each lesson is a revision. Each future recall that includes those lessons is compounding. The loop closes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Learn&lt;/strong&gt;: Execute a task → create skill + record episode (with what worked/failed)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recall&lt;/strong&gt;: Find matching skill → include its episodes as evidence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apply&lt;/strong&gt;: Model uses the procedure + the concrete lessons&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feedback&lt;/strong&gt;: Record a new episode for this application → update skill confidence&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compound&lt;/strong&gt;: Next recall is richer, it has more episodes, more lessons, more evidence&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's the recursive loop. That's the agricultural threshold. And the architecture for it already exists, it just needs the binding.&lt;/p&gt;

&lt;h2&gt;
  
  
  What About Track C?
&lt;/h2&gt;

&lt;p&gt;This also explains why Track C was the only track that showed improvement. Track C tasks are routing decisions, complex, multi-path choices where the model has to weigh trade-offs. These are exactly the kind of tasks where episodic evidence would help most.&lt;/p&gt;

&lt;p&gt;When someone says "last time we tried path A for a similar routing problem, it failed because of X, path B worked because of Y," that's genuinely new information. The model can't derive it from its weights. It's system-specific, run-specific, outcome-specific.&lt;/p&gt;

&lt;p&gt;The current brain helped Track C even without episodes because the tasks are hard enough that any additional context, even a vague abstract skill, provides a useful scaffold. But imagine Track C with Memory Bundles, the model would get both the abstract pattern AND the specific outcomes from previous routing decisions.&lt;/p&gt;

&lt;p&gt;Tracks A, B, D, and E didn't improve because the model already scores 9.3+/10 on them. It doesn't need help. No amount of memory, procedural, episodic, or otherwise, will improve a 9.5/10 response to a 10/10 response. The tasks aren't hard enough to require accumulated knowledge.&lt;/p&gt;

&lt;p&gt;This isn't a failure of the memory system. It's a boundary condition. Memory helps when the task exceeds single-shot capability. It doesn't help when the model is already near-perfect without it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'm Not Claiming
&lt;/h2&gt;

&lt;p&gt;I want to be careful here, because I've been burned before by getting ahead of my own evidence.&lt;/p&gt;

&lt;p&gt;I'm &lt;strong&gt;not&lt;/strong&gt; claiming that Memory Bundles will definitely show large improvements. I'm claiming that the current system stores memories that are too impoverished to be useful, and I now understand what richer memories should look like.&lt;/p&gt;

&lt;p&gt;I'm &lt;strong&gt;not&lt;/strong&gt; claiming the ceiling effect is the only problem. The pairwise-rubric disagreement at 62% vs +0.06 suggests position/length bias is still contaminating the pairwise results. That confound exists regardless of memory architecture.&lt;/p&gt;

&lt;p&gt;I'm &lt;strong&gt;not&lt;/strong&gt; claiming this is a new idea. Cognitive scientists have written about memory binding for decades. What's new (maybe) is applying it to agent memory systems where the default assumption seems to be that one type of memory, usually RAG-style document retrieval, is sufficient.&lt;/p&gt;

&lt;p&gt;And I'm &lt;strong&gt;not&lt;/strong&gt; pretending the community feedback didn't shape this thinking. When TechPulse Lab wrote that episodic and institutional memory matters more than procedural memory, they were describing exactly the gap I ended up finding. When Nova Elvaris pointed out that skills can only grow, never decay, that's the absence of failure episodes. When Kuro said memory maintenance matters more than storage, that's about binding quality, not storage quantity.&lt;/p&gt;

&lt;p&gt;I just didn't understand what they were telling me until the numbers forced me to look harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens Next
&lt;/h2&gt;

&lt;p&gt;The code changes needed are surprisingly small. The Episode system is already built, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/brain/episode.py" rel="noopener noreferrer"&gt;&lt;code&gt;episode.py&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/brain/episode_store.py" rel="noopener noreferrer"&gt;&lt;code&gt;episode_store.py&lt;/code&gt;&lt;/a&gt;, &lt;a href="https://github.com/marcosomma/orka-reasoning/blob/master/orka/brain/episode_recall.py" rel="noopener noreferrer"&gt;&lt;code&gt;episode_recall.py&lt;/code&gt;&lt;/a&gt; are all production-ready with tests. What's needed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Binding&lt;/strong&gt;: Add &lt;code&gt;episode_ids[]&lt;/code&gt; to Skill, add &lt;code&gt;skill_id&lt;/code&gt; to Episode. When &lt;code&gt;brain.learn()&lt;/code&gt; fires, it creates both and links them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unified recall&lt;/strong&gt;: When &lt;code&gt;brain.recall()&lt;/code&gt; finds a matching skill, it fetches the associated episodes automatically. The prompt template includes both the abstract procedure and the concrete lessons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transfer scoring&lt;/strong&gt;: Episode quality becomes a component of the transfer score. Skills with successful episodes score higher.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feedback loop&lt;/strong&gt;: &lt;code&gt;brain.feedback()&lt;/code&gt; records a new episode for the current application, so the skill's evidence base grows over time.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then re-run the benchmark. Specifically on Track C-difficulty tasks, where the model actually needs help.&lt;/p&gt;

&lt;p&gt;I'm not going to promise the numbers will be different this time. I've been wrong before, twice now, measured against my own benchmarks, published for everyone to see. But I understand something I didn't understand before: a memory without experience is just a note. A memory with experience is a skill.&lt;/p&gt;

&lt;p&gt;The plumbing metaphor from the first article still holds. But I was plumbing one pipe when the system needs at least four, all flowing into the same tap.&lt;/p&gt;




&lt;p&gt;All benchmark data, scripts, and results are publicly available in the &lt;a href="https://github.com/marcosomma/orka-reasoning/tree/master/examples/benchmark_v2" rel="noopener noreferrer"&gt;OrKa repository&lt;/a&gt;. The &lt;a href="https://github.com/marcosomma/orka-reasoning/tree/master/examples/benchmark_v2/results" rel="noopener noreferrer"&gt;full result files&lt;/a&gt; include every individual task response, judge score, and pairwise comparison. If you want to re-run the analysis: &lt;code&gt;python aggregate_benchmark.py --judge-tag local&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If you've worked on agent memory systems and found similar walls, or found ways through them, I'd genuinely like to hear about it. The comments on the first article were more useful than most papers I've read on the topic.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This is part of an ongoing series about building &lt;a href="https://github.com/marcosomma/orka-reasoning" rel="noopener noreferrer"&gt;OrKa&lt;/a&gt;, an open-source YAML-first agent orchestration framework. Previous installments: &lt;a href="https://dev.to/marco_somma_a9e88a3063f3/i-tried-to-turn-agent-memory-into-plumbing-instead-of-philosophy-1bpm"&gt;Part 1: Plumbing Instead of Philosophy&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>rag</category>
      <category>graphknowledge</category>
    </item>
  </channel>
</rss>
