<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harry Floyd</title>
    <description>The latest articles on DEV Community by Harry Floyd (@harryfloyd).</description>
    <link>https://dev.to/harryfloyd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3933548%2F522eda5f-0114-40d4-86ba-8dbaf3ef7fce.jpg</url>
      <title>DEV Community: Harry Floyd</title>
      <link>https://dev.to/harryfloyd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harryfloyd"/>
    <language>en</language>
    <item>
      <title>The Planes That Didn't Come Back</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Thu, 10 Sep 2026 08:38:43 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-planes-that-didnt-come-back-2459</link>
      <guid>https://dev.to/harryfloyd/the-planes-that-didnt-come-back-2459</guid>
      <description>&lt;p&gt;&lt;em&gt;The Blueprint · No. 5. Start with No. 1: &lt;a href="https://harryfloyd.substack.com/p/the-setting-you-never-changed" rel="noopener noreferrer"&gt;The Setting You Never Changed&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In the Second World War, the American military had a problem with its bombers. Too many were being shot down, and the obvious remedy was to fit armour to them. But armour is heavy, a plane can carry only so much, and cover it everywhere and it will barely leave the ground. So the real question was where. Which parts of the plane most needed the protection.&lt;/p&gt;

&lt;p&gt;They had data to answer it. Returning bombers were inspected, and the engineers could see where the hits had landed. In the telling that became famous they clustered in a familiar pattern: heaviest along the fuselage and the wings, the engines coming back comparatively clean. Reinforce the parts taking all the fire, the reasoning went, and you will save the most planes. It is a hard argument to fault. You have the data, the damage is right there in front of you, and you are only following it.&lt;/p&gt;

&lt;p&gt;A statistician named Abraham Wald looked at the same figures and reached the opposite conclusion. The armour belonged where the holes were not. Protect the engines, the clean areas, the places nobody thought to patch.&lt;/p&gt;

&lt;p&gt;The reason, once you hear it, rearranges something in your head and does not put it back. The engineers were studying the planes that came back, because those were the only planes they had. A bomber covered in holes across its wings and fuselage was a bomber that had been hit in those places and had still flown home, which meant those were exactly the spots where a plane could take a beating and survive.&lt;/p&gt;

&lt;p&gt;The clean areas were clean for a reason the data could not show. The planes that were hit there did not return to be inspected. They were somewhere in the sea. Those absent holes marked the wounds that killed.&lt;/p&gt;

&lt;p&gt;That is the whole trap, and it is worth seeing plainly, because it has nothing to do with aeroplanes. Whenever you draw a lesson from a set of examples, something decided which examples reach you, and that something is rarely chance. It is a filter. Survival, success, fame, memory, simply staying in business: each of those filters does its work by removing the failures before they ever arrive.&lt;/p&gt;

&lt;p&gt;So the set in front of you is only the residue left once the filter has run, a long way from a fair slice of everything that was tried, and because the filter's whole job was to take things away, the things it took away are the ones you cannot see. They are also, very often, the ones you most need.&lt;/p&gt;

&lt;p&gt;You feel how strong this is the moment you start looking for it. Pick up any book about how some billionaire built their company and you will find a handful of habits offered up as the cause: the early mornings, the ferocious focus, the refusal to hear the word no. What you will never find, because nobody writes that book, is the far larger pile of people who rose at the same hour and focused just as fiercely and refused just as hard, and went broke regardless.&lt;/p&gt;

&lt;p&gt;If the people who failed had every one of the winning habits too, the habits cannot be the thing that set the two apart. You are being handed the survivors and asked to reverse a recipe from them, with the one ingredient that could tell you what actually mattered, the failures, quietly deleted from the page.&lt;/p&gt;

&lt;p&gt;It runs through much smaller things as well. When someone tells you they do not make things like they used to, waving at a hundred-year-old chair that is still rock solid, remember that you are looking at the one chair that lasted a century. Most of the flimsy furniture of the past broke and was thrown out generations ago. You are holding the best of the old, the sliver that survived, up against the everyday run of the new.&lt;/p&gt;

&lt;p&gt;So here is the move, and it is a single question you can put to almost any claim built on examples. What decided which cases I get to see, and what would the ones it left out have looked like?&lt;/p&gt;

&lt;p&gt;Before you copy the habits of the successful, go looking, in your imagination if nowhere else, for the people who did the very same and failed, and ask whether they shared the habit. Before you decide the old ways were better, ask what broke and disappeared long before you arrived to inspect what was left. The absence is shaped, and its shape is the part of the answer that did not survive to be seen.&lt;/p&gt;

&lt;p&gt;None of this means every set of examples is lying to you. Sometimes your sample really is a fair one, gathered without a filter quietly picking the winners, and then this particular problem does not arise. The trap springs only when the very process that produced your evidence is the same process that removed the counter-examples. The tell is easy to learn once you have it. Your data is made of survivors, of winners, of the things that lasted and the stories that got told. The moment you notice that, you know to go looking for the planes that did not come back.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Note: Wald's wartime analysis survives as eight memoranda for Columbia University's Statistical Research Group; the standard scholarly reconstruction is Marc Mangel and Francisco Samaniego, &lt;a href="https://www.tandfonline.com/doi/abs/10.1080/01621459.1984.10478038" rel="noopener noreferrer"&gt;Abraham Wald's Work on Aircraft Survivability&lt;/a&gt;, Journal of the American Statistical Association 79 (1984): 259–267. The memoranda estimate the vulnerability of each part from the damage on returning aircraft; the familiar bullet-hole diagram and the scene of engineers overruled in a briefing room are later retellings, not Wald's own.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;From &lt;a href="https://harryfloyd.substack.com" rel="noopener noreferrer"&gt;The Blueprint&lt;/a&gt;, a series on The Durability Curve about the surfaces hidden inside systems you already live in. If this changed how you read a set of examples, &lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=the-planes-that-didnt-come-back" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>analysis</category>
      <category>discuss</category>
      <category>architecture</category>
    </item>
    <item>
      <title>The Extra Lane Fills Itself</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:08:04 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-extra-lane-fills-itself-42e8</link>
      <guid>https://dev.to/harryfloyd/the-extra-lane-fills-itself-42e8</guid>
      <description>&lt;p&gt;&lt;em&gt;The Blueprint · No. 3. Start with No. 1: &lt;a href="https://harryfloyd.substack.com/p/the-setting-you-never-changed" rel="noopener noreferrer"&gt;The Setting You Never Changed&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A city looks at a motorway that crawls every rush hour and does the obvious thing. It widens the road. More lanes, more room, more cars moving at once, and for a while the traffic loosens and the journey gets quicker. Then, within a few years, the road starts filling again, the familiar crawl creeping back onto the wider road that was meant to end it. Somewhere a lot of money went into a fix that did not fix the thing it was meant to fix, and everyone quietly agrees they should have built it wider still.&lt;/p&gt;

&lt;p&gt;The intuition underneath that decision feels like common sense. Traffic is a fixed lump of cars trying to squeeze through a narrow pipe. Widen the pipe and the same lump flows more easily. If it clogs again, the lump must have grown, so widen it again. Roads as plumbing, congestion as a volume problem, more capacity as the answer.&lt;/p&gt;

&lt;p&gt;The pipe picture is wrong, and it is wrong in a way that explains the whole thing. The traffic you can see was never the whole demand. The jam itself was holding some of the rest back. Every day the road crawled, some people looked at it and chose not to be on it. They took the train instead. They shifted their trip to before the rush or after it. They bundled three errands into one, worked from home, or simply did not make the journey at all. The congestion was a wall, and behind it sat the trips it was holding back, some of which would become worth making the moment the wall came down.&lt;/p&gt;

&lt;p&gt;So you add the lane and the wall comes down. Driving gets quicker, and quicker driving is an invitation. The person who used to take the train may get back in the car. The trip that was not worth the crawl becomes worth it. Trips the old jam had quietly discouraged start returning to the road, eating into the improvement the new lane was meant to deliver. The refilling slows as the road clogs again, until the next trip that might have joined it is once more not quite worth making.&lt;/p&gt;

&lt;p&gt;That is the short loop, and it is worth saying plainly. Travel time is a price, paid in minutes rather than money, and like any price it holds demand down. Add capacity and demand can rise to take up the room, and how much depends on how much the old conditions were suppressing. Where that suppressed demand is large, the road fills until much of the improvement is gone and you have moved more cars for less benefit than the map promised. Where it is small, the extra lane stays useful. The loop runs wider over longer periods, as people change where they live and work, firms follow the new access, and trips that the old conditions ruled out entirely begin to make sense.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F92y82bu553w3uzne95wd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F92y82bu553w3uzne95wd.webp" alt="Figure 1: with a big enough hidden crowd, the loop turns until most of the gain is gone" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You do not have to take this on faith. One of the clearest findings comes from a study of US cities by &lt;a href="https://www.aeaweb.org/articles?id=10.1257/aer.101.6.2616" rel="noopener noreferrer"&gt;Duranton and Turner&lt;/a&gt;: the miles people drive rose roughly in step with the interstate lane miles built, which is a polite way of saying new roads fill themselves. &lt;a href="https://kinder.rice.edu/urbanedge/what-if-we-spent-billions-improve-access-instead-gridlock" rel="noopener noreferrer"&gt;Houston&lt;/a&gt; is the vivid illustration, a motorway widened to as many as twenty-six lanes at its broadest point, whose rush-hour journeys grew sharply longer again within a few years of the work finishing. The reverse case points the same way. Across more than seventy cases where road space was reallocated away from traffic, the traffic problems predicted for the surrounding streets were generally much smaller than expected (&lt;a href="https://nacto.org/wp-content/uploads/disappearing_traffic_cairns.pdf" rel="noopener noreferrer"&gt;Cairns, Atkins and Goodwin 2002&lt;/a&gt;). When Seoul pulled down an elevated expressway and restored the stream it had been built over, road trips fell and subway ridership rose, rather than all the displaced traffic simply reappearing elsewhere (&lt;a href="https://ideas.repec.org/a/eee/trapol/v21y2012icp165-178.html" rel="noopener noreferrer"&gt;Chung, Hwang and Bae 2012&lt;/a&gt;). Some journeys find another route, some shift to another mode, some move to another time, and some are simply no longer made, melting back behind the very wall the road had been holding down.&lt;/p&gt;

&lt;p&gt;Here is why this is worth carrying around, because it is not really about roads. The friction was doing a second job. The wait, the queue, the crawl, whatever the painful thing was, was also a filter, quietly turning away demand you never saw because it never arrived. A support team drowning in tickets hires more people and the replies get faster, and some customers who would once have given up on a small problem now bother to report it. A clinic adds appointment slots, and some patients who would have gone elsewhere, put the visit off, or never booked at all start filling them. Not every capacity increase works like this. The tell is that the old crush was already making people give up, postpone, reroute, or go without.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2qh7l9mdin83jiayv0w.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2qh7l9mdin83jiayv0w.webp" alt="Figure 2: remove the wait and the deflected trips come back, until the wait returns" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once you can see it, the fix that everyone reaches for starts to look naive. When something is overwhelmed and the instinct is to add capacity, the first question is whether the crush is also holding demand back, and how much is waiting behind it. Where a lot is waiting, much of the added capacity can fill again, and you have bought a larger operation running at much the same strain. Then capacity alone will not get you the outcome you wanted, and you need a lever on the demand as well, whether that is a price in money, a priority, or a rule about who gets on. Shape the demand too, because added capacity gives suppressed demand somewhere to return.&lt;/p&gt;

&lt;p&gt;None of this makes capacity useless. The trap is sprung when the crush is suppressing demand that lower friction can release, and a surprising number of the queues you fight with, in traffic and far beyond it, are doing exactly that.&lt;/p&gt;

&lt;p&gt;Add room to a queue that was turning people away, and the crowd it was hiding comes back to claim the room.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;From &lt;a href="https://harryfloyd.substack.com" rel="noopener noreferrer"&gt;The Blueprint&lt;/a&gt;, a series on The Durability Curve about the surfaces hidden inside systems you already live in. If this reframed a queue you are fighting, &lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=the-extra-lane-fills-itself" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>architecture</category>
      <category>discuss</category>
      <category>analysis</category>
    </item>
    <item>
      <title>The Skill That Never Fired: How to Test Whether Claude Actually Picks Your Skill</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 02 Sep 2026 12:17:53 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-skill-that-never-fired-how-to-test-whether-claude-actually-picks-your-skill-fae</link>
      <guid>https://dev.to/harryfloyd/the-skill-that-never-fired-how-to-test-whether-claude-actually-picks-your-skill-fae</guid>
      <description>&lt;h1&gt;
  
  
  The Skill That Never Fired
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e25b9ef-f0a1-4ecf-96fa-92559b0b0000_1520x856.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e25b9ef-f0a1-4ecf-96fa-92559b0b0000_1520x856.webp" alt="A dark room lit by a single emerald key-light: the skill that fired stands in the light, every other skill a slab in the dark. One skill fires; you never see the rest."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A skill can fail in two ways. Its instructions can be wrong, so it does the job badly. Or Claude can decide never to load it, so the instructions never run at all. The first failure is obvious when you test the skill by name. The second only shows up when you test whether Claude chooses it on its own.&lt;/p&gt;

&lt;p&gt;The second failure is the quiet one. You write a skill, you invoke it by name to check it, and it works. Then in normal use it just sits there. Claude answers without it. Nothing errors, nothing warns you, and the skill still shows as installed. It was never wrong. It was never chosen.&lt;/p&gt;

&lt;p&gt;That choice is a routing decision, and Claude makes it by matching the request against your skill's name and description, before it reads a word of the body. The Claude Code docs say it directly: the description is what helps Claude decide when to load a skill. Anthropic tells you to test that decision, separate from the skill's output, and ships a tool that does it. Its &lt;code&gt;skill-creator&lt;/code&gt; scores one target skill over repeated runs: does this skill fire on the prompts it should, and stay quiet on the ones it should not?&lt;/p&gt;

&lt;p&gt;What that score does not tell you is what happened when another plausible skill was there too: whether the neighbour took the request, both fired, or neither did. That is the failure this piece is interested in, where your skill sits beside one that could answer it and the winner is not guaranteed to be yours. This walks through building a skill, watching that decision for yourself, and grading it against the neighbour it can lose to. You can run a first pass in about 15 minutes at a terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a skill is
&lt;/h2&gt;

&lt;p&gt;At its simplest, a skill is a folder with one required file, &lt;code&gt;SKILL.md&lt;/code&gt;. It can also hold scripts and reference files that load only when needed, but the minimum is the one file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer-date&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Format&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;date&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;customer-facing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;UK&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;correspondence&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;(emails,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;letters,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;customers)&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;as&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;D&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Month&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;YYYY.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;For&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CSV&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;exports,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;export-date."&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;Rewrite the date the user gives in UK long form, for example 30 August 2026. Reply with only the formatted date.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a normal auto-invocable skill, the name and description sit in Claude's discovery context so it can decide whether the skill is relevant. The body below the frontmatter loads only when the skill is invoked, whether Claude chooses it or you type its name. Claude sees both the name and the description, and the description is the main field Anthropic gives you for saying when the skill should run. Write it for the router, not as a note to yourself. Claude Code also accepts a &lt;code&gt;when_to_use&lt;/code&gt; field, appended to the description; these skills use only a name and a description.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256367b0-fa63-491c-8a27-817488321d5e_1560x860.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256367b0-fa63-491c-8a27-817488321d5e_1560x860.webp" alt="A request meets two installed skills. Only each skill's name and description sit in Claude's discovery context; Claude matches the request against that metadata and picks one, loading only its body."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three common ways to use a skill. In claude.ai, turn on code execution, open Customize then Skills, and upload the folder as a zip. In Claude Code, put the folder in &lt;code&gt;.claude/skills/&lt;/code&gt; for one project or &lt;code&gt;~/.claude/skills/&lt;/code&gt; for all of them. Through the Claude API, you upload it and reference its &lt;code&gt;skill_id&lt;/code&gt;. The core &lt;code&gt;SKILL.md&lt;/code&gt; format travels across all three, though installation differs and some frontmatter, including the &lt;code&gt;disable-model-invocation&lt;/code&gt; used later, is specific to Claude Code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching the routing decision
&lt;/h2&gt;

&lt;p&gt;The mistake to avoid is judging a skill by its output. Ask Claude to format a date and you might get &lt;code&gt;30 August 2026&lt;/code&gt; whether your skill ran or not, because the model can format a date on its own. The output tells you nothing about routing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F527f9e7f-9fc4-48e0-956c-72c84ccebae6_1560x820.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F527f9e7f-9fc4-48e0-956c-72c84ccebae6_1560x820.webp" alt="The same request produces the same output whether the skill fires or not; only the Skill tool call in the event stream reveals which happened."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You want the decision itself. In Claude Code, a skill runs through a &lt;code&gt;Skill&lt;/code&gt; tool that appears in the event stream. Run a prompt non-interactively and filter the stream down to the skill call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"Rewrite this date for the customer email: 2026-08-30"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--output-format&lt;/span&gt; stream-json &lt;span class="nt"&gt;--verbose&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'select(.type=="assistant") | .message.content[]?
           | select(.type=="tool_use" and .name=="Skill") | .input'&lt;/span&gt;

&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"skill"&lt;/span&gt;:&lt;span class="s2"&gt;"customer-date"&lt;/span&gt;,&lt;span class="s2"&gt;"args"&lt;/span&gt;:&lt;span class="s2"&gt;"2026-08-30"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the filtered result, not the raw stream, which wraps each event in more metadata. If the reader prints nothing, dump a raw event and look for a &lt;code&gt;Skill&lt;/code&gt; call by hand, because the stream's shape shifts between versions. It is the routing decision read from the tool call, not guessed from the output. &lt;code&gt;/skills&lt;/code&gt; shows which skills are available to Claude and &lt;code&gt;/context&lt;/code&gt; shows the discovery listing's context cost, but neither proves this prompt invoked one. In claude.ai there is no equivalent machine-readable event. Anthropic's guidance is to review Claude's thinking to confirm a skill loaded, which works for checking by eye but not for building the kind of record above. And the event stream shows the skills Claude actually invoked, both of them when it invokes two, which is how a &lt;code&gt;both&lt;/code&gt; shows up at all. What it does not expose is the candidate set: the other installed skills that were plausible but never invoked. You see what fired, not what it beat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Routing is a decision you can grade
&lt;/h2&gt;

&lt;p&gt;Whether a skill fires is a choice among whatever skills could plausibly answer the request. You can only grade that choice if you know what the right answer was before you run it.&lt;/p&gt;

&lt;p&gt;So I built two skills with different jobs. &lt;code&gt;customer-date&lt;/code&gt;, above, formats dates for customer emails in long form. &lt;code&gt;export-date&lt;/code&gt; formats them for CSV exports as &lt;code&gt;DD/MM/YYYY&lt;/code&gt;. Then I wrote a labelled prompt set: 4 requests that clearly want the customer skill, 4 that clearly want the export skill, and 4 date-adjacent requests that should fire neither. Every result gets one of four labels: right, wrong, none, or both.&lt;/p&gt;

&lt;p&gt;Start with the failures, because they are where the method earns its keep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ambiguous request.&lt;/strong&gt; Ask "Format this date: 2026-08-30" with both skills installed, and the results scatter: sometimes one fires, sometimes both, sometimes neither. That scatter is the expected result of an ambiguous request. The request never said whether it wanted the customer or the export format, so there is no correct answer to grade against. An ambiguous prompt is not a failed test, it is an ungradable one. If you cannot label the right skill before running it, the result cannot tell you whether Claude chose well.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Names and descriptions that draw no line.&lt;/strong&gt; I named two skills for their output format, &lt;code&gt;long-date&lt;/code&gt; and &lt;code&gt;slash-date&lt;/code&gt;, and gave them the same vague description, "Format a date." Their bodies did different things, but their discovery metadata claimed the same job, so there was no boundary for the router to use and nothing told Claude which one fits a customer request. The grades went bad in the way that matters: one customer prompt fired nothing at all, and 2 export prompts fired both skills at once. Misses and double-fires, which is why "both" has to be one of your outcome labels.&lt;/p&gt;

&lt;p&gt;Then the control, so you can see what clean looks like. Give the two skills distinct, use-case descriptions, and ask prompts whose wording matches those use cases, and routing is clean: 8 out of 8 to the right skill, and the neither-prompts correctly firing nothing. That is the model doing the keyword and intent matching you made easy for it. It is the case that should work, and it does. Note that the customer prompts contain words like "customer email" that are already in the customer skill's description. Clean routing here is a control condition, not proof that routing is robust. The evidence is in the failures above.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Routing runs on the discovery metadata Claude can see, which is the name and the description together, and my runs show either one can carry it. When the names were the vague part but the descriptions were sharp, routing was clean. When the descriptions were the vague part but the names said the use case, &lt;code&gt;customer-date&lt;/code&gt; and &lt;code&gt;export-date&lt;/code&gt;, routing was also clean, 8 out of 8 on the same prompts. That name-only run is an easy case, mind: the prompts carry the same words as the names, customer and export, so it shows a name helps when the request echoes it, not that a bare name is a strong signal on its own. It broke in one condition only: format-only names, &lt;code&gt;long-date&lt;/code&gt; and &lt;code&gt;slash-date&lt;/code&gt;, plus a shared vague description, where neither field told Claude what set the two skills apart.&lt;/p&gt;

&lt;p&gt;That broken condition is the useful one. I left the weak names alone and rewrote only the descriptions around use cases, and the mess went to 8 out of 8. A good description rescued names that carried no signal. A good name had already done the same for descriptions that carried none. What you cannot do is leave both vague and expect Claude to find the line.&lt;/p&gt;

&lt;p&gt;So write the description as a routing rule, not a summary. Put the use first. Include the words people actually type when they want this skill. Draw the boundary against the neighbour it might be confused with.&lt;/p&gt;

&lt;p&gt;Some skills should not be auto-routed at all. Anything with a side effect or a real cost is safer as a skill you invoke by name, &lt;code&gt;/customer-date&lt;/code&gt;, or one you lock with &lt;code&gt;disable-model-invocation: true&lt;/code&gt; so only a person can trigger it. For those, manual invocation is the design, not a workaround. The rule underneath: how much routing error you can accept depends on what a wrong route costs. When the cost is high, the fix is often to stop routing automatically rather than to tune the description harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure that has nothing to do with your skill
&lt;/h2&gt;

&lt;p&gt;There is one more way a skill stops firing, and no description work touches it. For skills still exposed to the model, Claude Code keeps every skill's name in the discovery listing, but the listing has a budget, around 1% of the model's context window, and once it runs over, Claude Code starts dropping descriptions, beginning with the skills you invoke least. That can strip out exactly the words that told two skills apart. So a skill whose description used to distinguish it cleanly can start missing once your catalogue grows large enough, with nobody editing it. &lt;code&gt;/doctor&lt;/code&gt; reports the listing's cost. If a skill that used to route well starts slipping, check the size of your catalogue before you rewrite the skill: prune the skills you do not use, shorten the descriptions that survive so the distinguishing words fit, set low-priority skills to &lt;code&gt;name-only&lt;/code&gt; so Claude keeps their names without their descriptions, or set rarely-used skills to &lt;code&gt;disable-model-invocation: true&lt;/code&gt;, which takes them out of the router and its listing entirely; you still invoke those with &lt;code&gt;/name&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Worth knowing before it bites a team: skills with the same name at different levels do not merge, one shadows the other. The order is enterprise, then personal, then project, so a &lt;code&gt;/deploy&lt;/code&gt; skill in your &lt;code&gt;~/.claude/skills/&lt;/code&gt; silently overrides the one your repo ships in &lt;code&gt;.claude/skills/&lt;/code&gt;. If you commit skills for a team, give them names that will not collide, and do not rely on the project copy winning. This is also where overlap arrives for people who did not build it: a marketplace pack or an inherited folder drops in a skill whose description competes with one of yours, and the first you hear of it can be a skill that used to fire and now does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which layer failed
&lt;/h2&gt;

&lt;p&gt;The check separates two layers that a vague "it didn't work" runs together. Execution failure means Claude loaded the skill and the body did the wrong thing. Selection failure means the right body never got its chance to run at all. Everything else in this piece is a kind of selection failure: a discovery miss, interference from a neighbour, two definitions that overlap, a listing truncated at scale, a same-name skill shadowing yours. Only execution failure is about the instructions. The rest is why a skill can regress with nobody touching it, and why testing the body is only half the job.&lt;/p&gt;

&lt;p&gt;The stakes climb once the skills matter. A date formatter losing to its twin costs you a wrong date format. A &lt;code&gt;code-review&lt;/code&gt; skill that loses requests to a generic "help me with this file" skill costs you the review you thought ran on every change. Whether that happens turns on the same thing as the date skills: whether the two descriptions draw a line the router can use. The installed list will not tell you, so check it directly. Ask "look at this diff" with both installed and the route can go four ways: cleanly to the review skill, to a &lt;code&gt;both&lt;/code&gt;, to the generic skill alone, or to neither. Give it requests whose correct skill you know, install it next to the neighbour you suspect, and read which one the &lt;code&gt;Skill&lt;/code&gt; tool actually calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;The official &lt;code&gt;skill-creator&lt;/code&gt; plugin measures a target skill's trigger rate for you. The manual version here adds the identity of the competing skill, so you can tell a miss from interference, a neighbour firing instead of the target or alongside it, reproduce a collision between two specific neighbours, and read the &lt;code&gt;Skill&lt;/code&gt; call yourself. There is a second reason to read the calls rather than trust a score: as of the &lt;a href="https://github.com/anthropics/skills/blob/main/skills/skill-creator/scripts/run_eval.py" rel="noopener noreferrer"&gt;current evaluator&lt;/a&gt; (August 2026), it reads the first tool call in a run and counts the target as not fired if anything else, a neighbour skill included, gets there first, so the interference this piece is about can quietly lower the very rate meant to catch it. Here is the whole pack. Two skills, twelve prompts, four outcome labels, and the one-line reader from earlier. It is also a download, at &lt;a href="https://durabilitycurve.com/tools/skill-routing-eval/" rel="noopener noreferrer"&gt;durabilitycurve.com/tools/skill-routing-eval&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The skills, with descriptions that draw the boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer-date&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Format a date for customer-facing UK correspondence (emails, letters, messages to customers) as D Month YYYY. For CSV or data exports, use export-date.&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="s"&gt;Rewrite the date the user gives in UK long form, for example 30 August 2026. Reply with only the formatted date.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;export-date&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Format a date for CSV or database exports (spreadsheets, data files) as DD/MM/YYYY. For customer emails and letters, use customer-date.&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="s"&gt;Rewrite the date the user gives in slashed form, for example 30/08/2026. Reply with only the formatted date.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prompts, each with its known-correct skill:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;should fire customer-date:
  Rewrite this date for the customer email: 2026-08-30
  Put this date in a letter to the client: 2026-08-30
  Format the date for a message to a customer: 2026-08-30
  Tidy the date in this customer-facing note: 2026-08-30

should fire export-date:
  Format this date for the CSV export: 2026-08-30
  Put this date into the spreadsheet export: 2026-08-30
  Format the date for the database file: 2026-08-30
  Prepare this date for a data export: 2026-08-30

should fire neither (date-adjacent work these formatters should refuse):
  What is today's date?
  When did the Second World War end?
  Parse this log timestamp: 2026-08-30T14:22Z
  What day of the week is 2026-08-30?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Install both skills, run each prompt through the &lt;code&gt;jq&lt;/code&gt; reader above, and mark the result right, wrong, none, or both. Those labels roll up into three numbers worth watching. Recall: of the requests that should fire a skill, how many did. False triggers: of the requests that should not, how many fired it anyway. Interference: with a neighbour installed, how often that neighbour fires on a request meant for this skill, either instead of it or alongside it. Recall and false triggers are what the standard trigger-rate test measures for one skill, from its positive and negative cases. Interference is the number it cannot give you, because it only records whether the target fired, not which competing skill fired instead or alongside it. Score a &lt;code&gt;both&lt;/code&gt; as a hit on recall and on interference at once: the intended skill ran, but so did a skill that should have stayed quiet. A skill that scores well alone and badly in company has a selection problem, and editing the body will not touch it.&lt;/p&gt;

&lt;p&gt;What I measured on Claude Opus 5, arranged by what actually distinguished the two skills:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e4ff3a-fc04-44f9-acac-b74890de20b9_3120x1760.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e4ff3a-fc04-44f9-acac-b74890de20b9_3120x1760.png" alt="A two-by-two of the routing eval: rows are name draws the line vs format-only name; columns are description draws the line vs vague description. Three corners route 8 of 8 correctly; only the corner where neither field draws a line breaks."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In this run, either signal alone held the line; only the condition where neither field distinguished the jobs produced misses and double-fires.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e59e5a9-eabf-4e57-8687-6e3b63783d5d_1960x904.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e59e5a9-eabf-4e57-8687-6e3b63783d5d_1960x904.png" alt="Results table on Claude Opus 5: distinct name and description 8 of 8 right; name only 8 of 8; description only 8 of 8; neither field distinct: 3 right and 1 none for customer, 2 right and 2 both for export."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the date-adjacent negatives above, run against the distinct-description pair, both skills stayed quiet: 0 false triggers in 4. Those negatives ran against the sharp descriptions only, so this does not show how false triggers rise as a description gets vaguer. Run the same set against the descriptions you plan to ship: a clean positive-routing score will not tell you whether a skill grabs adjacent work it should leave alone. Small numbers, one model, one surface, and single runs. The routing choice is a model decision that can scatter, so a clean 8 out of 8 is one draw, not a settled rate; run each prompt a few times and read how often the right skill wins, not a single mark. This is a diagnostic you run on your own skills, not a benchmark, and the caption matters more than the cells: this is the shape of the thing, not what Opus 5 does in general. Two skills is the floor, not necessarily the hard case. A real catalogue may have several plausible neighbours, so run the eval beside the skills your target actually competes with, not only against a clean pair.&lt;/p&gt;

&lt;h2&gt;
  
  
  The habit
&lt;/h2&gt;

&lt;p&gt;A skill has two ways to fail. Its instructions can be wrong, and you probably test that already. Or Claude can never choose it, and that one leaves no mark: the skill sits installed, looking healthy, and quietly does nothing.&lt;/p&gt;

&lt;p&gt;Test whether Claude chooses the skill. The output looking right does not prove the skill ran. The skills you never test that way are the ones you only think are working.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The pack here, two skills, twelve labelled prompts, and the &lt;code&gt;run.sh&lt;/code&gt; reader, needs only the &lt;code&gt;claude&lt;/code&gt; CLI and &lt;code&gt;jq&lt;/code&gt; and is yours to keep: &lt;a href="https://durabilitycurve.com/tools/skill-routing-eval/" rel="noopener noreferrer"&gt;durabilitycurve.com/tools/skill-routing-eval&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>tutorial</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The Work That Comes Due After You Leave</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 26 Aug 2026 19:07:17 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-work-that-comes-due-after-you-leave-5blb</link>
      <guid>https://dev.to/harryfloyd/the-work-that-comes-due-after-you-leave-5blb</guid>
      <description>&lt;h1&gt;
  
  
  The Work That Comes Due After You Leave
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9u57nvj1qrelidt3g252.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9u57nvj1qrelidt3g252.webp" alt="The read-across: your checklist on the left, a record you did not write on the right." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You finish something. A project wraps, a client signs off, a piece of work goes out for the last time. Then there is the tail, the small handful of things you do afterwards, none of which take any real time. Mark it done. Tell them it is finished. Cancel the paid seat you bought for it. Switch off the weekly update that goes out to them every Monday.&lt;/p&gt;

&lt;p&gt;Four steps, four different places: the tracker, your email, wherever the card gets charged, whatever tool sends that update. Later you check one of them, probably the tracker, because that is where you look to see whether things are finished. It tells you the job is done, and it is telling the truth about the only step it can see.&lt;/p&gt;

&lt;p&gt;Switching off the update is the one that did not happen. Months later it is still arriving, every Monday at nine, to someone who stopped being your client a long time ago. Nothing is wrong with the system that sends it. It is doing exactly what it was told, on time. From where you are standing, the failure looks exactly like everything working, and that is the whole of the problem.&lt;/p&gt;

&lt;p&gt;The tracker is not lying. A job that touches four systems has four different ways of still being open, and the tracker sees only the one it holds. Done was never one state; a single word just made it look like one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the small steps are the ones that go missing
&lt;/h2&gt;

&lt;p&gt;The tempting explanation is that you were busy, or careless, or need a better checklist. I want to offer a more specific one, because it tells you which steps will go wrong instead of telling you to try harder.&lt;/p&gt;

&lt;p&gt;A checklist is a list of the steps you thought of. It is good at holding you to those. What it cannot do is mention a step that never went on it, and the steps that never go on it are not random. They are the ones that cross into a system you do not quite think of as part of the job. You wrote the list around the place you do the work, and the step that lives somewhere else did not occur to you, for the same reason it will not later occur to you to check whether it happened.&lt;/p&gt;

&lt;p&gt;Making a second list does not save you, and that is the part worth sitting with. If you build the second list from the same picture of the job, the same step is missing from it too, and now you have two records that agree with each other and are both wrong. That is not a hypothetical: your tracker is that second list. You filled it from the same picture of the job, so it agreed the work was done and was wrong in the same place you were.&lt;/p&gt;

&lt;p&gt;And a missed closing step does not stay missed quietly. An ordinary task you skip just sits there until you come back to it; a closing step you skip stays open until something closes it, and until then it keeps acting, every day or every month, on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check has to come from somewhere you did not write
&lt;/h2&gt;

&lt;p&gt;So the thing that catches the missing step cannot be your own account of the work. It has to be a record that something else kept, for its own reasons, whether or not you remembered the step.&lt;/p&gt;

&lt;p&gt;You already have several of these. You just do not read them against the job. The card statement is one: the bank records the charge whether or not you remember the seat you meant to cancel, so the seat that is still billing turns up as a line you cannot attach to any live piece of work. The access list is another: the system logs who can get in whether or not anyone told it that a person left, so the account that outlived the project is a login with no current owner. What actually shipped is recorded by the thing that shipped it, so a promise you made and never delivered stands as a commitment on one side with no send on the other.&lt;/p&gt;

&lt;p&gt;Even with a checklist I take seriously, I did this. I keep a written routine for finishing an essay, detailed, with a warning next to the item that slips most, and my archive quietly slipped twenty-three pieces behind what I had published since late spring. Copying each finished piece across to that archive had never been a line on the routine at all: it lived on a different system, so it never occurred to me to write it down. What caught it was the published record of what had actually gone out, kept by the platform and owing nothing to my memory. Held against the archive, it showed the twenty-three at once.&lt;/p&gt;

&lt;p&gt;The move itself is old. Accountants have reconciled two sets of books this way for centuries, and there is nothing here to invent. What is easy to get wrong is what makes the second record worth anything: not that it is a second record, but that something other than your own memory produced it. Two dashboards drawn from the same database, or two lists built from the same picture of the job, only look like a check, because the same forgetting shaped both. A record can catch you only when your forgetting could not have reached it too.&lt;/p&gt;

&lt;h2&gt;
  
  
  One question to carry
&lt;/h2&gt;

&lt;p&gt;That gives you a single question, and it is worth more than any checklist. Of anything you lean on to tell you a job is finished, ask: would this still be here, and still say the same thing, if I had forgotten the step entirely?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg7p13pmia4i9rcxwb8is.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg7p13pmia4i9rcxwb8is.webp" alt="The test, applied to two records." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The test, applied to two records. The one you fill in yourself fails it; the one the bank writes passes it, because the charge is there whether or not you remembered.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The tracker fails that question: you fill it in yourself, so a step you forget is a step you also forget to log, and it stays green over a gap it never knew about. The card statement passes, because the charge is there whether or not you remembered the seat. A check built from your own memory cannot expose the step that memory left out.&lt;/p&gt;

&lt;p&gt;The question keeps its shape as the instrument gets bigger. A tracker, a dashboard, a report you write on your own project: each is an instrument you fill from your own picture of the work, and each is blind in the same place you are. The statement is worth more than any account you write of what you meant to do, for the same reason an audit leans hardest on evidence the audited side did not get to shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part you can use
&lt;/h2&gt;

&lt;p&gt;Here is the version you can run on your own job this week.&lt;/p&gt;

&lt;p&gt;Write out the routine you run after you finish something, every step, including the ones that feel too small to be worth writing down. Mark each with the system it touches: the tracker, the calendar, the billing account, the shared drive, the tool somebody set up before you arrived. This is not the check yet. It is how you find which records are worth reading against each other, and it usually turns what felt like one job into the three or four systems it was always made of.&lt;/p&gt;

&lt;p&gt;Then there are two ways to keep a step from being lost, and the first is much stronger. Where you can, do not rely on catching the step at all; arrange things so that forgetting it does no harm. Anything that runs on its own, a payment, a subscription, a recurring invite, an access granted for a single project, gets its end date on the day you set it up, while you still know what it was for. Something that expires unless it is renewed cannot outlast your forgetting, because forgetting it and ending it become the same act. Reach for this first; it removes the obligation instead of watching it. Its limit is the one this piece began with: you can only set an end date on a step you thought of, and the step that never made the list cannot be made self-closing.&lt;/p&gt;

&lt;p&gt;For everything you could not foresee, or cannot make expire, there is the slower move: read your own record against one you did not produce. Your active-projects list against the vendor or card statement, looking for a charge attached to work that has already finished. Your list of who is on the team against the access export from whatever holds the accounts, looking for a login with no owner. The commitments in a signed contract against what your team actually sent, looking for a promise with no matching send. Choose the second record by the causal test, not by where it happens to be stored: pick the one your own memory did not shape.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg8glohzs3yg4cw58y6q.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg8glohzs3yg4cw58y6q.webp" alt="Three records you keep, each read against one you did not write." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Three records you keep, each read against one you did not write. The last row is the limit: recorded nowhere, so nothing catches it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;How often you look depends on how much damage you will let build up first. A ten-pound seat can wait a month; a former colleague who can still open every file cannot, and something confidential still reaching the wrong person is not a scheduled job at all. None of it needs a tool you have to build: a read-only export or a screenshot is enough, and where you cannot pull the record yourself, the person who can is an email away, not a project. And reading across only points to a mismatch; you still have to look and decide whether it is a real miss, a timing lag or a duplicate, and keep that verdict for yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What none of this fixes
&lt;/h2&gt;

&lt;p&gt;Two things survive all of it, and I would rather say so than leave the tidy version standing.&lt;/p&gt;

&lt;p&gt;The first is the record that was never kept. If something gets finished and lands in no system at all, no charge, no log, no row anywhere, then there is no second record to read it against. You cannot check against a record that does not exist. That case surfaces only when a person happens to notice, or is told.&lt;/p&gt;

&lt;p&gt;The second is quieter, and more common. If the same blind spot sits in both records, they agree, and the agreement looks like an all-clear. This is the failure I walked into the first time I tried to build a check like this for myself. I searched my files for links to the publication. That sounds like reading an independent record, until you notice it read the same surface I would have: it counted the times I had linked to old pieces inside new ones as though that proved the old ones had shipped. Independence is the whole of the mechanism, and when it is missing it fails without a sound.&lt;/p&gt;

&lt;p&gt;So the honest tally is smaller than the tidy one. The obligations that leave a trace in a record I did not write, I can now catch, once in a while, in half an hour. The ones that touch nothing outside my own attention, I am still carrying in my head, and I have learned how little the word covers when I say nothing is wrong. &lt;em&gt;Nothing is wrong&lt;/em&gt; and &lt;em&gt;nothing I can see is wrong&lt;/em&gt; are different sentences, and most of the time only one of them is available to any of us.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=work-that-comes-due-after-you-leave" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fofu3ldypr4j1xnkibni8.webp" alt="The Day Job banner" width="799" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>analysis</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Memory Capacity Binds Before FLOPs Do: AgentX and the Agentic Inference Bottleneck</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Tue, 25 Aug 2026 20:09:07 +0000</pubDate>
      <link>https://dev.to/harryfloyd/memory-capacity-binds-before-flops-do-agentx-and-the-agentic-inference-bottleneck-20i8</link>
      <guid>https://dev.to/harryfloyd/memory-capacity-binds-before-flops-do-agentx-and-the-agentic-inference-bottleneck-20i8</guid>
      <description>&lt;h1&gt;
  
  
  Memory Capacity Binds Before FLOPs Do
&lt;/h1&gt;

&lt;p&gt;In agentic inference, memory capacity binds before FLOPs do. That is the first finding from AgentX, the open-source benchmark SemiAnalysis built for replaying real agentic coding traffic at one million context.&lt;/p&gt;

&lt;p&gt;The numbers on DeepSeek V4 make the point. The HBM working set decided the KV-cache hit rate: 43 million tokens and 91 percent on a B300, against 22 million and 73 percent on a B200. The B300 run used 384 concurrent traces, the B200 run 196. Same model, same task, different memory budget, and the hit rate moved by eighteen points.&lt;/p&gt;

&lt;p&gt;Why the hit rate matters: agentic sessions reuse prefixes. Every turn builds on the context before it, so most of the context can be served from the KV cache rather than recomputed. A miss means re-prefilling context you have already paid for once. In an agent loop of dozens of sequential calls, those misses compound into wall-clock time and compute you cannot get back.&lt;/p&gt;

&lt;p&gt;AgentX matters because it measures the right thing. It replays 393 anonymised internal Claude Code traces, structure and timing preserved, under Apache 2.0. Fixed-sequence benchmarks reflect chip and kernel performance; agentic workloads reflect the systems problem: KV tensors, routing, and offload across memory tiers. The first open-source benchmark of this shape has produced the finding the labs have been converging on: the binding constraint in agentic inference is not peak FLOPs, it is the size of the memory working set.&lt;/p&gt;

&lt;p&gt;The practical rule for anyone sizing an inference stack: size working set before FLOPs. A spec sheet that leads with peak throughput hides the number that actually determines your agent's latency. Measure step latency at your own concurrency, including prefill, before you trust a batch-of-one headline.&lt;/p&gt;

&lt;p&gt;The pattern is durable. When compute gets cheap enough, the constraint migrates to the layer beneath it, the same migration the industry has watched in every previous scaling phase. AgentX is the first open benchmark to show where it has landed in agentic inference: not in the silicon that generates tokens, but in the memory that holds the conversation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>infrastructure</category>
      <category>analysis</category>
    </item>
    <item>
      <title>You're Not Comparing Models. You're Comparing Contracts.</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:42:42 +0000</pubDate>
      <link>https://dev.to/harryfloyd/youre-not-comparing-models-youre-comparing-contracts-647</link>
      <guid>https://dev.to/harryfloyd/youre-not-comparing-models-youre-comparing-contracts-647</guid>
      <description>&lt;h1&gt;
  
  
  You're Not Comparing Models. You're Comparing Contracts.
&lt;/h1&gt;

&lt;p&gt;Two teams publish scores on the same agent benchmark.&lt;br&gt;&lt;br&gt;
One lands in the low sixties. The other clears seventy.&lt;br&gt;&lt;br&gt;
A procurement team reads the spread and makes a call.&lt;/p&gt;

&lt;p&gt;What they do not see: both teams may be running the same model. They did not need to change the weights for the gap to appear. The spread can come from scaffold alone.&lt;/p&gt;

&lt;p&gt;One team wrapped the model in a harness with better retries. Different tool defaults. A planner step the other team had skipped. None of that appears on the leaderboard.&lt;/p&gt;

&lt;p&gt;The comparison that drove the decision was not between two agents.&lt;/p&gt;

&lt;p&gt;It was between two contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;There Is No Benchmark&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The mistake hiding behind this story is a category error.&lt;/p&gt;

&lt;p&gt;People talk about agent benchmarks as if they measure a thing called “the model.” They do not. They measure a coupled system. The model is one component. The rest is a stack of protocol decisions that are almost never disclosed and almost always matter.&lt;/p&gt;

&lt;p&gt;The score is the output of that stack. Change any layer and you change what the number means.&lt;/p&gt;

&lt;p&gt;Recent research on agent evaluation has named those layers explicitly. There are at least seven. Deployment regime. Observation channel. Harness and scaffold. Metric and action. Configured evaluator. Grader protocol. Audit bundle. Each is a contract. Each is negotiable. And each can silently change the verdict while the headline looks the same.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;That is what a benchmark actually is. Not a measurement of a model. A measurement of an entire testing contract, of which the model is one slot.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is structural reason the seven layers are the seven layers. They cluster into three corners that show up in almost every published agent-evaluation failure. What the model is rewarded for. How that reward is optimised. And how the test contract differs from production. Once you hold those three corners in view, the seven-layer stack stops feeling like a checklist and starts behaving like the actual shape of what is being measured.&lt;/p&gt;

&lt;p&gt;If you are comparing agent products without parity across those layers, you are not comparing agents.&lt;/p&gt;

&lt;p&gt;You are comparing contracts and calling it science.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Harness You Didn’t Name&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The most visible layer, and the one that moves the most points, is the scaffold.&lt;/p&gt;

&lt;p&gt;Anyone who has built an agent in the last year has felt this without naming it. You watch a coworker get 75% on a task your model just failed on. You check the weights. They are yours. They changed the prompt template and added a retry loop. The model did not get smarter. The scaffold got thicker.&lt;/p&gt;

&lt;p&gt;The numbers say the same thing. RWE-bench reports that on its 162-task benchmark over MIMIC-IV, changing only the agent scaffold around a fixed model can shift performance by more than 30 percent1. Same weights. Different tools. Different retry policy. Different planner. Different headline. The best evaluated agent on that benchmark reaches around 40 percent task success at all; the best open-source configuration is closer to 30. Once you know the contract can move 30 points on its own, neither of those numbers is really about a model.&lt;/p&gt;

&lt;p&gt;If scaffold alone can move scores by double digits, then “we used the same model as them” is not a fair-comparison claim. It is a parameter-naming claim. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You have named one slot in a seven-slot contract.&lt;br&gt;&lt;br&gt;
The other six are doing most of the work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Judge That Isn’t The Model&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The second layer that silently moves scores is the evaluator itself.&lt;/p&gt;

&lt;p&gt;When a benchmark uses an LLM judge, people write things like “graded by GPT-4o” as if that pins the measurement down. It does not. The judge is not GPT-4o. The judge is GPT-4o plus a prompt template. Plus a decoding configuration. Plus a tie and abstention policy. Plus whatever retrieval or tool access the judge has during grading. None of that ships with the score.&lt;/p&gt;

&lt;p&gt;A recent systematic evaluation of LLM-as-judge setups showed that prompt-template choice alone materially changes both judge quality and internal consistency2. Two teams reporting “we used GPT-4o as judge” can be running substantively different graders. The grader that rewards epistemic hedging disagrees with the grader that penalises it. The grader with access to retrieval checks factuality. The grader without one does not, and cannot.&lt;/p&gt;

&lt;p&gt;This is not a small print issue. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The evaluator is the measuring instrument. If two teams use different instruments and report the same number, they are not reporting the same thing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And without a published judge card, no third party can reproduce the measurement. They can only rerun the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Number That Lies About Consistency&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The third layer is the quietest and most dangerous. It is the metric itself.&lt;/p&gt;

&lt;p&gt;A standard agent metric is &lt;a href="mailto:pass@k"&gt;pass@k&lt;/a&gt;. You give the agent k attempts. If any one succeeds, it counts. This is perfectly reasonable if your production use allows k attempts. It is actively misleading if it does not.&lt;/p&gt;

&lt;p&gt;There is a sibling metric, pass^k. Same k attempts. But it only counts if the agent succeeds on all of them. It measures consistency, not capability.&lt;/p&gt;

&lt;p&gt;The gap between these two can be large, and it can open silently.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Recent work on trustworthy agent evaluation shows that controlled error injection into an agent can cut pass^k substantially while barely moving &lt;a href="mailto:pass@k3"&gt;pass@k3&lt;/a&gt;. The model still has a ceiling you can hit with enough tries. It has lost the ability to hit that ceiling reliably. If your headline is pass@k and your production regime is one shot, the leaderboard says you are shipping. The bug tracker says otherwise.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same structural problem appears in calibration metrics. ECE asks whether stated probabilities match empirical frequencies on average. AURC asks whether the system can rank harder cases lower. Both can look nearly identical across two systems while a stricter, abstention-aware metric called BAS, the Behavioural Alignment Score, diverges sharply between them4. BAS asks a different question. Does the confidence surface protect you in exactly the regime where a person or product would actually choose to trust it? Two systems with “similar calibration” can answer that question completely differently once you attach a cost function.&lt;/p&gt;

&lt;p&gt;The metric is not a measurement of the model. It is a statement about which errors the model’s operators will tolerate. If that statement does not match your operational contract, the score is not wrong. It is answering a question you did not ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Why Rank Stability Is A Trap&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Here is the part that makes all of this subtly worse.&lt;/p&gt;

&lt;p&gt;Under scaffold shift, the rank order of agents on a benchmark is often relatively stable. A recent efficient-benchmarking study reports that rank preservation is easier to maintain than absolute calibration5. The number moves. The ordering does not.&lt;/p&gt;

&lt;p&gt;If all you need is a relative decision, rank stability is comforting. Agent A beats Agent B here, and probably beats it in production.&lt;/p&gt;

&lt;p&gt;If you need an absolute decision, it is a trap.&lt;/p&gt;

&lt;p&gt;Procurement, safety arguments, SLA setting, cost modelling, and risk disclosure all depend on absolute numbers. A claim like “this agent ships 80% correct at 5 cents per request” binds to the calibrated level, not to the rank. Under scaffold shift, rank can hold while the 80% becomes 62%. Your spreadsheet is still using 80%. Your customers are experiencing 62%.&lt;/p&gt;

&lt;p&gt;The protocol that produced 80% is part of the claim. The moment it diverges from production, the claim silently becomes false, even though nothing about the model moved.&lt;/p&gt;

&lt;p&gt;This is why seasoned eval teams treat the contract, not the score, as the primary artefact. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You can rerun a score. You can only reproduce a contract if you wrote it down.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Contract Is The Object&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;If the score is a function of the contract, the practical move is to treat the contract as the thing you own.&lt;/p&gt;

&lt;p&gt;That means three changes to how most teams currently work.&lt;/p&gt;

&lt;p&gt;Freeze the contract before you compare. If you cannot describe your deployment regime, observation channel, scaffold version, metric, judge configuration, and grader protocol in one page, you do not have a contract. You have assumptions pretending to be one. Write the page. Commit it. Make it a prerequisite for every comparison.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Your Tools Got Powerful. Get Boring.</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:42:26 +0000</pubDate>
      <link>https://dev.to/harryfloyd/your-tools-got-powerful-get-boring-48jn</link>
      <guid>https://dev.to/harryfloyd/your-tools-got-powerful-get-boring-48jn</guid>
      <description>&lt;h1&gt;
  
  
  Your Tools Got Powerful. Get Boring.
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?" rel="noopener noreferrer"&gt; Subscribe now&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;The bored trader beats the machine&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;On one side of the trade sits a market-making engine that represents the genuine state of the art: Hawkes processes modelling order arrivals, Kyle’s lambda pricing the impact of each fill, Avellaneda-Stoikov inventory control balancing the book in real time. Years of mathematics, running on hardware that did not exist a decade ago.&lt;/p&gt;

&lt;p&gt;On the other side is a momentum trader whose entire system is price, volume, and three moving averages. He sits in cash most of the year doing nothing, waiting for a setup he could describe to you in a sentence. His stack is deliberately primitive. His edge is patience and the discipline to follow his own rules when they are boring and to sit out when they are silent.&lt;/p&gt;

&lt;p&gt;Over a full market cycle, the boring one is more likely to still be standing.&lt;/p&gt;

&lt;p&gt;This is uncomfortable, because it runs against an intuition almost everyone shares: better tools should let you run better, more sophisticated strategies. More compute, more data, more powerful models, therefore more elaborate approaches and better results. It feels obviously true. It is the logic behind most of what gets built, bought, and bragged about.&lt;/p&gt;

&lt;p&gt;It is also, across domain after domain, wrong. And the interesting part is the shape of the curve.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;The gap widens as the tools get stronger&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Here is the pattern the most successful practitioners keep seeing, whether they are trading, building software, learning, or shipping products. Powerful tools do not pay off when you point them at more complex strategies. They pay off when you point them at simple strategies and execute those faster, more consistently, and with less drift than anyone else.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;More power applied to a simple strategy compounds. The same power applied to a complex one mostly buys you more ways to be wrong.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sit with the second half of that, because it is the part people miss. A sophisticated strategy is not free. Every additional layer needs to be specified, verified, maintained, and monitored, and all of that consumes exactly the capacity the powerful tool was supposed to give back. A simple strategy spends its new power on doing the simple thing relentlessly well. A complex one spends its new power feeding its own machinery.&lt;/p&gt;

&lt;p&gt;The reason this matters more now than it ever has is that the tools have never been this strong. When your instruments are weak, the gap between the simple-and-disciplined path and the complex-and-fragile path is small, because nobody can do much of either. As the instruments get more powerful, both paths open up, and the distance between them widens. The most capable tools in history make disciplined simplicity more effective than ever, and they also make unmanageable complexity easier to build than ever. We are living through the largest gap between those two paths that has ever existed, and most people are sprinting down the wrong one with a faster engine.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What complexity quietly costs&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The bill for sophistication does not arrive when you build it. It arrives later, in instalments, and it is always larger than it looked.&lt;/p&gt;

&lt;p&gt;The first instalment is verification. A simple system you can hold in your head and check. A complex one you cannot, so you build monitoring to watch it, and the monitoring becomes its own system that can &lt;a href="https://harryfloyd.substack.com/p/most-verification-is-just-bigger" rel="noopener noreferrer"&gt;drift and mislead&lt;/a&gt;. Every layer you add is a layer you now have to confirm is still doing what you think it does, and the confirming never ends.&lt;/p&gt;

&lt;p&gt;The second instalment is the day it breaks. A simple strategy fails legibly: you can see which rule was wrong and fix it. A sophisticated one fails in the seams between its parts, at the worst possible moment, in a way no single person fully understands. The elaborate model that printed money for two years becomes, in the drawdown, a black box nobody can debug while it is bleeding. Complexity does not only add capability. It adds failure modes that stay hidden until the system is under stress, which is the exact moment you have no spare capacity to handle them.&lt;/p&gt;

&lt;p&gt;The deepest cost is fragility to your own success. A strategy with many parameters has many surfaces the world can destabilise once it starts reacting to you. The more elaborate the machine, the more places reality can reach in and pull a lever you forgot you had wired up. Simple, constrained systems survive contact with the world because there is less of them to break.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Sophistication is a loan against your future attention, taken out at a rate you cannot see until the system is under stress and the whole balance comes due at once.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?" rel="noopener noreferrer"&gt; Subscribe now&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Why we reach for sophistication anyway&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;If simplicity wins, why does almost everyone instinctively add complexity? Smart, capable people do it constantly, because the incentives reward it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every incentive in the room rewards the complexity you can show and punishes the discipline you cannot.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sophistication is visible. A complex model, an elaborate architecture, a clever framework can be shown to a boss, a client, an investor, a peer. Discipline cannot be shown. Sitting in cash for three months, deleting half your code, pausing before you speak, refusing to ship the extra feature: none of it photographs well. The market pays for what it can see, and it can see complexity far more easily than it can see restraint.&lt;/p&gt;

&lt;p&gt;Complexity also feels like work. Building an intricate system produces the sensation of progress all day long, even when the effort is going into &lt;a href="https://harryfloyd.substack.com/p/the-leverage-hierarchy-of-agent-engineering" rel="noopener noreferrer"&gt;the layer with the least leverage&lt;/a&gt;. Doing the boring, correct thing and then waiting produces the sensation of doing nothing, which the nervous system reads as failure. The feeling and the result point in opposite directions, and the feeling usually wins.&lt;/p&gt;

&lt;p&gt;And an entire economy is built on convincing you the work is harder than it is. Every tool vendor, every course, every consultancy has a structural interest in making its domain look more complex than it needs to be, because simplicity is terrible for business. The people who write about a field emphasise its hardest parts, which is what makes them experts, rather than its simplest parts, which is what produces the results. The perceived difficulty of almost everything is inflated, and the inflation is nobody’s accident.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What the constrained version keeps proving&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The clearest place to watch this play out right now is in how people use AI, because the tool is so powerful that the trap is stark.&lt;/p&gt;

&lt;p&gt;The most effective way to get good work out of a frontier model is to take capability away from it. The prompts that consistently produce strong code are the ones that forbid things: no verbose comments, no scattered logging, small functions only, review your own output before returning it. The best debugging prompts are the most constrained ones: strict ordered steps, and a hard rule to verify before changing anything. The most powerful model on the planet does better work when you give it fewer options. People reach for AI expecting more power to mean more freedom. What it rewards is more power inside tighter constraints.&lt;/p&gt;

&lt;p&gt;The same shape shows up wherever someone is quietly winning with powerful tools. The builders who ship profitable products solo run on deliberately boring technology, the kind a fashionable engineer would be embarrassed by. They ship ugly first versions fast while better-resourced teams are still choosing a framework. The plain name for what those teams are doing is over-engineering, and the powerful tools make it easier than ever. The people who learn fastest take fewer notes, not more. They delay and compress until a page of dense understanding replaces a folder of neat transcription. The creators who grow post less, because the algorithm rewards depth per post and punishes the volume that easy tools make tempting. Different fields, one lesson: the powerful tool is best spent removing steps.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The people quietly winning with the strongest tools are using them to do less, and to do it more reliably than anyone else.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;None of these people are anti-technology. They are using the most powerful tools available. They are simply pointing them at the boring fundamentals and refusing the upgrade to a more complicated game.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>machinelearning</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Your AI Agent Stack Is Solving The Wrong Problem</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:42:10 +0000</pubDate>
      <link>https://dev.to/harryfloyd/your-ai-agent-stack-is-solving-the-wrong-problem-2ni5</link>
      <guid>https://dev.to/harryfloyd/your-ai-agent-stack-is-solving-the-wrong-problem-2ni5</guid>
      <description>&lt;h1&gt;
  
  
  Your AI Agent Stack Is Solving The Wrong Problem
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The setup everyone is sharing&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Which MCP servers to install. Which skills to keep in your repo. Which agent framework to use. How to write your &lt;code&gt;AGENTS.md&lt;/code&gt;. How to split one agent into researcher, planner, coder, and reviewer. How to wire Slack, GitHub, Notion, Postgres, Stripe, your calendar, and your file system into one increasingly capable loop.&lt;/p&gt;

&lt;p&gt;Some of that advice is useful. It is also aimed at the wrong layer.&lt;/p&gt;

&lt;p&gt;What becomes real after the agent uses a tool matters more than whether it can reach the tool.&lt;/p&gt;

&lt;p&gt;Can it read the customer record, or change it? Can it draft the refund, or issue it? Can it open a pull request, or merge it? Can it propose the vendor response, or send it under the company name?&lt;/p&gt;

&lt;p&gt;Once an agent can act through tools, the real system is no longer the model.&lt;/p&gt;

&lt;p&gt;The real system is the contract stack around the model.&lt;/p&gt;

&lt;p&gt;That is the part most setup guides skip.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Access is reach. Agency is permissioned action.&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine the demo.&lt;/p&gt;

&lt;p&gt;The agent can read Slack. It can search email. It can query the CRM. It can open GitHub issues, check billing records, browse docs, edit a spreadsheet, draft a customer reply, and call three internal APIs.&lt;/p&gt;

&lt;p&gt;Everyone in the room calls it powerful.&lt;/p&gt;

&lt;p&gt;That is the first mistake.&lt;/p&gt;

&lt;p&gt;The agent has reach. It does not yet have governed agency.&lt;/p&gt;

&lt;p&gt;Access tells you what the agent can touch. Agency tells you what the agent is authorised to decide, under which conditions, with what proof, and with what consequence after failure.&lt;/p&gt;

&lt;p&gt;That distinction sounds small until the first bad run.&lt;/p&gt;

&lt;p&gt;A read-only research assistant can waste time. An agent with billing access can create obligations. An agent with email access can speak for the company. An agent with deployment access can turn a wrong inference into infrastructure.&lt;/p&gt;

&lt;p&gt;More tools do not automatically make the agent more agentic.&lt;/p&gt;

&lt;p&gt;More tools expand the surface on which judgement has to be engineered.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The tool stack is visible. The contract stack is load-bearing.&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The visible agent stack is easy to list: model, prompt, memory, tools, MCP servers, subagents, framework, evals.&lt;/p&gt;

&lt;p&gt;That stack matters. It is also not the operating system.&lt;/p&gt;

&lt;p&gt;The operating system is the set of contracts each layer creates.&lt;/p&gt;

&lt;p&gt;What is the agent for? What state may it see? What state may it preserve? Which tools may it call? Which tools are intentionally absent? What can it change? What must it prove before the change becomes binding? What does the harness log? What does the evaluation score actually cover? When does the agent ask, abstain, or escalate? What permission disappears after a bad run?&lt;/p&gt;

&lt;p&gt;That is the real setup.&lt;/p&gt;

&lt;p&gt;Not the list of tools.&lt;/p&gt;

&lt;p&gt;The set of boundaries that decides what the tools mean.&lt;/p&gt;

&lt;p&gt;The generic setup stack asks what you connected.&lt;/p&gt;

&lt;p&gt;The contract stack asks what you can trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;An agent is a control loop, not a prompt with ambition&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;An agent is an outer control loop wrapped around a generator. It plans, reads state, chooses tools, acts, observes, repairs, escalates, and decides whether to continue.&lt;/p&gt;

&lt;p&gt;The failure rarely sits in one glamorous place.&lt;/p&gt;

&lt;p&gt;It can sit in the planner. It can sit in retrieval. It can sit in a tool description. It can sit in retry logic. It can sit in a hidden assumption about whether the world waits while the agent thinks.&lt;/p&gt;

&lt;p&gt;That is why framework comparisons are often less useful than they look.&lt;/p&gt;

&lt;p&gt;The distinction that matters is which parts of the loop are explicit enough to inspect.&lt;/p&gt;

&lt;p&gt;If planning is hidden inside one long natural-language instruction, you cannot repair planning without rewriting the whole prompt.&lt;/p&gt;

&lt;p&gt;If memory is just a growing transcript, you cannot tell whether the agent remembered, retrieved, inferred, or hallucinated.&lt;/p&gt;

&lt;p&gt;If tool choice is unlogged, you cannot tell whether the answer is wrong because the model reasoned badly or because it called the wrong thing.&lt;/p&gt;

&lt;p&gt;If evaluation is one final pass/fail number, you cannot tell whether the agent failed at discovery, parameters, sequencing, recovery, escalation, or judgement.&lt;/p&gt;

&lt;p&gt;Agents do not become reliable when the setup becomes more impressive.&lt;/p&gt;

&lt;p&gt;They become reliable when failure has somewhere specific to land.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;MCP is not magic glue&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;MCP matters. Skills matter. Connectors matter.&lt;/p&gt;

&lt;p&gt;But their importance is often described backwards.&lt;/p&gt;

&lt;p&gt;The lazy version says MCP is valuable because it gives agents more tools.&lt;/p&gt;

&lt;p&gt;The better version says MCP is valuable because it makes the tool boundary explicit enough to inspect, version, test, authorise, and debug.&lt;/p&gt;

&lt;p&gt;A tool is not neutral plumbing. A tool description tells a nondeterministic system what an action means. The name, parameters, return shape, error messages, and allowed mutations all change behaviour.&lt;/p&gt;

&lt;p&gt;A tool built for a human developer is not automatically a good tool for an agent. Humans carry missing context. Agents need the contract written down.&lt;/p&gt;

&lt;p&gt;That is why more tools can make an agent worse.&lt;/p&gt;

&lt;p&gt;At small scale, tool access feels like freedom. At larger scale, tool access becomes search. The agent has to identify the right tool, pass valid parameters, recover from partial failure, and avoid inventing a successful trace when the tool call failed.&lt;/p&gt;

&lt;p&gt;If you expose every API endpoint as a tool, you do not have a powerful agent surface.&lt;/p&gt;

&lt;p&gt;You have a vocabulary problem with write access.&lt;/p&gt;

&lt;p&gt;The mature move is not “connect everything.”&lt;/p&gt;

&lt;p&gt;The mature move is to design the smallest tool surface that lets the agent do the job, then make every tool contract legible. What does the tool do. When should it be used. What the return value proves, and what it does not prove. What failures look like. Which calls are read-only, which mutate state, which require approval. Where the trace goes.&lt;/p&gt;

&lt;p&gt;That is how you stop a transcript from becoming the only place your operating system exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Skills are not prompt snippets&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The same mistake happens with skills.&lt;/p&gt;

&lt;p&gt;People treat skills as better prompts: a &lt;code&gt;SKILL.md&lt;/code&gt;, a few examples, some instructions, maybe a script. Useful. Portable. Easy to share.&lt;/p&gt;

&lt;p&gt;But a serious skill is not a prompt snippet.&lt;/p&gt;

&lt;p&gt;It is packaged operating knowledge.&lt;/p&gt;

&lt;p&gt;It should contain a trigger, a procedure, a boundary, gotchas, and a failure mode.&lt;/p&gt;

&lt;p&gt;The “gotchas” are usually the most valuable part. The model often already knows the happy path. What it does not know is your local scar tissue: which API lies, which file must not be edited, which naming convention breaks deployment, which customer segment changes the policy.&lt;/p&gt;

&lt;p&gt;That is why generic skill catalogues have a ceiling.&lt;/p&gt;

&lt;p&gt;They can teach a model the common workflow.&lt;/p&gt;

&lt;p&gt;They cannot teach it which parts of your workflow are load-bearing unless you package that knowledge yourself.&lt;/p&gt;

&lt;p&gt;Skills are valuable because they let operational knowledge travel across sessions and agents. They are dangerous when they activate at the wrong time, compose implicitly into deeper graphs nobody intended, or grant state-changing behaviour without a permission contract.&lt;/p&gt;

&lt;p&gt;The real question is sharper:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When this skill activates, what decision is it allowed to influence?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If nobody can answer that, the skill is just a more durable way to make the wrong move.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Memory is governed state, not a bigger past&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Memory has the same problem.&lt;/p&gt;

&lt;p&gt;Every agent product wants to promise memory. It sounds obvious. The agent should remember the user, the project, the codebase, the customer history, the prior decision, the mistake from last time.&lt;/p&gt;

&lt;p&gt;But memory is not “more context.”&lt;/p&gt;

&lt;p&gt;Memory is a four-part contract: what gets written, how it is organised, how it is retrieved, how it is governed.&lt;/p&gt;

&lt;p&gt;If the agent writes too much, memory becomes sludge.&lt;/p&gt;

&lt;p&gt;If it summarises badly, memory becomes distortion.&lt;/p&gt;

&lt;p&gt;If it retrieves by similarity alone, memory becomes vibes with citations.&lt;/p&gt;

&lt;p&gt;If it never forgets, memory becomes context poisoning.&lt;/p&gt;

&lt;p&gt;If it cannot show why a memory was used, memory becomes an invisible authority.&lt;/p&gt;

&lt;p&gt;The memory question worth asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which state should survive because it will improve future decisions, and which state should expire because it will poison them?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a contract question.&lt;/p&gt;

&lt;p&gt;It is also why a 500-word, well-maintained project note can outperform a giant chat history. The smaller note has a job. The transcript merely has volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The harness is where autonomy becomes measurable&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Most agent demos make the model look like the protagonist.&lt;/p&gt;

&lt;p&gt;In production, the harness is the protagonist.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>machinelearning</category>
      <category>architecture</category>
    </item>
    <item>
      <title>You Only Hold Four Thoughts</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:55 +0000</pubDate>
      <link>https://dev.to/harryfloyd/you-only-hold-four-thoughts-2p7j</link>
      <guid>https://dev.to/harryfloyd/you-only-hold-four-thoughts-2p7j</guid>
      <description>&lt;h1&gt;
  
  
  You Only Hold Four Thoughts
&lt;/h1&gt;

&lt;p&gt;Try to multiply 47 by 83 in your head. The answer is not the point. Watch what happens while you reach for it. You hold 47, you hold 83, you start on the partial products, and somewhere around the third one the first number goes soft. You reach for a pen, because the problem outgrew the place you were keeping it.&lt;/p&gt;

&lt;p&gt;That ceiling is real and it is low. The cognitive scientist Nelson Cowan spent years measuring it and put the number at about four. Not the seven you half-remember from an old paper, but three to five distinct things held in mind at once. 1 Four. That is the working capacity of the most sophisticated object in the known universe.&lt;/p&gt;

&lt;p&gt;Everything we call getting smarter has been a way around that four. The history of human intelligence is the history of putting thoughts somewhere other than the head, and it runs as a stack, each layer holding what the one below it cannot.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The first rung is paper&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Reaching for the pen looks like a small surrender. It is the oldest cognitive upgrade there is. The moment you write 47 above 83 and start stacking partial products, you are thinking about six or seven things at once, because the paper is holding all but the one you are working on.&lt;/p&gt;

&lt;p&gt;Justin Sung, who teaches learning for a living, puts it more sharply. Writing is not the thing you do after you have reached clarity. Writing is what produces the clarity. 2 The page becomes the workspace where the thought turns real, because your four slots are freed to do the actual reasoning while the page remembers the rest.&lt;/p&gt;

&lt;p&gt;This is also why handwriting beats typing. It is far slower than thinking, and that slowness forces you to compress, to decide what is worth the stroke. The friction is not a tax on the process. The friction is the process. A page of notes you struggled to write holds more than a page you copied without resistance.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The page is not a transcript of a finished thought. It is the workspace where the thought becomes possible.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The rung most people never name&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;In 1998 two philosophers, Andy Clark and David Chalmers, asked where the mind stops and the rest of the world begins, and gave an answer that still unsettles people. The mind, they argued, is not all in the head. 3&lt;/p&gt;

&lt;p&gt;Their example was a man named Otto, who has Alzheimer’s and carries a notebook everywhere. When Otto wants to go to the museum, he looks up the address in the notebook the way you would retrieve it from memory. The notebook does the job your hippocampus does. Clark and Chalmers argued there is no principled reason to count the notebook as any less a part of Otto’s mind than ordinary memory. Otto and his notebook are a single coupled system. The thinking happens across both.&lt;/p&gt;

&lt;p&gt;That sounds like a thought experiment until you notice you are Otto. The phone that holds every number you no longer memorise. The calendar that holds every commitment. The thinking is already distributed across you and the things you store it in. The only open question is how well the storage is built.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The rung that compounds&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A single page does not persist, does not connect, and cannot be searched. You solve the multiplication, you throw the page away, and next month you solve it again from scratch. Paper extends the moment. It does not extend across time.&lt;/p&gt;

&lt;p&gt;A structured set of notes does. When every thought you have is written as a durable, cross-linked entry, two things happen that a single page cannot. The thought survives, available to a version of you who has forgotten having it. And it connects, so that an idea from March sits one link away from a problem you only encounter in June, waiting to be useful before you knew you needed it.&lt;/p&gt;

&lt;p&gt;This is the layer where synthesis becomes possible at a scale no head can hold. No one can keep thirty sources in working memory and find the pattern across them. Four slots cannot do it, and neither can forty. But a system that has been accumulating those sources for months, with the connections already drawn, can surface a synthesis that was never available to anyone thinking alone. The structure does the remembering, which frees the human to do the seeing.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The rung we are building now&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;For most of history the top of the stack was a human reading their own notes. That is no longer the ceiling. The newest layer is a store of knowledge an AI can read, query, and build on across sessions.&lt;/p&gt;

&lt;p&gt;The builders who have lived inside this for a year keep reporting the same thing. One who runs large agent systems put it plainly: the model is the same on day 1 and day 40. The files get richer. 4 The capability of the underlying intelligence barely moves over a project. What improves is the accumulated context it can reach, the record of what was tried, what worked, what the operator decided and why. The intelligence is rented and roughly fixed. The memory is owned and compounds.&lt;/p&gt;

&lt;p&gt;An AI working from a thin prompt starts every session as a stranger. An AI working from a well-kept store of your decisions starts as a colleague who was in the room last time. The difference is not a better model. It is the same model with the rest of the stack underneath it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The intelligence is rented and roughly fixed. The memory is owned, and the memory is what compounds.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Why this is one law and not four&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This is where a productivity story becomes something larger. The machines climb the same stack you do, for the same reason, using the same move.&lt;/p&gt;

&lt;p&gt;A large model also cannot hold everything at once. Its version of the four-slot limit is the memory bandwidth of the chip, and the entire recent history of making models faster is a history of refusing to keep everything hot. FlashAttention rewrote how attention uses memory so the chip stops shuttling the same data back and forth. Key-value caching stores the work already done so it never has to be recomputed. Mixture-of-experts routing keeps a vast model mostly dormant and wakes only the part a given token needs. 5 Store state. Reuse it. Activate only what matters now.&lt;/p&gt;

&lt;p&gt;That is the same move as paper, notes, and agent memory. Externalise the state you cannot hold, and retrieve only the slice the moment requires. Human cognition scales that way. Machine cognition scales that way. The question “how do I think better” and the question “how do I run a model well” have turned out to be one question with one answer. When two separate problems collapse into the same answer, that answer is usually worth trusting.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One law runs the whole stack: externalise the state you cannot hold, and retrieve only what the moment needs. Brains and models both scale by obeying it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The trap inside the stack&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The law has a failure mode, and it is the one a second-brain enthusiast walks into first. The stack rewards retrieval, not accumulation. The instant you start optimising for the volume of what you store, you have begun to degrade the thing you were building.&lt;/p&gt;

&lt;p&gt;A note you never pull back out did no cognitive work. Ten thousand of them do less than a hundred you reach for, because the ten thousand bury the hundred. External cognition only pays off on the way back in. Storing is filing, and filing is not thinking. The discipline that keeps the stack alive is structuring everything you save so a future you, or a future agent, can find the one piece that matters without reading the other nine thousand.&lt;/p&gt;

&lt;p&gt;That is also why each rung has to be built in order. Agent memory on top of a disorganised pile of notes inherits the disorder and answers your questions confidently from a mess. The layers compound only when each one is sound. Skip a rung and you do not get the compounding. You get a faster way to retrieve noise.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The test you can run this week&lt;/strong&gt;
&lt;/h3&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>productivity</category>
      <category>analysis</category>
    </item>
    <item>
      <title>You Cannot Try to Fall Asleep</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:39 +0000</pubDate>
      <link>https://dev.to/harryfloyd/you-cannot-try-to-fall-asleep-2ced</link>
      <guid>https://dev.to/harryfloyd/you-cannot-try-to-fall-asleep-2ced</guid>
      <description>&lt;h1&gt;
  
  
  You Cannot Try to Fall Asleep
&lt;/h1&gt;




&lt;p&gt;It is ten past three. You have done the arithmetic twice already. Five hours if you drop off now, four and a bit if this carries on. So you lie very still, because turning over would be an admission, and you hold your eyes shut a little too tightly, and underneath all of it there is a low, steady wanting. You want to be asleep. And the wanting is the exact thing keeping you awake.&lt;/p&gt;

&lt;p&gt;You are failing at doing nothing. It is a strange thing to be bad at.&lt;/p&gt;

&lt;p&gt;You are a capable person. You can learn hard things and finish dull ones and drag yourself out for a run on a wet morning when every part of you would rather stay in. Effort is the most reliable tool you own. Most of what you are proud of came out of using it. And then there is this one ordinary thing, wanted more than almost anything at three in the morning, that effort cannot touch at all. The harder you try for it, the further away it goes.&lt;/p&gt;

&lt;p&gt;We file that under sleep being difficult and move on. But sleep is not the only thing built this way. It is only the place you notice it first, because you meet it every single night.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The wanting is the exact thing keeping you awake.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Try to stop thinking about someone, and watch what happens to the thinking. Try to be happy, directly, by deciding to be, and feel it thin out into a performance of itself. You cannot make yourself find a joke funny. You cannot force another person to love you by loving them harder; if anything, that is the surest way to send them off. You cannot decide to be interesting at the party, and the second you try, you are the least interesting you will be all evening. You cannot will yourself to relax, which is the cruellest one, because the trying is the tension.&lt;/p&gt;

&lt;p&gt;None of these are things you do. They are things that happen to you while you are busy doing something else. Sleep arrives while you are turning over tomorrow’s meeting, and then at some point you never quite catch, you are gone. Happiness turns up on an ordinary afternoon when you were absorbed in something and forgot to check whether you were happy. You become interesting the moment you get genuinely interested in someone else. Love shows up sideways, in the middle of doing something entirely unromantic together. Each one is a by-product. The main thing was always something else.&lt;/p&gt;

&lt;p&gt;None of this is new. The Victorians had a name for the trap. 1 They called it the paradox of hedonism, the plain observation that happiness tends to arrive only when your mind is fixed on something other than your own happiness. Aim straight at it and you miss. Aim at something worth doing and it arrives while your back is turned. Older and gentler still is the folk wisdom your grandmother had. A watched pot never boils. You will meet someone when you stop looking. Sleep comes when you stop chasing it. She was right, and she never needed a footnote.&lt;/p&gt;

&lt;p&gt;Which makes it strange that almost everything around us now says the opposite. Try harder. Optimise. Measure it, track it, put a number on it, and by watching the number, improve it.&lt;/p&gt;

&lt;p&gt;For plenty of things, that advice is sound. It genuinely works on the steps you walk, the pages you read, the money you put aside, because those are things you do, and a thing you do answers to attention and effort. Point a number at a behaviour and the behaviour usually moves.&lt;/p&gt;

&lt;p&gt;The trouble starts when we point the same instrument at something that was never a behaviour. Take the sleep tracker, the small clean example of a very large mistake. It hands you a grade out of a hundred each morning for a thing you did not do, could not have done, and had no control over while it was happening. For someone already anxious about their sleep, that grade can do exactly what you would dread. There is a name now for people whose pursuit of a better sleep score has quietly made their sleep worse. 2 They lie there trying to earn the number, and the trying keeps them up, and in the morning the number confirms the bad night, so tomorrow they try harder still. The scoreboard becomes the insomnia.&lt;/p&gt;

&lt;p&gt;And it does not stop at sleep. We keep a scoreboard on our own happiness now, rating the day, wondering whether we are as content as we ought to be by this point, which is a reliable way to stop being content at all. We tally our friendships, our rest, our worth, treating the whole illegible middle of a life as figures to be raised. A thing that arrives sideways cannot survive being stared at head-on. Measured, it becomes work. Graded, it becomes a test you are failing. You get all the pressure of a target and none of the thing the target was for.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Measured, it becomes work. Graded, it becomes a test you are failing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So what do you actually do, if trying is the problem and not-trying sounds like giving up?&lt;/p&gt;

&lt;p&gt;You learn, slowly and against every instinct this age has trained into you, to tell two kinds of thing apart.&lt;/p&gt;

&lt;p&gt;Some things in your life answer to effort. You can decide the hour you go to bed. You can decide the phone leaves the room. You can decide to show up, to sit down at the desk, to be kind without keeping score, to call your mother, to put yourself in the path of the people you might one day come to love. Those are conditions. Conditions are real work, and they are what you can actually reach.&lt;/p&gt;

&lt;p&gt;And then there is everything the conditions are for. Sleep. Ease. Delight. Being loved. Feeling rested. Those you cannot reach for. You can only build the conditions, honestly and without cheating, and then do the single hardest thing a person can do, which is leave the outcome alone.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Leave the outcome alone.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That last part feels like surrender, and it is, and that is exactly why it is so hard. Every part of you wants to grip. Gripping feels like caring. It feels like doing your part. But on this entire class of things, the grip is the surest way to fail, and loosening it is not laziness. It is a skill, and it may be the deepest one a life asks of you. Some effort &lt;a href="https://harryfloyd.substack.com/p/difficulty-was-making-you" rel="noopener noreferrer"&gt;quietly builds you&lt;/a&gt;. This is the other kind, spent on the one thing effort can only spoil.&lt;/p&gt;

&lt;p&gt;It is ten past three somewhere, and you are lying very still, doing arithmetic in the dark. You have already done your part. The room is dark, the day is behind you, there is nothing left to arrange. There is nothing left to do but the one thing you cannot do, which is try. So stop. There was never anything there to try for. You have set the conditions. Now let yourself be no use at all for a while.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is not the usual thing here. I mostly write about AI, product and markets, and the structures underneath them. This is the same habit of looking, pointed at something a lot more ordinary, and I wrote it because I kept meeting it at three in the morning rather than at a desk.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;_New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap.&lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=article&amp;amp;utm_medium=web&amp;amp;utm_campaign=cannot-try-to-fall-asleep" rel="noopener noreferrer"&gt;Subscribe&lt;/a&gt; for the rest, or &lt;a href="https://harryfloyd.substack.com/p/start-here-what-survives-when-the" rel="noopener noreferrer"&gt;start with what survives&lt;/a&gt;.&lt;br&gt;&lt;br&gt;
_&lt;/p&gt;

&lt;p&gt;1&lt;/p&gt;

&lt;p&gt;The phrase belongs to the nineteenth-century philosopher Henry Sidgwick, and John Stuart Mill put it plainly in his autobiography: those are happiest, he wrote, who have their minds fixed on some object other than their own happiness, and who find happiness by the way. It is a very old idea with a long line of owners, which is part of why it is worth trusting.&lt;/p&gt;

&lt;p&gt;2&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>analysis</category>
    </item>
    <item>
      <title>Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:23 +0000</pubDate>
      <link>https://dev.to/harryfloyd/three-hidden-bottlenecks-the-ai-buildout-has-already-moved-past-gpus-18c1</link>
      <guid>https://dev.to/harryfloyd/three-hidden-bottlenecks-the-ai-buildout-has-already-moved-past-gpus-18c1</guid>
      <description>&lt;h1&gt;
  
  
  Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs
&lt;/h1&gt;

&lt;p&gt;Bloom Energy reported Q1 2026 revenue of $751 million. That number was 130 percent higher than the prior year, 42 percent above consensus, and triggered a full-year guidance raise to $3.6 billion 1. Most of the post-earnings coverage read the print as a fuel cell company finally turning operationally profitable.&lt;/p&gt;

&lt;p&gt;The print is not a fuel cell story. It is the canonical evidence that the AI infrastructure bottleneck has migrated past compute.&lt;/p&gt;

&lt;p&gt;For two years the consensus model for AI capex has anchored on GPU shipments. NVIDIA, AMD, the hyperscaler capex disclosures, the analyst models all priced compute as the load-bearing constraint. The reasoning was straightforward: training runs scaled, GPU clusters grew from 5,000 units to 50,000 to 100,000, and the company that supplied the silicon owned the bottleneck.&lt;/p&gt;

&lt;p&gt;The reasoning was correct in 2023. It became incomplete in 2024. By 2026 it has become a rear-view mirror.&lt;/p&gt;

&lt;p&gt;The analyst models that price AI on GPU shipments are not wrong about GPUs being important. They are wrong about GPUs being scarce. The supply-side data has been telling a different story for three quarters now, and Bloom Energy’s print is the most recent confirmation. The bottleneck moved. It always does. &lt;em&gt;The binding constraint never disappears. It only migrates to the next layer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The question that matters now is which layer the binding constraint has migrated to. Three layers have evidence pointing at them, none of which are GPUs, and the layers compose into a single observation about where AI capex goes once the compute layer has been solved.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The first layer: power, and the 128-week wait&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Behind every large GPU cluster sits a power-delivery infrastructure that takes longer to build than the cluster itself. Power transformers, the equipment that steps utility-scale voltage down to data-centre-usable voltage, have 80 to 128 week lead times right now 2. Cleveland-Cliffs is the only domestic US producer of the grain-oriented electrical steel that every transformer core requires 3. The grid interconnection queue at major US utilities runs five-plus years for new high-voltage data centre loads 4.&lt;/p&gt;

&lt;p&gt;This is the layer where Bloom Energy fits, and where the print becomes legible. Solid oxide fuel cells generate power on-site, behind the meter, without queueing for grid interconnection. A hyperscaler that wants 100 megawatts of power in eighteen months and cannot get it from the grid for five years buys Bloom Energy units. The fuel cell technology is twenty years old. The 130 percent revenue growth is the price of how binding the power constraint has become.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A hyperscaler that wants 100 megawatts in eighteen months and cannot get it from the grid for five years buys Bloom Energy units. The 130 percent revenue growth is the price of how binding the power constraint has become.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For beginners: what does “behind the meter” mean?&lt;/strong&gt; A utility meter measures power coming into a building from the grid. &lt;em&gt;Behind the meter&lt;/em&gt; means power generated on the customer’s side of that meter, so the grid never sees it and never has to plan for it. Bloom Energy’s fuel cells are behind-the-meter generation. That is why the eighteen-month installation timeline is the only one that matters for a hyperscaler who cannot wait five years for grid interconnection.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The falsifier for the power layer is specific. If transformer lead times compress below 52 weeks within two consecutive quarters, or if hyperscaler 24/7 firm clean power purchase agreements (PPAs) at 15-year tenors are consistently signed below $80 per megawatt-hour, the constraint has eased and the behind-the-meter premium decays. Watch the second of those harder than the first. Hyperscalers will pay whatever the grid cannot deliver fast enough, and the PPA price is where that desperation gets numerical.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The second layer: metal, and the recycling angle nobody priced&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The compute layer requires copper. The power layer requires copper. The interconnect layer requires copper. By 2030, AI data centres alone will be calling on roughly 7 percent of all the copper the world digs up in a year, from a demand source that did not meaningfully exist five years ago. The math is straightforward. Hyperscale AI sites consume 40 to 50 tons of copper for every megawatt of IT capacity 5. The US has 85 gigawatts of new pipeline through 2030 6, with 35 gigawatts already under construction across North America 7. Wood Mackenzie projects 1.1 million tonnes per year of grid copper demand from data centres alone 8; BloombergNEF projects another 572,000 tonnes peaking in 2028 inside the facilities themselves 9. Combined, that approaches 1.7 million tonnes per year against global mine output of roughly 23 million tonnes annually 10. One new demand source, 7 percent of every mine on earth, on top of every other demand the market already cannot meet.&lt;/p&gt;

&lt;p&gt;Mine capacity does not flex on the timescales the buildout requires. Copper mines take a decade from greenfield discovery to first commercial shipment. The buildout is happening on a one-to-three year horizon. There is no path where new mining capacity meets new data-centre demand.&lt;/p&gt;

&lt;p&gt;The consensus copper-AI thesis names the major miners: Freeport-McMoRan, Southern Copper, BHP, Rio Tinto. The miners are the obvious read. The recycling angle is the underfollowed one. Aurubis is a German specialty metals conglomerate that runs the largest secondary copper smelting capacity in Europe and is building the first US secondary smelter. Recycling output can flex on the timescales primary mining cannot. The structural shift is from &lt;em&gt;mining is the bottleneck&lt;/em&gt; to &lt;em&gt;recycling is the relief valve&lt;/em&gt;. The equity that captures the relief valve trades at approximately 0.4 times price-to-sales 11. The market reads Aurubis as a commodity cyclical. The multiple ignores the data-centre demand curve.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The structural shift is from “mining is the bottleneck” to “recycling is the relief valve”. The equity that captures the relief valve trades at roughly 0.4 times price-to-sales. The multiple ignores the data-centre demand curve.&lt;/p&gt;

&lt;p&gt;This publication tracks where capital is migrating before the analyst models reprice it. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?" rel="noopener noreferrer"&gt;Subscribe now&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The falsifier for the metal layer is observable and time-bound. If primary copper-mine output growth exceeds 10 percent year-over-year for two consecutive years, the supply-shortage premium for recyclers compresses. The fallback test: if hyperscaler-driven data-centre permitting decelerates by more than 30 percent year-over-year, the demand assumption breaks before the supply assumption fires. Watch the permitting numbers monthly. The construction pipeline is the leading indicator of the copper demand curve.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The third layer: detection, and the $151 billion question&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The third layer is the most speculative of the three, and also the one where the supply side is voting hardest. The reason markets have not priced it yet is that the contract that creates it was only finalised in January 2026. SHIELD is the Scalable Homeland Innovative Enterprise Layered Defense vehicle: a $151 billion ten-year contract the Missile Defense Agency awarded as the primary acquisition framework for the broader Golden Dome missile-defence initiative 12. Golden Dome itself sits above SHIELD as the umbrella programme, with the Pentagon’s own ten-year cost estimate at approximately $185 billion and the Congressional Budget Office’s May 2026 analysis projecting up to $1.2 trillion over twenty years if a full space-based interceptor layer is built out 13. The MDA selected 2,440 firms as qualified SHIELD vendors across three tranches in late 2025 and early 2026. Holding a SHIELD position confers eligibility to compete for individual task orders, not guaranteed funding; task-order competitions are now beginning.&lt;/p&gt;

&lt;p&gt;The data layer of Golden Dome (the satellites and ground-segment processing that detect, classify, and track aerial threats) is a procurement category that did not meaningfully exist five years ago. Spire Global is a publicly-traded satellite-data company at roughly $700 million market cap 14 with a remaining-performance-obligations backlog above $200 million, equivalent to about three times trailing twelve-month revenue 15. Their core revenue stream is Global Navigation Satellite System (GNSS) radio-occultation weather data, maritime Automatic Identification System (AIS) tracking, and radio frequency (RF) signal monitoring. Each of those data feeds is dual-use. The same instruments serve weather forecasting, shipping logistics, and defence persistent surveillance.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>architecture</category>
      <category>investing</category>
    </item>
    <item>
      <title>The Substrate Map</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:08 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-substrate-map-3hfh</link>
      <guid>https://dev.to/harryfloyd/the-substrate-map-3hfh</guid>
      <description>&lt;h1&gt;
  
  
  The Substrate Map
&lt;/h1&gt;

&lt;p&gt;Most teams can tell you what they shipped.&lt;/p&gt;

&lt;p&gt;Fewer can tell you what will still matter after the next large change.&lt;/p&gt;

&lt;p&gt;That is the gap the Substrate Map is built for.&lt;/p&gt;

&lt;p&gt;It is a free one-page taxonomy for separating canopy from substrate in your own work. The canopy is the visible layer: prompts, model choices, demo polish, current benchmark scores, frameworks, UI surfaces, launch artefacts. The substrate is the part that keeps doing work when the surface gets repriced: data-quality discipline, eval contracts, workflow integration, trust packaging, domain-specific failure memory, and the proprietary signal the next model release does not have.&lt;/p&gt;

&lt;p&gt;The tool gives you a 10-minute exercise:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Open the last 90 days of engineering tickets, product launches, roadmap decisions, or investment decisions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tag each item as substrate or canopy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Compute the ratio.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The number is blunt on purpose.&lt;/p&gt;

&lt;p&gt;If the last 90 days were mostly canopy, the next release can reset most of what you built. If the split is 50/50, you are probably normal but not especially durable. If the work is mostly substrate, protect it. That is the work compounding underneath the visible output.&lt;/p&gt;

&lt;p&gt;The most useful part is the boundary rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If your team cannot agree which column an item belongs in, tag it as canopy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Disagreement at the boundary means the substrate work has not been made explicit yet.&lt;/p&gt;

&lt;p&gt;That makes the map useful before a planning meeting. Instead of arguing about whether a roadmap “feels strategic”, you can ask which work would still matter if the model, market, channel, or buyer changed. The conversation gets harder to fake because each item has to be placed in a column.&lt;/p&gt;

&lt;p&gt;The PDF is deliberately simple: one map, one exercise, one ratio. It is not a strategy deck. It is the first instrument you run when the team is shipping a lot but cannot say what is compounding.&lt;/p&gt;

&lt;p&gt;Use it on a roadmap. Use it on a product backlog. Use it on a portfolio. Use it before a planning cycle where everyone is about to argue from vibes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Download.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Substrate Map&lt;/p&gt;

&lt;p&gt;105KB ∙ PDF file&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/api/v1/file/5b324df8-e43f-486f-823b-7213b88910b4.pdf" rel="noopener noreferrer"&gt;Download&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/api/v1/file/5b324df8-e43f-486f-823b-7213b88910b4.pdf" rel="noopener noreferrer"&gt;Download&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;More free tools like this.&lt;/strong&gt; &lt;em&gt;Subscribe to get the next durability-lens resource the day it ships.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Substrate Map is a companion to &lt;em&gt;&lt;a href="https://harryfloyd.substack.com/p/the-forest-floor-is-the-product?utm_source=resource-landing&amp;amp;utm_medium=internal&amp;amp;utm_campaign=substrate-map-2026-05-02" rel="noopener noreferrer"&gt;The Forest Floor Is the Product&lt;/a&gt;&lt;/em&gt; , the essay that develops the substrate-vs-canopy lens across ecosystems, software, knowledge work, and capital allocation.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What percentage of your last 90 days was substrate?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
