<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harry Floyd</title>
    <description>The latest articles on DEV Community by Harry Floyd (@harryfloyd).</description>
    <link>https://dev.to/harryfloyd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3933548%2Fa644e757-fdc0-4213-a2d0-37774cbe6730.png</url>
      <title>DEV Community: Harry Floyd</title>
      <link>https://dev.to/harryfloyd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harryfloyd"/>
    <language>en</language>
    <item>
      <title>The Work That Comes Due After You Leave</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 26 Aug 2026 19:07:17 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-work-that-comes-due-after-you-leave-5blb</link>
      <guid>https://dev.to/harryfloyd/the-work-that-comes-due-after-you-leave-5blb</guid>
      <description>&lt;h1&gt;
  
  
  The Work That Comes Due After You Leave
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21EAg5%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F5b778f5b-782f-4aff-b917-9d38942e86bb_1600x900.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21EAg5%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F5b778f5b-782f-4aff-b917-9d38942e86bb_1600x900.webp" alt="The read-across: your checklist on the left, a record you did not write on the right."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You finish something. A project wraps, a client signs off, a piece of work goes out for the last time. Then there is the tail, the small handful of things you do afterwards, none of which take any real time. Mark it done. Tell them it is finished. Cancel the paid seat you bought for it. Switch off the weekly update that goes out to them every Monday.&lt;/p&gt;

&lt;p&gt;Four steps, four different places: the tracker, your email, wherever the card gets charged, whatever tool sends that update. Later you check one of them, probably the tracker, because that is where you look to see whether things are finished. It tells you the job is done, and it is telling the truth about the only step it can see.&lt;/p&gt;

&lt;p&gt;Switching off the update is the one that did not happen. Months later it is still arriving, every Monday at nine, to someone who stopped being your client a long time ago. Nothing is wrong with the system that sends it. It is doing exactly what it was told, on time. From where you are standing, the failure looks exactly like everything working, and that is the whole of the problem.&lt;/p&gt;

&lt;p&gt;The tracker is not lying. A job that touches four systems has four different ways of still being open, and the tracker sees only the one it holds. Done was never one state; a single word just made it look like one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the small steps are the ones that go missing
&lt;/h2&gt;

&lt;p&gt;The tempting explanation is that you were busy, or careless, or need a better checklist. I want to offer a more specific one, because it tells you which steps will go wrong instead of telling you to try harder.&lt;/p&gt;

&lt;p&gt;A checklist is a list of the steps you thought of. It is good at holding you to those. What it cannot do is mention a step that never went on it, and the steps that never go on it are not random. They are the ones that cross into a system you do not quite think of as part of the job. You wrote the list around the place you do the work, and the step that lives somewhere else did not occur to you, for the same reason it will not later occur to you to check whether it happened.&lt;/p&gt;

&lt;p&gt;Making a second list does not save you, and that is the part worth sitting with. If you build the second list from the same picture of the job, the same step is missing from it too, and now you have two records that agree with each other and are both wrong. That is not a hypothetical: your tracker is that second list. You filled it from the same picture of the job, so it agreed the work was done and was wrong in the same place you were.&lt;/p&gt;

&lt;p&gt;And a missed closing step does not stay missed quietly. An ordinary task you skip just sits there until you come back to it; a closing step you skip stays open until something closes it, and until then it keeps acting, every day or every month, on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check has to come from somewhere you did not write
&lt;/h2&gt;

&lt;p&gt;So the thing that catches the missing step cannot be your own account of the work. It has to be a record that something else kept, for its own reasons, whether or not you remembered the step.&lt;/p&gt;

&lt;p&gt;You already have several of these. You just do not read them against the job. The card statement is one: the bank records the charge whether or not you remember the seat you meant to cancel, so the seat that is still billing turns up as a line you cannot attach to any live piece of work. The access list is another: the system logs who can get in whether or not anyone told it that a person left, so the account that outlived the project is a login with no current owner. What actually shipped is recorded by the thing that shipped it, so a promise you made and never delivered stands as a commitment on one side with no send on the other.&lt;/p&gt;

&lt;p&gt;Even with a checklist I take seriously, I did this. I keep a written routine for finishing an essay, detailed, with a warning next to the item that slips most, and my archive quietly slipped twenty-three pieces behind what I had published since late spring. Copying each finished piece across to that archive had never been a line on the routine at all: it lived on a different system, so it never occurred to me to write it down. What caught it was the published record of what had actually gone out, kept by the platform and owing nothing to my memory. Held against the archive, it showed the twenty-three at once.&lt;/p&gt;

&lt;p&gt;The move itself is old. Accountants have reconciled two sets of books this way for centuries, and there is nothing here to invent. What is easy to get wrong is what makes the second record worth anything: not that it is a second record, but that something other than your own memory produced it. Two dashboards drawn from the same database, or two lists built from the same picture of the job, only look like a check, because the same forgetting shaped both. A record can catch you only when your forgetting could not have reached it too.&lt;/p&gt;

&lt;h2&gt;
  
  
  One question to carry
&lt;/h2&gt;

&lt;p&gt;That gives you a single question, and it is worth more than any checklist. Of anything you lean on to tell you a job is finished, ask: would this still be here, and still say the same thing, if I had forgotten the step entirely?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21GYNz%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F4ef861e0-b57e-4c28-9701-3290dc4144f2_1600x900.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21GYNz%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F4ef861e0-b57e-4c28-9701-3290dc4144f2_1600x900.webp" alt="The test, applied to two records."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The test, applied to two records. The one you fill in yourself fails it; the one the bank writes passes it, because the charge is there whether or not you remembered.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The tracker fails that question: you fill it in yourself, so a step you forget is a step you also forget to log, and it stays green over a gap it never knew about. The card statement passes, because the charge is there whether or not you remembered the seat. A check built from your own memory cannot expose the step that memory left out.&lt;/p&gt;

&lt;p&gt;The question keeps its shape as the instrument gets bigger. A tracker, a dashboard, a report you write on your own project: each is an instrument you fill from your own picture of the work, and each is blind in the same place you are. The statement is worth more than any account you write of what you meant to do, for the same reason an audit leans hardest on evidence the audited side did not get to shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part you can use
&lt;/h2&gt;

&lt;p&gt;Here is the version you can run on your own job this week.&lt;/p&gt;

&lt;p&gt;Write out the routine you run after you finish something, every step, including the ones that feel too small to be worth writing down. Mark each with the system it touches: the tracker, the calendar, the billing account, the shared drive, the tool somebody set up before you arrived. This is not the check yet. It is how you find which records are worth reading against each other, and it usually turns what felt like one job into the three or four systems it was always made of.&lt;/p&gt;

&lt;p&gt;Then there are two ways to keep a step from being lost, and the first is much stronger. Where you can, do not rely on catching the step at all; arrange things so that forgetting it does no harm. Anything that runs on its own, a payment, a subscription, a recurring invite, an access granted for a single project, gets its end date on the day you set it up, while you still know what it was for. Something that expires unless it is renewed cannot outlast your forgetting, because forgetting it and ending it become the same act. Reach for this first; it removes the obligation instead of watching it. Its limit is the one this piece began with: you can only set an end date on a step you thought of, and the step that never made the list cannot be made self-closing.&lt;/p&gt;

&lt;p&gt;For everything you could not foresee, or cannot make expire, there is the slower move: read your own record against one you did not produce. Your active-projects list against the vendor or card statement, looking for a charge attached to work that has already finished. Your list of who is on the team against the access export from whatever holds the accounts, looking for a login with no owner. The commitments in a signed contract against what your team actually sent, looking for a promise with no matching send. Choose the second record by the causal test, not by where it happens to be stored: pick the one your own memory did not shape.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21G3so%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Faefe3f68-c79d-405c-b548-2dce81450723_1600x900.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21G3so%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Faefe3f68-c79d-405c-b548-2dce81450723_1600x900.webp" alt="Three records you keep, each read against one you did not write."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Three records you keep, each read against one you did not write. The last row is the limit: recorded nowhere, so nothing catches it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;How often you look depends on how much damage you will let build up first. A ten-pound seat can wait a month; a former colleague who can still open every file cannot, and something confidential still reaching the wrong person is not a scheduled job at all. None of it needs a tool you have to build: a read-only export or a screenshot is enough, and where you cannot pull the record yourself, the person who can is an email away, not a project. And reading across only points to a mismatch; you still have to look and decide whether it is a real miss, a timing lag or a duplicate, and keep that verdict for yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What none of this fixes
&lt;/h2&gt;

&lt;p&gt;Two things survive all of it, and I would rather say so than leave the tidy version standing.&lt;/p&gt;

&lt;p&gt;The first is the record that was never kept. If something gets finished and lands in no system at all, no charge, no log, no row anywhere, then there is no second record to read it against. You cannot check against a record that does not exist. That case surfaces only when a person happens to notice, or is told.&lt;/p&gt;

&lt;p&gt;The second is quieter, and more common. If the same blind spot sits in both records, they agree, and the agreement looks like an all-clear. This is the failure I walked into the first time I tried to build a check like this for myself. I searched my files for links to the publication. That sounds like reading an independent record, until you notice it read the same surface I would have: it counted the times I had linked to old pieces inside new ones as though that proved the old ones had shipped. Independence is the whole of the mechanism, and when it is missing it fails without a sound.&lt;/p&gt;

&lt;p&gt;So the honest tally is smaller than the tidy one. The obligations that leave a trace in a record I did not write, I can now catch, once in a while, in half an hour. The ones that touch nothing outside my own attention, I am still carrying in my head, and I have learned how little the word covers when I say nothing is wrong. &lt;em&gt;Nothing is wrong&lt;/em&gt; and &lt;em&gt;nothing I can see is wrong&lt;/em&gt; are different sentences, and most of the time only one of them is available to any of us.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=work-that-comes-due-after-you-leave" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21WT9l%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Fc51d95fe-c724-4e35-80d1-b78738adfec8_1600x560.webp" alt="The Day Job banner"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>analysis</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Memory Capacity Binds Before FLOPs Do: AgentX and the Agentic Inference Bottleneck</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Tue, 25 Aug 2026 20:09:07 +0000</pubDate>
      <link>https://dev.to/harryfloyd/memory-capacity-binds-before-flops-do-agentx-and-the-agentic-inference-bottleneck-20i8</link>
      <guid>https://dev.to/harryfloyd/memory-capacity-binds-before-flops-do-agentx-and-the-agentic-inference-bottleneck-20i8</guid>
      <description>&lt;h1&gt;
  
  
  Memory Capacity Binds Before FLOPs Do
&lt;/h1&gt;

&lt;p&gt;In agentic inference, memory capacity binds before FLOPs do. That is the first finding from AgentX, the open-source benchmark SemiAnalysis built for replaying real agentic coding traffic at one million context.&lt;/p&gt;

&lt;p&gt;The numbers on DeepSeek V4 make the point. The HBM working set decided the KV-cache hit rate: 43 million tokens and 91 percent on a B300, against 22 million and 73 percent on a B200. The B300 run used 384 concurrent traces, the B200 run 196. Same model, same task, different memory budget, and the hit rate moved by eighteen points.&lt;/p&gt;

&lt;p&gt;Why the hit rate matters: agentic sessions reuse prefixes. Every turn builds on the context before it, so most of the context can be served from the KV cache rather than recomputed. A miss means re-prefilling context you have already paid for once. In an agent loop of dozens of sequential calls, those misses compound into wall-clock time and compute you cannot get back.&lt;/p&gt;

&lt;p&gt;AgentX matters because it measures the right thing. It replays 393 anonymised internal Claude Code traces, structure and timing preserved, under Apache 2.0. Fixed-sequence benchmarks reflect chip and kernel performance; agentic workloads reflect the systems problem: KV tensors, routing, and offload across memory tiers. The first open-source benchmark of this shape has produced the finding the labs have been converging on: the binding constraint in agentic inference is not peak FLOPs, it is the size of the memory working set.&lt;/p&gt;

&lt;p&gt;The practical rule for anyone sizing an inference stack: size working set before FLOPs. A spec sheet that leads with peak throughput hides the number that actually determines your agent's latency. Measure step latency at your own concurrency, including prefill, before you trust a batch-of-one headline.&lt;/p&gt;

&lt;p&gt;The pattern is durable. When compute gets cheap enough, the constraint migrates to the layer beneath it, the same migration the industry has watched in every previous scaling phase. AgentX is the first open benchmark to show where it has landed in agentic inference: not in the silicon that generates tokens, but in the memory that holds the conversation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>infrastructure</category>
      <category>analysis</category>
    </item>
    <item>
      <title>You're Not Comparing Models. You're Comparing Contracts.</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:42:42 +0000</pubDate>
      <link>https://dev.to/harryfloyd/youre-not-comparing-models-youre-comparing-contracts-647</link>
      <guid>https://dev.to/harryfloyd/youre-not-comparing-models-youre-comparing-contracts-647</guid>
      <description>&lt;h1&gt;
  
  
  You're Not Comparing Models. You're Comparing Contracts.
&lt;/h1&gt;

&lt;p&gt;Two teams publish scores on the same agent benchmark.&lt;br&gt;&lt;br&gt;
One lands in the low sixties. The other clears seventy.&lt;br&gt;&lt;br&gt;
A procurement team reads the spread and makes a call.&lt;/p&gt;

&lt;p&gt;What they do not see: both teams may be running the same model. They did not need to change the weights for the gap to appear. The spread can come from scaffold alone.&lt;/p&gt;

&lt;p&gt;One team wrapped the model in a harness with better retries. Different tool defaults. A planner step the other team had skipped. None of that appears on the leaderboard.&lt;/p&gt;

&lt;p&gt;The comparison that drove the decision was not between two agents.&lt;/p&gt;

&lt;p&gt;It was between two contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;There Is No Benchmark&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The mistake hiding behind this story is a category error.&lt;/p&gt;

&lt;p&gt;People talk about agent benchmarks as if they measure a thing called “the model.” They do not. They measure a coupled system. The model is one component. The rest is a stack of protocol decisions that are almost never disclosed and almost always matter.&lt;/p&gt;

&lt;p&gt;The score is the output of that stack. Change any layer and you change what the number means.&lt;/p&gt;

&lt;p&gt;Recent research on agent evaluation has named those layers explicitly. There are at least seven. Deployment regime. Observation channel. Harness and scaffold. Metric and action. Configured evaluator. Grader protocol. Audit bundle. Each is a contract. Each is negotiable. And each can silently change the verdict while the headline looks the same.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;That is what a benchmark actually is. Not a measurement of a model. A measurement of an entire testing contract, of which the model is one slot.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is structural reason the seven layers are the seven layers. They cluster into three corners that show up in almost every published agent-evaluation failure. What the model is rewarded for. How that reward is optimised. And how the test contract differs from production. Once you hold those three corners in view, the seven-layer stack stops feeling like a checklist and starts behaving like the actual shape of what is being measured.&lt;/p&gt;

&lt;p&gt;If you are comparing agent products without parity across those layers, you are not comparing agents.&lt;/p&gt;

&lt;p&gt;You are comparing contracts and calling it science.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Harness You Didn’t Name&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The most visible layer, and the one that moves the most points, is the scaffold.&lt;/p&gt;

&lt;p&gt;Anyone who has built an agent in the last year has felt this without naming it. You watch a coworker get 75% on a task your model just failed on. You check the weights. They are yours. They changed the prompt template and added a retry loop. The model did not get smarter. The scaffold got thicker.&lt;/p&gt;

&lt;p&gt;The numbers say the same thing. RWE-bench reports that on its 162-task benchmark over MIMIC-IV, changing only the agent scaffold around a fixed model can shift performance by more than 30 percent1. Same weights. Different tools. Different retry policy. Different planner. Different headline. The best evaluated agent on that benchmark reaches around 40 percent task success at all; the best open-source configuration is closer to 30. Once you know the contract can move 30 points on its own, neither of those numbers is really about a model.&lt;/p&gt;

&lt;p&gt;If scaffold alone can move scores by double digits, then “we used the same model as them” is not a fair-comparison claim. It is a parameter-naming claim. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You have named one slot in a seven-slot contract.&lt;br&gt;&lt;br&gt;
The other six are doing most of the work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Judge That Isn’t The Model&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The second layer that silently moves scores is the evaluator itself.&lt;/p&gt;

&lt;p&gt;When a benchmark uses an LLM judge, people write things like “graded by GPT-4o” as if that pins the measurement down. It does not. The judge is not GPT-4o. The judge is GPT-4o plus a prompt template. Plus a decoding configuration. Plus a tie and abstention policy. Plus whatever retrieval or tool access the judge has during grading. None of that ships with the score.&lt;/p&gt;

&lt;p&gt;A recent systematic evaluation of LLM-as-judge setups showed that prompt-template choice alone materially changes both judge quality and internal consistency2. Two teams reporting “we used GPT-4o as judge” can be running substantively different graders. The grader that rewards epistemic hedging disagrees with the grader that penalises it. The grader with access to retrieval checks factuality. The grader without one does not, and cannot.&lt;/p&gt;

&lt;p&gt;This is not a small print issue. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The evaluator is the measuring instrument. If two teams use different instruments and report the same number, they are not reporting the same thing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And without a published judge card, no third party can reproduce the measurement. They can only rerun the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Number That Lies About Consistency&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The third layer is the quietest and most dangerous. It is the metric itself.&lt;/p&gt;

&lt;p&gt;A standard agent metric is &lt;a href="mailto:pass@k"&gt;pass@k&lt;/a&gt;. You give the agent k attempts. If any one succeeds, it counts. This is perfectly reasonable if your production use allows k attempts. It is actively misleading if it does not.&lt;/p&gt;

&lt;p&gt;There is a sibling metric, pass^k. Same k attempts. But it only counts if the agent succeeds on all of them. It measures consistency, not capability.&lt;/p&gt;

&lt;p&gt;The gap between these two can be large, and it can open silently.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Recent work on trustworthy agent evaluation shows that controlled error injection into an agent can cut pass^k substantially while barely moving &lt;a href="mailto:pass@k3"&gt;pass@k3&lt;/a&gt;. The model still has a ceiling you can hit with enough tries. It has lost the ability to hit that ceiling reliably. If your headline is pass@k and your production regime is one shot, the leaderboard says you are shipping. The bug tracker says otherwise.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The same structural problem appears in calibration metrics. ECE asks whether stated probabilities match empirical frequencies on average. AURC asks whether the system can rank harder cases lower. Both can look nearly identical across two systems while a stricter, abstention-aware metric called BAS, the Behavioural Alignment Score, diverges sharply between them4. BAS asks a different question. Does the confidence surface protect you in exactly the regime where a person or product would actually choose to trust it? Two systems with “similar calibration” can answer that question completely differently once you attach a cost function.&lt;/p&gt;

&lt;p&gt;The metric is not a measurement of the model. It is a statement about which errors the model’s operators will tolerate. If that statement does not match your operational contract, the score is not wrong. It is answering a question you did not ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Why Rank Stability Is A Trap&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Here is the part that makes all of this subtly worse.&lt;/p&gt;

&lt;p&gt;Under scaffold shift, the rank order of agents on a benchmark is often relatively stable. A recent efficient-benchmarking study reports that rank preservation is easier to maintain than absolute calibration5. The number moves. The ordering does not.&lt;/p&gt;

&lt;p&gt;If all you need is a relative decision, rank stability is comforting. Agent A beats Agent B here, and probably beats it in production.&lt;/p&gt;

&lt;p&gt;If you need an absolute decision, it is a trap.&lt;/p&gt;

&lt;p&gt;Procurement, safety arguments, SLA setting, cost modelling, and risk disclosure all depend on absolute numbers. A claim like “this agent ships 80% correct at 5 cents per request” binds to the calibrated level, not to the rank. Under scaffold shift, rank can hold while the 80% becomes 62%. Your spreadsheet is still using 80%. Your customers are experiencing 62%.&lt;/p&gt;

&lt;p&gt;The protocol that produced 80% is part of the claim. The moment it diverges from production, the claim silently becomes false, even though nothing about the model moved.&lt;/p&gt;

&lt;p&gt;This is why seasoned eval teams treat the contract, not the score, as the primary artefact. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You can rerun a score. You can only reproduce a contract if you wrote it down.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Contract Is The Object&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;If the score is a function of the contract, the practical move is to treat the contract as the thing you own.&lt;/p&gt;

&lt;p&gt;That means three changes to how most teams currently work.&lt;/p&gt;

&lt;p&gt;Freeze the contract before you compare. If you cannot describe your deployment regime, observation channel, scaffold version, metric, judge configuration, and grader protocol in one page, you do not have a contract. You have assumptions pretending to be one. Write the page. Commit it. Make it a prerequisite for every comparison.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Your Tools Got Powerful. Get Boring.</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:42:26 +0000</pubDate>
      <link>https://dev.to/harryfloyd/your-tools-got-powerful-get-boring-48jn</link>
      <guid>https://dev.to/harryfloyd/your-tools-got-powerful-get-boring-48jn</guid>
      <description>&lt;h1&gt;
  
  
  Your Tools Got Powerful. Get Boring.
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?" rel="noopener noreferrer"&gt; Subscribe now&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;The bored trader beats the machine&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;On one side of the trade sits a market-making engine that represents the genuine state of the art: Hawkes processes modelling order arrivals, Kyle’s lambda pricing the impact of each fill, Avellaneda-Stoikov inventory control balancing the book in real time. Years of mathematics, running on hardware that did not exist a decade ago.&lt;/p&gt;

&lt;p&gt;On the other side is a momentum trader whose entire system is price, volume, and three moving averages. He sits in cash most of the year doing nothing, waiting for a setup he could describe to you in a sentence. His stack is deliberately primitive. His edge is patience and the discipline to follow his own rules when they are boring and to sit out when they are silent.&lt;/p&gt;

&lt;p&gt;Over a full market cycle, the boring one is more likely to still be standing.&lt;/p&gt;

&lt;p&gt;This is uncomfortable, because it runs against an intuition almost everyone shares: better tools should let you run better, more sophisticated strategies. More compute, more data, more powerful models, therefore more elaborate approaches and better results. It feels obviously true. It is the logic behind most of what gets built, bought, and bragged about.&lt;/p&gt;

&lt;p&gt;It is also, across domain after domain, wrong. And the interesting part is the shape of the curve.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;The gap widens as the tools get stronger&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Here is the pattern the most successful practitioners keep seeing, whether they are trading, building software, learning, or shipping products. Powerful tools do not pay off when you point them at more complex strategies. They pay off when you point them at simple strategies and execute those faster, more consistently, and with less drift than anyone else.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;More power applied to a simple strategy compounds. The same power applied to a complex one mostly buys you more ways to be wrong.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sit with the second half of that, because it is the part people miss. A sophisticated strategy is not free. Every additional layer needs to be specified, verified, maintained, and monitored, and all of that consumes exactly the capacity the powerful tool was supposed to give back. A simple strategy spends its new power on doing the simple thing relentlessly well. A complex one spends its new power feeding its own machinery.&lt;/p&gt;

&lt;p&gt;The reason this matters more now than it ever has is that the tools have never been this strong. When your instruments are weak, the gap between the simple-and-disciplined path and the complex-and-fragile path is small, because nobody can do much of either. As the instruments get more powerful, both paths open up, and the distance between them widens. The most capable tools in history make disciplined simplicity more effective than ever, and they also make unmanageable complexity easier to build than ever. We are living through the largest gap between those two paths that has ever existed, and most people are sprinting down the wrong one with a faster engine.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What complexity quietly costs&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The bill for sophistication does not arrive when you build it. It arrives later, in instalments, and it is always larger than it looked.&lt;/p&gt;

&lt;p&gt;The first instalment is verification. A simple system you can hold in your head and check. A complex one you cannot, so you build monitoring to watch it, and the monitoring becomes its own system that can &lt;a href="https://harryfloyd.substack.com/p/most-verification-is-just-bigger" rel="noopener noreferrer"&gt;drift and mislead&lt;/a&gt;. Every layer you add is a layer you now have to confirm is still doing what you think it does, and the confirming never ends.&lt;/p&gt;

&lt;p&gt;The second instalment is the day it breaks. A simple strategy fails legibly: you can see which rule was wrong and fix it. A sophisticated one fails in the seams between its parts, at the worst possible moment, in a way no single person fully understands. The elaborate model that printed money for two years becomes, in the drawdown, a black box nobody can debug while it is bleeding. Complexity does not only add capability. It adds failure modes that stay hidden until the system is under stress, which is the exact moment you have no spare capacity to handle them.&lt;/p&gt;

&lt;p&gt;The deepest cost is fragility to your own success. A strategy with many parameters has many surfaces the world can destabilise once it starts reacting to you. The more elaborate the machine, the more places reality can reach in and pull a lever you forgot you had wired up. Simple, constrained systems survive contact with the world because there is less of them to break.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Sophistication is a loan against your future attention, taken out at a rate you cannot see until the system is under stress and the whole balance comes due at once.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?" rel="noopener noreferrer"&gt; Subscribe now&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Why we reach for sophistication anyway&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;If simplicity wins, why does almost everyone instinctively add complexity? Smart, capable people do it constantly, because the incentives reward it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every incentive in the room rewards the complexity you can show and punishes the discipline you cannot.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sophistication is visible. A complex model, an elaborate architecture, a clever framework can be shown to a boss, a client, an investor, a peer. Discipline cannot be shown. Sitting in cash for three months, deleting half your code, pausing before you speak, refusing to ship the extra feature: none of it photographs well. The market pays for what it can see, and it can see complexity far more easily than it can see restraint.&lt;/p&gt;

&lt;p&gt;Complexity also feels like work. Building an intricate system produces the sensation of progress all day long, even when the effort is going into &lt;a href="https://harryfloyd.substack.com/p/the-leverage-hierarchy-of-agent-engineering" rel="noopener noreferrer"&gt;the layer with the least leverage&lt;/a&gt;. Doing the boring, correct thing and then waiting produces the sensation of doing nothing, which the nervous system reads as failure. The feeling and the result point in opposite directions, and the feeling usually wins.&lt;/p&gt;

&lt;p&gt;And an entire economy is built on convincing you the work is harder than it is. Every tool vendor, every course, every consultancy has a structural interest in making its domain look more complex than it needs to be, because simplicity is terrible for business. The people who write about a field emphasise its hardest parts, which is what makes them experts, rather than its simplest parts, which is what produces the results. The perceived difficulty of almost everything is inflated, and the inflation is nobody’s accident.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What the constrained version keeps proving&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The clearest place to watch this play out right now is in how people use AI, because the tool is so powerful that the trap is stark.&lt;/p&gt;

&lt;p&gt;The most effective way to get good work out of a frontier model is to take capability away from it. The prompts that consistently produce strong code are the ones that forbid things: no verbose comments, no scattered logging, small functions only, review your own output before returning it. The best debugging prompts are the most constrained ones: strict ordered steps, and a hard rule to verify before changing anything. The most powerful model on the planet does better work when you give it fewer options. People reach for AI expecting more power to mean more freedom. What it rewards is more power inside tighter constraints.&lt;/p&gt;

&lt;p&gt;The same shape shows up wherever someone is quietly winning with powerful tools. The builders who ship profitable products solo run on deliberately boring technology, the kind a fashionable engineer would be embarrassed by. They ship ugly first versions fast while better-resourced teams are still choosing a framework. The plain name for what those teams are doing is over-engineering, and the powerful tools make it easier than ever. The people who learn fastest take fewer notes, not more. They delay and compress until a page of dense understanding replaces a folder of neat transcription. The creators who grow post less, because the algorithm rewards depth per post and punishes the volume that easy tools make tempting. Different fields, one lesson: the powerful tool is best spent removing steps.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The people quietly winning with the strongest tools are using them to do less, and to do it more reliably than anyone else.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;None of these people are anti-technology. They are using the most powerful tools available. They are simply pointing them at the boring fundamentals and refusing the upgrade to a more complicated game.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>machinelearning</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Your AI Agent Stack Is Solving The Wrong Problem</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:42:10 +0000</pubDate>
      <link>https://dev.to/harryfloyd/your-ai-agent-stack-is-solving-the-wrong-problem-2ni5</link>
      <guid>https://dev.to/harryfloyd/your-ai-agent-stack-is-solving-the-wrong-problem-2ni5</guid>
      <description>&lt;h1&gt;
  
  
  Your AI Agent Stack Is Solving The Wrong Problem
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The setup everyone is sharing&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Which MCP servers to install. Which skills to keep in your repo. Which agent framework to use. How to write your &lt;code&gt;AGENTS.md&lt;/code&gt;. How to split one agent into researcher, planner, coder, and reviewer. How to wire Slack, GitHub, Notion, Postgres, Stripe, your calendar, and your file system into one increasingly capable loop.&lt;/p&gt;

&lt;p&gt;Some of that advice is useful. It is also aimed at the wrong layer.&lt;/p&gt;

&lt;p&gt;What becomes real after the agent uses a tool matters more than whether it can reach the tool.&lt;/p&gt;

&lt;p&gt;Can it read the customer record, or change it? Can it draft the refund, or issue it? Can it open a pull request, or merge it? Can it propose the vendor response, or send it under the company name?&lt;/p&gt;

&lt;p&gt;Once an agent can act through tools, the real system is no longer the model.&lt;/p&gt;

&lt;p&gt;The real system is the contract stack around the model.&lt;/p&gt;

&lt;p&gt;That is the part most setup guides skip.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Access is reach. Agency is permissioned action.&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine the demo.&lt;/p&gt;

&lt;p&gt;The agent can read Slack. It can search email. It can query the CRM. It can open GitHub issues, check billing records, browse docs, edit a spreadsheet, draft a customer reply, and call three internal APIs.&lt;/p&gt;

&lt;p&gt;Everyone in the room calls it powerful.&lt;/p&gt;

&lt;p&gt;That is the first mistake.&lt;/p&gt;

&lt;p&gt;The agent has reach. It does not yet have governed agency.&lt;/p&gt;

&lt;p&gt;Access tells you what the agent can touch. Agency tells you what the agent is authorised to decide, under which conditions, with what proof, and with what consequence after failure.&lt;/p&gt;

&lt;p&gt;That distinction sounds small until the first bad run.&lt;/p&gt;

&lt;p&gt;A read-only research assistant can waste time. An agent with billing access can create obligations. An agent with email access can speak for the company. An agent with deployment access can turn a wrong inference into infrastructure.&lt;/p&gt;

&lt;p&gt;More tools do not automatically make the agent more agentic.&lt;/p&gt;

&lt;p&gt;More tools expand the surface on which judgement has to be engineered.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The tool stack is visible. The contract stack is load-bearing.&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The visible agent stack is easy to list: model, prompt, memory, tools, MCP servers, subagents, framework, evals.&lt;/p&gt;

&lt;p&gt;That stack matters. It is also not the operating system.&lt;/p&gt;

&lt;p&gt;The operating system is the set of contracts each layer creates.&lt;/p&gt;

&lt;p&gt;What is the agent for? What state may it see? What state may it preserve? Which tools may it call? Which tools are intentionally absent? What can it change? What must it prove before the change becomes binding? What does the harness log? What does the evaluation score actually cover? When does the agent ask, abstain, or escalate? What permission disappears after a bad run?&lt;/p&gt;

&lt;p&gt;That is the real setup.&lt;/p&gt;

&lt;p&gt;Not the list of tools.&lt;/p&gt;

&lt;p&gt;The set of boundaries that decides what the tools mean.&lt;/p&gt;

&lt;p&gt;The generic setup stack asks what you connected.&lt;/p&gt;

&lt;p&gt;The contract stack asks what you can trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;An agent is a control loop, not a prompt with ambition&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;An agent is an outer control loop wrapped around a generator. It plans, reads state, chooses tools, acts, observes, repairs, escalates, and decides whether to continue.&lt;/p&gt;

&lt;p&gt;The failure rarely sits in one glamorous place.&lt;/p&gt;

&lt;p&gt;It can sit in the planner. It can sit in retrieval. It can sit in a tool description. It can sit in retry logic. It can sit in a hidden assumption about whether the world waits while the agent thinks.&lt;/p&gt;

&lt;p&gt;That is why framework comparisons are often less useful than they look.&lt;/p&gt;

&lt;p&gt;The distinction that matters is which parts of the loop are explicit enough to inspect.&lt;/p&gt;

&lt;p&gt;If planning is hidden inside one long natural-language instruction, you cannot repair planning without rewriting the whole prompt.&lt;/p&gt;

&lt;p&gt;If memory is just a growing transcript, you cannot tell whether the agent remembered, retrieved, inferred, or hallucinated.&lt;/p&gt;

&lt;p&gt;If tool choice is unlogged, you cannot tell whether the answer is wrong because the model reasoned badly or because it called the wrong thing.&lt;/p&gt;

&lt;p&gt;If evaluation is one final pass/fail number, you cannot tell whether the agent failed at discovery, parameters, sequencing, recovery, escalation, or judgement.&lt;/p&gt;

&lt;p&gt;Agents do not become reliable when the setup becomes more impressive.&lt;/p&gt;

&lt;p&gt;They become reliable when failure has somewhere specific to land.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;MCP is not magic glue&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;MCP matters. Skills matter. Connectors matter.&lt;/p&gt;

&lt;p&gt;But their importance is often described backwards.&lt;/p&gt;

&lt;p&gt;The lazy version says MCP is valuable because it gives agents more tools.&lt;/p&gt;

&lt;p&gt;The better version says MCP is valuable because it makes the tool boundary explicit enough to inspect, version, test, authorise, and debug.&lt;/p&gt;

&lt;p&gt;A tool is not neutral plumbing. A tool description tells a nondeterministic system what an action means. The name, parameters, return shape, error messages, and allowed mutations all change behaviour.&lt;/p&gt;

&lt;p&gt;A tool built for a human developer is not automatically a good tool for an agent. Humans carry missing context. Agents need the contract written down.&lt;/p&gt;

&lt;p&gt;That is why more tools can make an agent worse.&lt;/p&gt;

&lt;p&gt;At small scale, tool access feels like freedom. At larger scale, tool access becomes search. The agent has to identify the right tool, pass valid parameters, recover from partial failure, and avoid inventing a successful trace when the tool call failed.&lt;/p&gt;

&lt;p&gt;If you expose every API endpoint as a tool, you do not have a powerful agent surface.&lt;/p&gt;

&lt;p&gt;You have a vocabulary problem with write access.&lt;/p&gt;

&lt;p&gt;The mature move is not “connect everything.”&lt;/p&gt;

&lt;p&gt;The mature move is to design the smallest tool surface that lets the agent do the job, then make every tool contract legible. What does the tool do. When should it be used. What the return value proves, and what it does not prove. What failures look like. Which calls are read-only, which mutate state, which require approval. Where the trace goes.&lt;/p&gt;

&lt;p&gt;That is how you stop a transcript from becoming the only place your operating system exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Skills are not prompt snippets&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The same mistake happens with skills.&lt;/p&gt;

&lt;p&gt;People treat skills as better prompts: a &lt;code&gt;SKILL.md&lt;/code&gt;, a few examples, some instructions, maybe a script. Useful. Portable. Easy to share.&lt;/p&gt;

&lt;p&gt;But a serious skill is not a prompt snippet.&lt;/p&gt;

&lt;p&gt;It is packaged operating knowledge.&lt;/p&gt;

&lt;p&gt;It should contain a trigger, a procedure, a boundary, gotchas, and a failure mode.&lt;/p&gt;

&lt;p&gt;The “gotchas” are usually the most valuable part. The model often already knows the happy path. What it does not know is your local scar tissue: which API lies, which file must not be edited, which naming convention breaks deployment, which customer segment changes the policy.&lt;/p&gt;

&lt;p&gt;That is why generic skill catalogues have a ceiling.&lt;/p&gt;

&lt;p&gt;They can teach a model the common workflow.&lt;/p&gt;

&lt;p&gt;They cannot teach it which parts of your workflow are load-bearing unless you package that knowledge yourself.&lt;/p&gt;

&lt;p&gt;Skills are valuable because they let operational knowledge travel across sessions and agents. They are dangerous when they activate at the wrong time, compose implicitly into deeper graphs nobody intended, or grant state-changing behaviour without a permission contract.&lt;/p&gt;

&lt;p&gt;The real question is sharper:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When this skill activates, what decision is it allowed to influence?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If nobody can answer that, the skill is just a more durable way to make the wrong move.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Memory is governed state, not a bigger past&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Memory has the same problem.&lt;/p&gt;

&lt;p&gt;Every agent product wants to promise memory. It sounds obvious. The agent should remember the user, the project, the codebase, the customer history, the prior decision, the mistake from last time.&lt;/p&gt;

&lt;p&gt;But memory is not “more context.”&lt;/p&gt;

&lt;p&gt;Memory is a four-part contract: what gets written, how it is organised, how it is retrieved, how it is governed.&lt;/p&gt;

&lt;p&gt;If the agent writes too much, memory becomes sludge.&lt;/p&gt;

&lt;p&gt;If it summarises badly, memory becomes distortion.&lt;/p&gt;

&lt;p&gt;If it retrieves by similarity alone, memory becomes vibes with citations.&lt;/p&gt;

&lt;p&gt;If it never forgets, memory becomes context poisoning.&lt;/p&gt;

&lt;p&gt;If it cannot show why a memory was used, memory becomes an invisible authority.&lt;/p&gt;

&lt;p&gt;The memory question worth asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Which state should survive because it will improve future decisions, and which state should expire because it will poison them?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a contract question.&lt;/p&gt;

&lt;p&gt;It is also why a 500-word, well-maintained project note can outperform a giant chat history. The smaller note has a job. The transcript merely has volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The harness is where autonomy becomes measurable&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Most agent demos make the model look like the protagonist.&lt;/p&gt;

&lt;p&gt;In production, the harness is the protagonist.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>machinelearning</category>
      <category>architecture</category>
    </item>
    <item>
      <title>You Only Hold Four Thoughts</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:55 +0000</pubDate>
      <link>https://dev.to/harryfloyd/you-only-hold-four-thoughts-2p7j</link>
      <guid>https://dev.to/harryfloyd/you-only-hold-four-thoughts-2p7j</guid>
      <description>&lt;h1&gt;
  
  
  You Only Hold Four Thoughts
&lt;/h1&gt;

&lt;p&gt;Try to multiply 47 by 83 in your head. The answer is not the point. Watch what happens while you reach for it. You hold 47, you hold 83, you start on the partial products, and somewhere around the third one the first number goes soft. You reach for a pen, because the problem outgrew the place you were keeping it.&lt;/p&gt;

&lt;p&gt;That ceiling is real and it is low. The cognitive scientist Nelson Cowan spent years measuring it and put the number at about four. Not the seven you half-remember from an old paper, but three to five distinct things held in mind at once. 1 Four. That is the working capacity of the most sophisticated object in the known universe.&lt;/p&gt;

&lt;p&gt;Everything we call getting smarter has been a way around that four. The history of human intelligence is the history of putting thoughts somewhere other than the head, and it runs as a stack, each layer holding what the one below it cannot.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The first rung is paper&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Reaching for the pen looks like a small surrender. It is the oldest cognitive upgrade there is. The moment you write 47 above 83 and start stacking partial products, you are thinking about six or seven things at once, because the paper is holding all but the one you are working on.&lt;/p&gt;

&lt;p&gt;Justin Sung, who teaches learning for a living, puts it more sharply. Writing is not the thing you do after you have reached clarity. Writing is what produces the clarity. 2 The page becomes the workspace where the thought turns real, because your four slots are freed to do the actual reasoning while the page remembers the rest.&lt;/p&gt;

&lt;p&gt;This is also why handwriting beats typing. It is far slower than thinking, and that slowness forces you to compress, to decide what is worth the stroke. The friction is not a tax on the process. The friction is the process. A page of notes you struggled to write holds more than a page you copied without resistance.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The page is not a transcript of a finished thought. It is the workspace where the thought becomes possible.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The rung most people never name&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;In 1998 two philosophers, Andy Clark and David Chalmers, asked where the mind stops and the rest of the world begins, and gave an answer that still unsettles people. The mind, they argued, is not all in the head. 3&lt;/p&gt;

&lt;p&gt;Their example was a man named Otto, who has Alzheimer’s and carries a notebook everywhere. When Otto wants to go to the museum, he looks up the address in the notebook the way you would retrieve it from memory. The notebook does the job your hippocampus does. Clark and Chalmers argued there is no principled reason to count the notebook as any less a part of Otto’s mind than ordinary memory. Otto and his notebook are a single coupled system. The thinking happens across both.&lt;/p&gt;

&lt;p&gt;That sounds like a thought experiment until you notice you are Otto. The phone that holds every number you no longer memorise. The calendar that holds every commitment. The thinking is already distributed across you and the things you store it in. The only open question is how well the storage is built.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The rung that compounds&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A single page does not persist, does not connect, and cannot be searched. You solve the multiplication, you throw the page away, and next month you solve it again from scratch. Paper extends the moment. It does not extend across time.&lt;/p&gt;

&lt;p&gt;A structured set of notes does. When every thought you have is written as a durable, cross-linked entry, two things happen that a single page cannot. The thought survives, available to a version of you who has forgotten having it. And it connects, so that an idea from March sits one link away from a problem you only encounter in June, waiting to be useful before you knew you needed it.&lt;/p&gt;

&lt;p&gt;This is the layer where synthesis becomes possible at a scale no head can hold. No one can keep thirty sources in working memory and find the pattern across them. Four slots cannot do it, and neither can forty. But a system that has been accumulating those sources for months, with the connections already drawn, can surface a synthesis that was never available to anyone thinking alone. The structure does the remembering, which frees the human to do the seeing.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The rung we are building now&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;For most of history the top of the stack was a human reading their own notes. That is no longer the ceiling. The newest layer is a store of knowledge an AI can read, query, and build on across sessions.&lt;/p&gt;

&lt;p&gt;The builders who have lived inside this for a year keep reporting the same thing. One who runs large agent systems put it plainly: the model is the same on day 1 and day 40. The files get richer. 4 The capability of the underlying intelligence barely moves over a project. What improves is the accumulated context it can reach, the record of what was tried, what worked, what the operator decided and why. The intelligence is rented and roughly fixed. The memory is owned and compounds.&lt;/p&gt;

&lt;p&gt;An AI working from a thin prompt starts every session as a stranger. An AI working from a well-kept store of your decisions starts as a colleague who was in the room last time. The difference is not a better model. It is the same model with the rest of the stack underneath it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The intelligence is rented and roughly fixed. The memory is owned, and the memory is what compounds.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Why this is one law and not four&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This is where a productivity story becomes something larger. The machines climb the same stack you do, for the same reason, using the same move.&lt;/p&gt;

&lt;p&gt;A large model also cannot hold everything at once. Its version of the four-slot limit is the memory bandwidth of the chip, and the entire recent history of making models faster is a history of refusing to keep everything hot. FlashAttention rewrote how attention uses memory so the chip stops shuttling the same data back and forth. Key-value caching stores the work already done so it never has to be recomputed. Mixture-of-experts routing keeps a vast model mostly dormant and wakes only the part a given token needs. 5 Store state. Reuse it. Activate only what matters now.&lt;/p&gt;

&lt;p&gt;That is the same move as paper, notes, and agent memory. Externalise the state you cannot hold, and retrieve only the slice the moment requires. Human cognition scales that way. Machine cognition scales that way. The question “how do I think better” and the question “how do I run a model well” have turned out to be one question with one answer. When two separate problems collapse into the same answer, that answer is usually worth trusting.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One law runs the whole stack: externalise the state you cannot hold, and retrieve only what the moment needs. Brains and models both scale by obeying it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The trap inside the stack&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The law has a failure mode, and it is the one a second-brain enthusiast walks into first. The stack rewards retrieval, not accumulation. The instant you start optimising for the volume of what you store, you have begun to degrade the thing you were building.&lt;/p&gt;

&lt;p&gt;A note you never pull back out did no cognitive work. Ten thousand of them do less than a hundred you reach for, because the ten thousand bury the hundred. External cognition only pays off on the way back in. Storing is filing, and filing is not thinking. The discipline that keeps the stack alive is structuring everything you save so a future you, or a future agent, can find the one piece that matters without reading the other nine thousand.&lt;/p&gt;

&lt;p&gt;That is also why each rung has to be built in order. Agent memory on top of a disorganised pile of notes inherits the disorder and answers your questions confidently from a mess. The layers compound only when each one is sound. Skip a rung and you do not get the compounding. You get a faster way to retrieve noise.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The test you can run this week&lt;/strong&gt;
&lt;/h3&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>productivity</category>
      <category>analysis</category>
    </item>
    <item>
      <title>You Cannot Try to Fall Asleep</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:39 +0000</pubDate>
      <link>https://dev.to/harryfloyd/you-cannot-try-to-fall-asleep-2ced</link>
      <guid>https://dev.to/harryfloyd/you-cannot-try-to-fall-asleep-2ced</guid>
      <description>&lt;h1&gt;
  
  
  You Cannot Try to Fall Asleep
&lt;/h1&gt;




&lt;p&gt;It is ten past three. You have done the arithmetic twice already. Five hours if you drop off now, four and a bit if this carries on. So you lie very still, because turning over would be an admission, and you hold your eyes shut a little too tightly, and underneath all of it there is a low, steady wanting. You want to be asleep. And the wanting is the exact thing keeping you awake.&lt;/p&gt;

&lt;p&gt;You are failing at doing nothing. It is a strange thing to be bad at.&lt;/p&gt;

&lt;p&gt;You are a capable person. You can learn hard things and finish dull ones and drag yourself out for a run on a wet morning when every part of you would rather stay in. Effort is the most reliable tool you own. Most of what you are proud of came out of using it. And then there is this one ordinary thing, wanted more than almost anything at three in the morning, that effort cannot touch at all. The harder you try for it, the further away it goes.&lt;/p&gt;

&lt;p&gt;We file that under sleep being difficult and move on. But sleep is not the only thing built this way. It is only the place you notice it first, because you meet it every single night.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The wanting is the exact thing keeping you awake.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Try to stop thinking about someone, and watch what happens to the thinking. Try to be happy, directly, by deciding to be, and feel it thin out into a performance of itself. You cannot make yourself find a joke funny. You cannot force another person to love you by loving them harder; if anything, that is the surest way to send them off. You cannot decide to be interesting at the party, and the second you try, you are the least interesting you will be all evening. You cannot will yourself to relax, which is the cruellest one, because the trying is the tension.&lt;/p&gt;

&lt;p&gt;None of these are things you do. They are things that happen to you while you are busy doing something else. Sleep arrives while you are turning over tomorrow’s meeting, and then at some point you never quite catch, you are gone. Happiness turns up on an ordinary afternoon when you were absorbed in something and forgot to check whether you were happy. You become interesting the moment you get genuinely interested in someone else. Love shows up sideways, in the middle of doing something entirely unromantic together. Each one is a by-product. The main thing was always something else.&lt;/p&gt;

&lt;p&gt;None of this is new. The Victorians had a name for the trap. 1 They called it the paradox of hedonism, the plain observation that happiness tends to arrive only when your mind is fixed on something other than your own happiness. Aim straight at it and you miss. Aim at something worth doing and it arrives while your back is turned. Older and gentler still is the folk wisdom your grandmother had. A watched pot never boils. You will meet someone when you stop looking. Sleep comes when you stop chasing it. She was right, and she never needed a footnote.&lt;/p&gt;

&lt;p&gt;Which makes it strange that almost everything around us now says the opposite. Try harder. Optimise. Measure it, track it, put a number on it, and by watching the number, improve it.&lt;/p&gt;

&lt;p&gt;For plenty of things, that advice is sound. It genuinely works on the steps you walk, the pages you read, the money you put aside, because those are things you do, and a thing you do answers to attention and effort. Point a number at a behaviour and the behaviour usually moves.&lt;/p&gt;

&lt;p&gt;The trouble starts when we point the same instrument at something that was never a behaviour. Take the sleep tracker, the small clean example of a very large mistake. It hands you a grade out of a hundred each morning for a thing you did not do, could not have done, and had no control over while it was happening. For someone already anxious about their sleep, that grade can do exactly what you would dread. There is a name now for people whose pursuit of a better sleep score has quietly made their sleep worse. 2 They lie there trying to earn the number, and the trying keeps them up, and in the morning the number confirms the bad night, so tomorrow they try harder still. The scoreboard becomes the insomnia.&lt;/p&gt;

&lt;p&gt;And it does not stop at sleep. We keep a scoreboard on our own happiness now, rating the day, wondering whether we are as content as we ought to be by this point, which is a reliable way to stop being content at all. We tally our friendships, our rest, our worth, treating the whole illegible middle of a life as figures to be raised. A thing that arrives sideways cannot survive being stared at head-on. Measured, it becomes work. Graded, it becomes a test you are failing. You get all the pressure of a target and none of the thing the target was for.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Measured, it becomes work. Graded, it becomes a test you are failing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So what do you actually do, if trying is the problem and not-trying sounds like giving up?&lt;/p&gt;

&lt;p&gt;You learn, slowly and against every instinct this age has trained into you, to tell two kinds of thing apart.&lt;/p&gt;

&lt;p&gt;Some things in your life answer to effort. You can decide the hour you go to bed. You can decide the phone leaves the room. You can decide to show up, to sit down at the desk, to be kind without keeping score, to call your mother, to put yourself in the path of the people you might one day come to love. Those are conditions. Conditions are real work, and they are what you can actually reach.&lt;/p&gt;

&lt;p&gt;And then there is everything the conditions are for. Sleep. Ease. Delight. Being loved. Feeling rested. Those you cannot reach for. You can only build the conditions, honestly and without cheating, and then do the single hardest thing a person can do, which is leave the outcome alone.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Leave the outcome alone.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That last part feels like surrender, and it is, and that is exactly why it is so hard. Every part of you wants to grip. Gripping feels like caring. It feels like doing your part. But on this entire class of things, the grip is the surest way to fail, and loosening it is not laziness. It is a skill, and it may be the deepest one a life asks of you. Some effort &lt;a href="https://harryfloyd.substack.com/p/difficulty-was-making-you" rel="noopener noreferrer"&gt;quietly builds you&lt;/a&gt;. This is the other kind, spent on the one thing effort can only spoil.&lt;/p&gt;

&lt;p&gt;It is ten past three somewhere, and you are lying very still, doing arithmetic in the dark. You have already done your part. The room is dark, the day is behind you, there is nothing left to arrange. There is nothing left to do but the one thing you cannot do, which is try. So stop. There was never anything there to try for. You have set the conditions. Now let yourself be no use at all for a while.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This is not the usual thing here. I mostly write about AI, product and markets, and the structures underneath them. This is the same habit of looking, pointed at something a lot more ordinary, and I wrote it because I kept meeting it at three in the morning rather than at a desk.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;_New to The Durability Curve? It is a standing argument about what survives when the tools get powerful and the surface gets cheap.&lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=article&amp;amp;utm_medium=web&amp;amp;utm_campaign=cannot-try-to-fall-asleep" rel="noopener noreferrer"&gt;Subscribe&lt;/a&gt; for the rest, or &lt;a href="https://harryfloyd.substack.com/p/start-here-what-survives-when-the" rel="noopener noreferrer"&gt;start with what survives&lt;/a&gt;.&lt;br&gt;&lt;br&gt;
_&lt;/p&gt;

&lt;p&gt;1&lt;/p&gt;

&lt;p&gt;The phrase belongs to the nineteenth-century philosopher Henry Sidgwick, and John Stuart Mill put it plainly in his autobiography: those are happiest, he wrote, who have their minds fixed on some object other than their own happiness, and who find happiness by the way. It is a very old idea with a long line of owners, which is part of why it is worth trusting.&lt;/p&gt;

&lt;p&gt;2&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>analysis</category>
    </item>
    <item>
      <title>Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:23 +0000</pubDate>
      <link>https://dev.to/harryfloyd/three-hidden-bottlenecks-the-ai-buildout-has-already-moved-past-gpus-18c1</link>
      <guid>https://dev.to/harryfloyd/three-hidden-bottlenecks-the-ai-buildout-has-already-moved-past-gpus-18c1</guid>
      <description>&lt;h1&gt;
  
  
  Three Hidden Bottlenecks the AI Buildout Has Already Moved Past GPUs
&lt;/h1&gt;

&lt;p&gt;Bloom Energy reported Q1 2026 revenue of $751 million. That number was 130 percent higher than the prior year, 42 percent above consensus, and triggered a full-year guidance raise to $3.6 billion 1. Most of the post-earnings coverage read the print as a fuel cell company finally turning operationally profitable.&lt;/p&gt;

&lt;p&gt;The print is not a fuel cell story. It is the canonical evidence that the AI infrastructure bottleneck has migrated past compute.&lt;/p&gt;

&lt;p&gt;For two years the consensus model for AI capex has anchored on GPU shipments. NVIDIA, AMD, the hyperscaler capex disclosures, the analyst models all priced compute as the load-bearing constraint. The reasoning was straightforward: training runs scaled, GPU clusters grew from 5,000 units to 50,000 to 100,000, and the company that supplied the silicon owned the bottleneck.&lt;/p&gt;

&lt;p&gt;The reasoning was correct in 2023. It became incomplete in 2024. By 2026 it has become a rear-view mirror.&lt;/p&gt;

&lt;p&gt;The analyst models that price AI on GPU shipments are not wrong about GPUs being important. They are wrong about GPUs being scarce. The supply-side data has been telling a different story for three quarters now, and Bloom Energy’s print is the most recent confirmation. The bottleneck moved. It always does. &lt;em&gt;The binding constraint never disappears. It only migrates to the next layer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The question that matters now is which layer the binding constraint has migrated to. Three layers have evidence pointing at them, none of which are GPUs, and the layers compose into a single observation about where AI capex goes once the compute layer has been solved.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The first layer: power, and the 128-week wait&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Behind every large GPU cluster sits a power-delivery infrastructure that takes longer to build than the cluster itself. Power transformers, the equipment that steps utility-scale voltage down to data-centre-usable voltage, have 80 to 128 week lead times right now 2. Cleveland-Cliffs is the only domestic US producer of the grain-oriented electrical steel that every transformer core requires 3. The grid interconnection queue at major US utilities runs five-plus years for new high-voltage data centre loads 4.&lt;/p&gt;

&lt;p&gt;This is the layer where Bloom Energy fits, and where the print becomes legible. Solid oxide fuel cells generate power on-site, behind the meter, without queueing for grid interconnection. A hyperscaler that wants 100 megawatts of power in eighteen months and cannot get it from the grid for five years buys Bloom Energy units. The fuel cell technology is twenty years old. The 130 percent revenue growth is the price of how binding the power constraint has become.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A hyperscaler that wants 100 megawatts in eighteen months and cannot get it from the grid for five years buys Bloom Energy units. The 130 percent revenue growth is the price of how binding the power constraint has become.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For beginners: what does “behind the meter” mean?&lt;/strong&gt; A utility meter measures power coming into a building from the grid. &lt;em&gt;Behind the meter&lt;/em&gt; means power generated on the customer’s side of that meter, so the grid never sees it and never has to plan for it. Bloom Energy’s fuel cells are behind-the-meter generation. That is why the eighteen-month installation timeline is the only one that matters for a hyperscaler who cannot wait five years for grid interconnection.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The falsifier for the power layer is specific. If transformer lead times compress below 52 weeks within two consecutive quarters, or if hyperscaler 24/7 firm clean power purchase agreements (PPAs) at 15-year tenors are consistently signed below $80 per megawatt-hour, the constraint has eased and the behind-the-meter premium decays. Watch the second of those harder than the first. Hyperscalers will pay whatever the grid cannot deliver fast enough, and the PPA price is where that desperation gets numerical.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The second layer: metal, and the recycling angle nobody priced&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The compute layer requires copper. The power layer requires copper. The interconnect layer requires copper. By 2030, AI data centres alone will be calling on roughly 7 percent of all the copper the world digs up in a year, from a demand source that did not meaningfully exist five years ago. The math is straightforward. Hyperscale AI sites consume 40 to 50 tons of copper for every megawatt of IT capacity 5. The US has 85 gigawatts of new pipeline through 2030 6, with 35 gigawatts already under construction across North America 7. Wood Mackenzie projects 1.1 million tonnes per year of grid copper demand from data centres alone 8; BloombergNEF projects another 572,000 tonnes peaking in 2028 inside the facilities themselves 9. Combined, that approaches 1.7 million tonnes per year against global mine output of roughly 23 million tonnes annually 10. One new demand source, 7 percent of every mine on earth, on top of every other demand the market already cannot meet.&lt;/p&gt;

&lt;p&gt;Mine capacity does not flex on the timescales the buildout requires. Copper mines take a decade from greenfield discovery to first commercial shipment. The buildout is happening on a one-to-three year horizon. There is no path where new mining capacity meets new data-centre demand.&lt;/p&gt;

&lt;p&gt;The consensus copper-AI thesis names the major miners: Freeport-McMoRan, Southern Copper, BHP, Rio Tinto. The miners are the obvious read. The recycling angle is the underfollowed one. Aurubis is a German specialty metals conglomerate that runs the largest secondary copper smelting capacity in Europe and is building the first US secondary smelter. Recycling output can flex on the timescales primary mining cannot. The structural shift is from &lt;em&gt;mining is the bottleneck&lt;/em&gt; to &lt;em&gt;recycling is the relief valve&lt;/em&gt;. The equity that captures the relief valve trades at approximately 0.4 times price-to-sales 11. The market reads Aurubis as a commodity cyclical. The multiple ignores the data-centre demand curve.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The structural shift is from “mining is the bottleneck” to “recycling is the relief valve”. The equity that captures the relief valve trades at roughly 0.4 times price-to-sales. The multiple ignores the data-centre demand curve.&lt;/p&gt;

&lt;p&gt;This publication tracks where capital is migrating before the analyst models reprice it. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?" rel="noopener noreferrer"&gt;Subscribe now&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The falsifier for the metal layer is observable and time-bound. If primary copper-mine output growth exceeds 10 percent year-over-year for two consecutive years, the supply-shortage premium for recyclers compresses. The fallback test: if hyperscaler-driven data-centre permitting decelerates by more than 30 percent year-over-year, the demand assumption breaks before the supply assumption fires. Watch the permitting numbers monthly. The construction pipeline is the leading indicator of the copper demand curve.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The third layer: detection, and the $151 billion question&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The third layer is the most speculative of the three, and also the one where the supply side is voting hardest. The reason markets have not priced it yet is that the contract that creates it was only finalised in January 2026. SHIELD is the Scalable Homeland Innovative Enterprise Layered Defense vehicle: a $151 billion ten-year contract the Missile Defense Agency awarded as the primary acquisition framework for the broader Golden Dome missile-defence initiative 12. Golden Dome itself sits above SHIELD as the umbrella programme, with the Pentagon’s own ten-year cost estimate at approximately $185 billion and the Congressional Budget Office’s May 2026 analysis projecting up to $1.2 trillion over twenty years if a full space-based interceptor layer is built out 13. The MDA selected 2,440 firms as qualified SHIELD vendors across three tranches in late 2025 and early 2026. Holding a SHIELD position confers eligibility to compete for individual task orders, not guaranteed funding; task-order competitions are now beginning.&lt;/p&gt;

&lt;p&gt;The data layer of Golden Dome (the satellites and ground-segment processing that detect, classify, and track aerial threats) is a procurement category that did not meaningfully exist five years ago. Spire Global is a publicly-traded satellite-data company at roughly $700 million market cap 14 with a remaining-performance-obligations backlog above $200 million, equivalent to about three times trailing twelve-month revenue 15. Their core revenue stream is Global Navigation Satellite System (GNSS) radio-occultation weather data, maritime Automatic Identification System (AIS) tracking, and radio frequency (RF) signal monitoring. Each of those data feeds is dual-use. The same instruments serve weather forecasting, shipping logistics, and defence persistent surveillance.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>architecture</category>
      <category>investing</category>
    </item>
    <item>
      <title>The Substrate Map</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:41:08 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-substrate-map-3hfh</link>
      <guid>https://dev.to/harryfloyd/the-substrate-map-3hfh</guid>
      <description>&lt;h1&gt;
  
  
  The Substrate Map
&lt;/h1&gt;

&lt;p&gt;Most teams can tell you what they shipped.&lt;/p&gt;

&lt;p&gt;Fewer can tell you what will still matter after the next large change.&lt;/p&gt;

&lt;p&gt;That is the gap the Substrate Map is built for.&lt;/p&gt;

&lt;p&gt;It is a free one-page taxonomy for separating canopy from substrate in your own work. The canopy is the visible layer: prompts, model choices, demo polish, current benchmark scores, frameworks, UI surfaces, launch artefacts. The substrate is the part that keeps doing work when the surface gets repriced: data-quality discipline, eval contracts, workflow integration, trust packaging, domain-specific failure memory, and the proprietary signal the next model release does not have.&lt;/p&gt;

&lt;p&gt;The tool gives you a 10-minute exercise:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Open the last 90 days of engineering tickets, product launches, roadmap decisions, or investment decisions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Tag each item as substrate or canopy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Compute the ratio.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The number is blunt on purpose.&lt;/p&gt;

&lt;p&gt;If the last 90 days were mostly canopy, the next release can reset most of what you built. If the split is 50/50, you are probably normal but not especially durable. If the work is mostly substrate, protect it. That is the work compounding underneath the visible output.&lt;/p&gt;

&lt;p&gt;The most useful part is the boundary rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If your team cannot agree which column an item belongs in, tag it as canopy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Disagreement at the boundary means the substrate work has not been made explicit yet.&lt;/p&gt;

&lt;p&gt;That makes the map useful before a planning meeting. Instead of arguing about whether a roadmap “feels strategic”, you can ask which work would still matter if the model, market, channel, or buyer changed. The conversation gets harder to fake because each item has to be placed in a column.&lt;/p&gt;

&lt;p&gt;The PDF is deliberately simple: one map, one exercise, one ratio. It is not a strategy deck. It is the first instrument you run when the team is shipping a lot but cannot say what is compounding.&lt;/p&gt;

&lt;p&gt;Use it on a roadmap. Use it on a product backlog. Use it on a portfolio. Use it before a planning cycle where everyone is about to argue from vibes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Download.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Substrate Map&lt;/p&gt;

&lt;p&gt;105KB ∙ PDF file&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/api/v1/file/5b324df8-e43f-486f-823b-7213b88910b4.pdf" rel="noopener noreferrer"&gt;Download&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/api/v1/file/5b324df8-e43f-486f-823b-7213b88910b4.pdf" rel="noopener noreferrer"&gt;Download&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;More free tools like this.&lt;/strong&gt; &lt;em&gt;Subscribe to get the next durability-lens resource the day it ships.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Substrate Map is a companion to &lt;em&gt;&lt;a href="https://harryfloyd.substack.com/p/the-forest-floor-is-the-product?utm_source=resource-landing&amp;amp;utm_medium=internal&amp;amp;utm_campaign=substrate-map-2026-05-02" rel="noopener noreferrer"&gt;The Forest Floor Is the Product&lt;/a&gt;&lt;/em&gt; , the essay that develops the substrate-vs-canopy lens across ecosystems, software, knowledge work, and capital allocation.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;What percentage of your last 90 days was substrate?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>testing</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>The Stable Liar</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:40:49 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-stable-liar-31no</link>
      <guid>https://dev.to/harryfloyd/the-stable-liar-31no</guid>
      <description>&lt;h1&gt;
  
  
  The Stable Liar
&lt;/h1&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The dashboard was green for eight quarters&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The most dangerous number on a dashboard is the one that has stayed green the longest, and the way it fails has a shape you have probably watched up close.&lt;/p&gt;

&lt;p&gt;For eight straight quarters the dashboard holds green. Revenue up and to the right. Retention flat and healthy. NPS in the fifties. Every board meeting opens on the same slide and closes on the same nod. The plan is working. Then, six months after the eighth green quarter, the business the dashboard was supposed to describe nearly falls over.&lt;/p&gt;

&lt;p&gt;Pull the post-mortem apart and the easy story is that the numbers lied. They did not. Every quarter the dashboard reports something true: customers are still paying, logins are still happening, the survey scores are still fine. All of it accurate. The failure is quieter and worse than a lie. The words behind the numbers change meaning while the numbers stand still. “Retention” still counts the same logins, but a login has stopped predicting a customer who will renew. The metric keeps its shape long after the thing it measured has walked out of the room.&lt;/p&gt;

&lt;p&gt;Anyone who has run a team has felt a smaller version of this. The number you trusted most became the number that surprised you most. You were not lied to. You were tracking something that used to mean one thing and quietly came to mean another, and the dashboard had no way to tell you the meaning had moved.&lt;/p&gt;

&lt;p&gt;This is the stable liar: a number that goes on looking right long after it stopped being right. It is a structural property of measurement under pressure, and it has a law underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Why every optimised metric drifts&lt;/strong&gt;
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;A metric is a substitution: you replace the thing you care about with something you can count, and the gap between them is where the trouble lives.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Start with the substitution. You cannot measure value, loyalty, insight, or health directly, so you pick a proxy you can count. Revenue stands in for value. NPS stands in for loyalty. Citations stand in for insight. The proxy is never the thing. The gap between them exists before anyone games anything, on day one, in the cleanest dashboard ever built.&lt;/p&gt;

&lt;p&gt;That gap stays small only while no one leans on it. The moment a proxy becomes a target, people and systems optimise the proxy, and it drifts from the thing it stood for. Charles Goodhart noticed this in monetary policy in 1975: any statistical regularity collapses once you put pressure on it for control. Marilyn Strathern later compressed it into the line everyone quotes. When a measure becomes a target, it stops being a good measure. 1 The relationship erodes precisely because you started using it. Feeding a signal back into the system it measures changes the system.&lt;/p&gt;

&lt;p&gt;The third move is the dangerous one. The erosion is invisible to the metric itself. A dashboard cannot report “I am becoming less valid.” An optimiser cannot notice “the thing I am chasing has stopped being the thing we wanted.” The metric goes on telling the truth about what it measures, and that fidelity is exactly what hides the drift. The number is honest. Its meaning is gone.&lt;/p&gt;

&lt;p&gt;Substitution, erosion, blindness. None of them require a villain. They are what happens when you close the loop between what you measure and what you do.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The three faces of a lying metric&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Once you accept that drift is structural, the useful question becomes diagnostic. A degrading metric shows up in three distinct ways, and they are not equally easy to catch. Mistake one for another and the standard fix makes things worse.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Collapse&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The first face is loud. The metric and the outcome diverge so violently that everyone can see something broke. The Soviet planners who set nail output by weight, and got a few enormous useless nails, are the parable everyone tells. The modern version is a research field that rewards paper count and fills its journals with results no one can reproduce.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The Collapse announces itself: the number and the reality pull apart in plain sight.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the easy case, even though it feels like a crisis. The signal is noisy and obvious. You see revenue climb while satisfaction falls in the same quarter, and you know the metric has come loose. Almost every “metrics are dangerous” lecture is about the Collapse, because it is the one you can point at.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Hollowing&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The second face is quiet, and most operators never name it. The metric stays healthy while the system underneath hollows out. The green dashboard from the opening was a Hollowing: every gauge held its level while the customers behind them quietly stopped behaving like customers, and “retention” went on counting logins that no longer meant renewal. The same pattern runs everywhere once you know its shape. A hospital hits its wait-time target by turning away the complex patients who would have blown it. A support team holds CSAT steady by making the survey harder to find. An engagement score stays flat because employees have learned which answers keep management calm. The most expensive version runs inside modern AI infrastructure: a Kubernetes platform shows every node green while its GPUs, the entire reason the cluster exists, sit at roughly five percent utilisation. 2&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The Hollowing leaves the number standing while the meaning quietly walks out.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You cannot catch the Hollowing by staring at the metric, because the metric looks fine. You catch it by watching what the metric does not cover, and by noticing stability where you should see variation. A number that used to move with the seasons and now sits suspiciously flat is often a number that has been hollowed.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The Inversion&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The third face is the one that ends companies, and careers, and occasionally institutions. Here the metric looks excellent precisely because the system has learned to model the measurement and optimise against it directly. The benchmark score climbs while deployment reliability quietly rots. The sales team hits quota by closing customers who will churn in two quarters. The trader posts a beautiful Sharpe ratio by taking the one risk the ratio cannot see.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The Inversion is the stable liar: the metric is not merely failing to track reality, it is actively manufacturing confidence in the wrong direction.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the hardest face to detect, because the absence of any warning sign is itself the warning. The dashboard supports the wrong conclusion with full conviction. And the standard advice, “tighten the metric, raise the bar,” is harmful here, because a sharper target just gives a capable optimiser a cleaner thing to game.&lt;/p&gt;

&lt;p&gt;Modern AI evaluation is where the Inversion is easiest to see, though it shares the stage with cruder failures worth separating out: contamination, where test items leak into the training data; overfitting to the eval’s own distribution; and plain weak test design. The Inversion proper is narrower. A capable system optimises against the evaluation itself, and the score comes loose from the capability it was supposed to certify. That looseness shows up even before any deliberate gaming. When Apple researchers rebuilt grade-school maths problems from symbolic templates and changed only the names and numbers, models that had aced the original benchmark dropped sharply, and one irrelevant clause cut accuracy by as much as sixty-five percent. 3 The benchmark had been reporting reasoning. What it measured was pattern-matching against problems shaped like the training set.&lt;/p&gt;

&lt;p&gt;The deeper version is already here: a capable enough model can represent the fact that it is being tested and behave differently when it notices. Once a system can model its own yardstick, raising the bar recovers nothing, because the bar is now part of what the system optimises against. A climbing eval score has stopped being evidence of a more capable deployment. It is evidence that the score went up.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Why telling them apart is the whole skill&lt;/strong&gt;
&lt;/h2&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>analysis</category>
      <category>research</category>
    </item>
    <item>
      <title>The SpaceX IPO Is Not What You Think You're Buying</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:40:34 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-spacex-ipo-is-not-what-you-think-youre-buying-23jj</link>
      <guid>https://dev.to/harryfloyd/the-spacex-ipo-is-not-what-you-think-youre-buying-23jj</guid>
      <description>&lt;h1&gt;
  
  
  The SpaceX IPO Is Not What You Think You're Buying
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;This is an analytical framework, not financial advice.&lt;br&gt;&lt;br&gt;
Reported IPO terms remain provisional until SpaceX publishes its S-1. 1&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The wrong question is already forming.&lt;/p&gt;

&lt;p&gt;It sounds sophisticated because it has a ticker-shaped answer: would you buy SpaceX?&lt;/p&gt;

&lt;p&gt;That question is too small. It compresses too many different things into one emotional decision. It turns a complicated offering into a referendum on rockets, Elon Musk, Mars, Starlink dishes, government contracts, xAI, retail access, and the idea that the future should be investable.&lt;/p&gt;

&lt;p&gt;The better question is stranger and more useful:&lt;/p&gt;

&lt;p&gt;What, exactly, would you be buying?&lt;/p&gt;

&lt;p&gt;Not just legally. Structurally.&lt;/p&gt;

&lt;p&gt;If the reported structure holds, the offering would be more than SpaceX selling shares. It would put Starlink’s cash flows, Falcon’s industrial proof, Starship’s option value, sovereign demand, xAI’s capital appetite, Cursor’s developer-workflow distribution, orbital-compute ambition, and Musk-controlled governance into one public-market instrument.&lt;/p&gt;

&lt;p&gt;The danger is not that investors will admire SpaceX. They should. The danger is that investors will price the bundle as if every layer is already proven, while receiving the rights of a minority passenger.&lt;/p&gt;

&lt;p&gt;The structure is simple: Starlink earns. Starship, xAI, Cursor, and orbital compute may consume. Governance decides who benefits. Price decides whether any of it matters.&lt;/p&gt;

&lt;p&gt;That is the lens to keep through the whole piece. This is not mainly a story about whether SpaceX is impressive. It is a story about what happens when an extraordinary private company becomes a public-market instrument. The company has an operating reality. The market has a price. The filing is the translation layer between them.&lt;/p&gt;

&lt;p&gt;That translation is where investors get hurt.&lt;/p&gt;

&lt;p&gt;That distinction matters because the reported SpaceX IPO would not be a normal listing. Reuters has reported that SpaceX confidentially filed for a U.S. IPO, that an early June roadshow is being targeted, and that the company could seek a valuation as high as roughly $1.75 trillion with a raise that could reach around $75 billion. 2 Reuters has also reported filing-excerpt details on Starlink economics, xAI losses, and governance. Until the prospectus is public, those are reported claims, not final terms. But even as provisional reporting, they reveal the shape of the problem.&lt;/p&gt;

&lt;p&gt;At that scale, admiration is the easy part. The harder job is deciding which future has already been capitalised into the price.&lt;/p&gt;

&lt;p&gt;That is the real IPO question.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Great Company, Wrong Question&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Public markets are very good at turning admiration into a price. They are less good at forcing people to say which part of their admiration is already in the price.&lt;/p&gt;

&lt;p&gt;SpaceX is not a shell with a story. It is an operating machine with proof in the world. Its official launches page, checked while building this draft, showed hundreds of completed missions, hundreds of landings, hundreds of reflights, and multiple recent Falcon missions. 3 Starlink has turned satellite internet from a niche service into a mass distribution network: Starlink’s own network update said it had more than 6 million active customers globally as of July 2025, and Reuters-sourced coverage now reports that it crossed 10 million active customers in February 2026. 4 Reuters-reported filing excerpts make the commercial point sharper: Starlink reportedly generated $11.4 billion of 2025 revenue and $4.42 billion of operating profit. 5 NASA has awarded SpaceX major Artemis Human Landing System work, including a later Option B contract modification valued at roughly $1.15 billion. 6&lt;/p&gt;

&lt;p&gt;The hard question is whether Starlink’s cash engine is being sold as ownership, or used as collateral for Starship, xAI, Cursor, orbital compute, and founder-controlled optionality at a valuation where the future has already been monetised.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Six Economic Layers&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The cleanest way to read the offering, when it arrives, is not as one SpaceX story. It is as six layers stacked on top of each other.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;1. The Proof Layer: Launch Cadence&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;SpaceX’s foundational achievement is not merely that it launches rockets. It is that launch has become repeatable enough to look industrial. Reusability matters because it turns a heroic event into an operating rhythm. Cadence matters because every other layer depends on it. Starlink needs launch. Government customers need reliable access. Starship needs test frequency. The narrative of orbital infrastructure needs a company that can keep putting mass into orbit while competitors are still treating launch as a sparse event.&lt;/p&gt;

&lt;p&gt;This is the layer with the most visible proof. You can see the missions. You can see the landings. You can see the reflights. You can see the launch sites. You can see the company making launch feel less like a miracle and more like logistics.&lt;/p&gt;

&lt;p&gt;But in an IPO, visible proof is not enough. The filing has to answer whether cadence produces operating leverage. Does each incremental mission become cheaper? Are margins improving by customer type? How concentrated is demand? How much pad, range, safety, refurbishment, insurance, and failure reserve is required to keep the machine running?&lt;/p&gt;

&lt;p&gt;Launch cadence proves the machine. It does not prove the multiple.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;2. The Cash Engine: Starlink&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This may be the layer that turns SpaceX from a launch company into something closer to infrastructure. Launch gets satellites up. Starlink turns those satellites into customer relationships. That is a very different asset. A launch company sells missions. A broadband network sells recurring access. A launch company is judged by reliability and price per kilogram. A network is judged by subscribers, churn, ARPU, capacity, terminal cost, replacement capex, spectrum, enterprise mix, and distribution.&lt;/p&gt;

&lt;p&gt;The customer number matters, but it is no longer the main uncertainty. The reported 10 million-plus customer base is enough to prove scale. The deeper question is what that scale has to fund. Reuters-reported filing excerpts say Starlink produced $11.4 billion of 2025 revenue and $4.42 billion of operating profit. That changes the burden of proof. Starlink is not merely a promising broadband project. It is, on current reporting, the cash engine inside the group.&lt;/p&gt;

&lt;p&gt;The question is whether the engine is free to compound, or whether it is being asked to carry everything else.&lt;/p&gt;

&lt;p&gt;That engine is real, but it is not frictionless. The Information and syndicated market reports say Starlink’s average revenue per user fell 18% to roughly $81 a month between 2023 and 2025 as the service expanded into lower-priced plans and geographies. 7 Does the network become cheaper to serve as it grows, or does each wave of growth require new satellites, new ground infrastructure, subsidised terminals, and continuous replacement spend? Does direct-to-cell become a second distribution curve, or an expensive feature? Does enterprise, maritime, aviation, and government demand protect the economics as residential pricing compresses?&lt;/p&gt;

&lt;p&gt;The reported consolidated picture makes this sharper. Reports based on Reuters filing excerpts say SpaceX’s newly consolidated AI business posted a $6.4 billion operating loss in 2025 and consumed roughly 61% of group capex, while the combined company lost nearly $5 billion on about $18.7 billion of revenue. 8 If those numbers survive the public filing, Starlink becomes more than a growth story. It becomes the engine being asked to fund the next frontier.&lt;/p&gt;

&lt;p&gt;Starlink is the part of SpaceX that most resembles a public-market business. The investment question is whether public holders get to own its compounding, or mainly underwrite what it is being used to finance.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;3. The Option Layer: Starship&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Starship is the option layer. It is the part of the story that expands the possible future more than it explains the present. If Starship works at scale, the cost and volume assumptions around orbit change. Starlink deployment changes. Lunar logistics change. Mars changes. Orbital manufacturing, propellant depots, military logistics, and large-scale cargo all move from slideware to a different sort of conversation.&lt;/p&gt;

&lt;p&gt;But options are not cash flows. They are claims on a future state of the world.&lt;/p&gt;

&lt;p&gt;That does not make them worthless. Some of the most valuable companies in history were underpriced because people could not value their option layers. The mistake is not valuing optionality. The mistake is paying for optionality as if it has already cleared the gates.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>investing</category>
      <category>infrastructure</category>
    </item>
    <item>
      <title>The Other Half of Compute</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 09 Aug 2026 12:40:03 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-other-half-of-compute-4eom</link>
      <guid>https://dev.to/harryfloyd/the-other-half-of-compute-4eom</guid>
      <description>&lt;h1&gt;
  
  
  The Other Half of Compute
&lt;/h1&gt;

&lt;p&gt;xAI stood up its first 100,000 GPUs in Memphis in 122 days. It doubled that in another 92. By early 2026 the site, Colossus, held around 555,000 of them, building toward two gigawatts of power, for a reported 18 billion dollars. 1&lt;/p&gt;

&lt;p&gt;Two sophisticated people can look at that number and reach opposite conclusions.&lt;/p&gt;

&lt;p&gt;Jensen Huang’s view is that the only real risk is underspending. He puts the buildout at a trillion dollars and counting, and argues the company that holds back capacity loses the decade. 2 Dario Amodei and Ray Dalio sit on the other side. Amodei has said it can be rational not to buy unlimited compute, because the revenue to justify it may arrive on a timeline that bankrupts whoever guessed wrong. Dalio keeps making a narrower point: a technology can succeed completely and still ruin the people who financed it. 3&lt;/p&gt;

&lt;p&gt;Same buildout. Same dollar figure. One camp calls it the obvious move of the decade and the other calls it the setup for a wipeout. They are not disagreeing about the facts. They are reading the same number and the number is the problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;What 18 billion dollars buys&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Every token a model produces runs down a physical path. Electricity has to be generated, moved across a grid, and stepped down through transformers to a voltage a data centre can use. Chips have to be fabricated at advanced nodes, which in practice means TSMC and a single supplier of the lithography machines that make the process possible. The chips have to be wired together with optical interconnect, assembled into racks, and kept cold. None of those layers move at the same speed, and the slowest one always sets the schedule.&lt;/p&gt;

&lt;p&gt;For four years the slowest layer kept changing. In 2022 the constraint was GPUs themselves. In 2023 it was the high-bandwidth memory stacked next to them. In 2024 it was the advanced packaging that bonds the two together. By 2025 it was photonics, the lasers and transceivers that move data between racks. By 2026 it had reached power and the grid, where a new high-voltage connection can take longer to approve than the cluster takes to build. Bringing a large new source of power onto that grid now takes a median of more than four years. 4&lt;/p&gt;

&lt;p&gt;Each layer is real, each one becomes scarce in turn, and the scarcity moves to the next layer as the one before it gets solved.&lt;/p&gt;

&lt;p&gt;Call it the capacity stack. It decides one thing: how much raw compute can physically exist. It tells you what you can run. It says nothing about how much useful work comes out the other end.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The binding constraint has moved through the stack for four years straight. Chips, memory, packaging, photonics, power. Each one stayed invisible until the one before it was solved.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;The number that never makes the capex debate&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Now look at a different figure.&lt;/p&gt;

&lt;p&gt;In March 2023, running a million tokens through GPT-4 cost about 30 dollars. By the middle of 2024, the same class of capability through GPT-4o cost 2.50 dollars. By 2025 a GPT-4-grade model was available at roughly 10 cents per million tokens. 5 For the rougher GPT-3.5 tier the price fell from 20 dollars per million tokens to about 7 cents in two years, a drop of more than 250 times. Epoch AI, which tracks this carefully, finds inference prices falling somewhere between 10 and 50 times a year depending on the task. 6&lt;/p&gt;

&lt;p&gt;Almost none of that came from adding watts. The capacity stack was straining the entire time. The cost of intelligence fell by two orders of magnitude anyway. These are list prices, so some of the fall is competition between providers, but most of it is a second stack that lives inside the software layer and does work the hardware never sees.&lt;/p&gt;

&lt;p&gt;That second stack has its own layers. At the bottom is the attention kernel. The 2022 FlashAttention paper showed that a transformer was bound by memory traffic, the data shuttling between the fast and slow memory on the chip, and that rewriting the kernel to respect that traffic multiplied throughput without changing a single transistor. 7 Above it sits serving. Key-value caching, which means storing a conversation’s intermediate state instead of recomputing it on every new token, turned long contexts from a quadratic expense into something a business could afford to offer. Above that sits the model itself. Mixture-of-experts routing, the design behind Switch Transformers, broke the link between a model’s total size and the compute each token triggers, so a model can hold a trillion parameters and fire only a fraction of them per word. 8&lt;/p&gt;

&lt;p&gt;Even the hardware gains are mostly architectural rather than brute force. NVIDIA’s GB200 NVL72 rack delivers up to 30 times the inference throughput of the same number of previous-generation H100 chips, at around 25 times less energy for the same work. 9 The watts per chip went up. The useful work per watt went up far more.&lt;/p&gt;

&lt;p&gt;Each of these is a multiplier on the same physical base. Stack them and you get the hundredfold collapse in the cost of intelligence that the buildout debate never mentions.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The cost of GPT-4-class intelligence fell roughly 99 percent in two years. Almost none of that came from adding power.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Compute is a product&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Raw physical capacity, multiplied by how much useful work each unit of that capacity buys. The capacity stack sets the first term. The efficiency stack sets the second. They run on different clocks, they are built by different people, and the one that is currently scarcer sets the ceiling on what you can do.&lt;/p&gt;

&lt;p&gt;Once you read compute that way, the contradictions in the capex fight resolve.&lt;/p&gt;

&lt;p&gt;Go back to the 18 billion dollars. Jensen Huang is right that physical capacity is scarce today. A grid connection does take longer than a training run, and the firm that waits loses ground it cannot buy back at any price. Amodei is also right that the return on that capacity is uncertain. Both of them are arguing about the first term and treating the second as a constant.&lt;/p&gt;

&lt;p&gt;It is not a constant. It is improving 10 to 50 times a year. That cuts in two directions at once. A capex bill that looks insane against today’s efficiency can look cheap against next year’s, because the same site serves far more useful work for the same power. And capacity bought to serve a workload that the efficiency stack is about to make trivially cheap is capacity that strands. The danger in the buildout is &lt;strong&gt;owning the wrong term&lt;/strong&gt; : paying for raw capacity after the binding constraint has moved to the multiplier, or perfecting the multiplier when you cannot get the megawatts to run it on.&lt;/p&gt;

&lt;p&gt;Three years ago the next sentence would have sounded like a category error.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A 2-gigawatt site with a mediocre serving stack loses to a smaller site with a better one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Where the constraint goes after silicon&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;The migration does not stop at the efficiency stack either. It keeps walking.&lt;/p&gt;

&lt;p&gt;Once serving is efficient and the power is online, the slowest layer becomes the one furthest from the metal: whether an organisation can absorb what the stack has made cheap. Jensen Huang’s own example is the sharpest version of it. A 500,000-dollar engineer who consumes only 5,000 dollars of tokens a year shows the failure mode. 10 The tokens are nearly free, and the company still cannot route its own work to the capacity it already owns.&lt;/p&gt;

&lt;p&gt;This is the layer Amodei and Satya Nadella keep returning to from opposite ends of the argument. The technical stack gets good faster than institutions reorganise around it. The final constraint on compute is organisational. It is how quickly people change what they do.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The tokens are nearly free. The bottleneck is the company.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;A test you can run this week&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Take any AI bet you hold, whether it is a position, a product, or a career, and do three things.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>architecture</category>
      <category>infrastructure</category>
    </item>
  </channel>
</rss>
