<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harry Floyd</title>
    <description>The latest articles on DEV Community by Harry Floyd (@harryfloyd).</description>
    <link>https://dev.to/harryfloyd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3933548%2F522eda5f-0114-40d4-86ba-8dbaf3ef7fce.jpg</url>
      <title>DEV Community: Harry Floyd</title>
      <link>https://dev.to/harryfloyd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harryfloyd"/>
    <language>en</language>
    <item>
      <title>You Built a Number You Will Not Trust</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Thu, 08 Oct 2026 07:54:20 +0000</pubDate>
      <link>https://dev.to/harryfloyd/you-built-a-number-you-will-not-trust-5hai</link>
      <guid>https://dev.to/harryfloyd/you-built-a-number-you-will-not-trust-5hai</guid>
      <description>&lt;p&gt;&lt;em&gt;Some things you build get graded by the world. Some only ever hand you back your own guess wearing a decimal. Telling the two apart before you start is the cheapest hour you will spend.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F85tponf00ykzogv5bh50.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F85tponf00ykzogv5bh50.webp" alt="Cover art for You Built a Number You Will Not Trust" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is probably something you are part of the way through building. A spreadsheet that will score your sales leads so you know which to ring first. A scorecard that will rank the people you interviewed. A formula that will tell you which product line to drop. You have the columns and the weights, and it more or less works. And some part of you already knows you will not trust the number it gives you. When it puts the wrong lead at the top, you will override it, the way you did last time.&lt;/p&gt;

&lt;p&gt;That thing was built well and it is not going to help you, and the reason is not that you got the weights wrong. You could tune them all weekend. The reason is that you never said what would make its answer right. You built a machine to rank the leads before you had settled what a good ranking even was, so when its number surprises you, you have no honest reason to trust the surprise over your own read, and you override it. Working that out before you start is worth more than any amount of tuning, because it changes what you build.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "it works" does not tell you
&lt;/h2&gt;

&lt;p&gt;When a thing you made runs, you feel finished. The columns add up, the report generates, the score comes out. But working only tells you the thing does what you built it to do. It says nothing about the two questions that can kill the thing before you waste a weekend on it: whether you could ever tell that its answer was wrong, and whether it beats what you would have done without it. You can ask both before you write a line, and most of the cost of a wasted build is the cost of not asking.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one people skip: what would tell you it was wrong
&lt;/h2&gt;

&lt;p&gt;Start with whether you could ever tell the answer was wrong, because that is the question that gets skipped, and it is easy to hear it as a question about whether the tool runs. Those are not the same: a tool can run perfectly and still hand you an answer you could never catch out. A total, a count, a due date, an amount owed: each has an obvious check, because there is an answer outside the tool to compare it with. Those you can build and know you built them well.&lt;/p&gt;

&lt;p&gt;Past those, it comes down to one thing that is easy to miss: whether the outcome that grades the answer arrives on its own, or whether acting on the answer is what decides which outcome you ever see. A forecast can be the first kind, as long as the thing you are predicting arrives regardless of what you do with the prediction. Predict how many people will show up or how many inbound calls will arrive, and the real number turns up whatever you guessed. Footfall, call volume, next week's demand: the world returns a verdict on these whether you like it or not, which makes them possible to test and improve even when the answers are hard. You always get to find out.&lt;/p&gt;

&lt;p&gt;The lead score looks like one of those and is not. The outcome you get to see is the one the score chose for you. You ring the leads at the top, some of them buy, and the ranking looks confirmed, but the leads it sent to the bottom you never ring, so they never get the chance to prove it wrong. It can bury half your real buyers and still show you a good week, because the ones it buried are the ones you will never check. The number decides which evidence you ever see, and the mistakes that matter are the ones it never shows you. That is the dangerous kind of number: the one whose very use hides its own mistakes, so the more you lean on it the less you can see them.&lt;/p&gt;

&lt;p&gt;A tool like that leaves you two choices, and neither is simply to trust the score. You can keep the decision for yourself and build the thing that lays the evidence out, the renewal date, the value, the last time you spoke, so the call takes seconds and the decision stays where it can actually be made. Or, if you want the score, you keep a corner of the world honest: at random, ring some of the leads it tells you to skip, so it cannot set its own exam and then pass it. What you cannot do is act on it everywhere and call the result proof.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Futacfap53i5g5av0l67w.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Futacfap53i5g5av0l67w.webp" alt="One decision, before you build for it. If the outcome arrives on its own (tomorrow's demand), build the thing and keep score. If acting on the answer picks the evidence you see (the leads you never ring), hold a slice of the world back to test it, or build the aid that lays the evidence out; the call stays yours." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the question that matters is not whether the answer is a fact or a judgement. Plenty of recommendations can be tested against what actually happens, over enough cases. The question is whether you would ever find out this one was wrong, for the options it turns down as much as the ones it takes. When nothing would, you have automated the decision without earning any reason to trust it.&lt;/p&gt;

&lt;p&gt;I have watched myself skip this question, so I know why it gets skipped. Asking is quick and uncomfortable, because the honest answer is sometimes that nothing would ever tell you, and that ends the project in the first ten minutes. Building is slow and feels like progress, and every hour in makes the thing harder to walk away from. So the quick question loses to the comfortable work, and you find out at the end.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it cost me to learn this
&lt;/h2&gt;

&lt;p&gt;I learned this by paying for it twice in one evening. I had built a small tool that read a page of written instructions, the kind you might write to hand a job over to someone, and sorted each line into two piles: the lines that told you to do something, and the lines that were only describing the setup. It ran. Then I tested it against a page it had never seen, marking the lines myself first, and on the ones it was willing to call, it agreed with me a little over half the time. The tempting lesson is that the tool was bad. The real lesson was worse. I had never pinned down what the right pile was, and I could have: written down the rule for what counts as an instruction, had someone else label the same page, checked whether we agreed. That would have given me something to hold the tool to. I never did it, so the tool's score could tell me how often it matched me, but not whether it had learned a rule worth trusting.&lt;/p&gt;

&lt;p&gt;The target was buildable. I had just built everything except the target.&lt;/p&gt;

&lt;p&gt;So I rebuilt it as the other kind of tool, one that only counts, showing each section's share of the page and letting me tick what to cut. That has an obvious check, and I built it properly, with a list of thirty-four ways it could go wrong, twenty-seven test cases worked by hand, seven real pages it passed cleanly. It was correct. It was also useless, because the thing it replaced was opening the file and reading it, and it beat that by nothing. The first question had sent me to a tool that could at least be right. The second question killed it anyway. That second failure is the easier one to miss, because a tool that is correct feels like a tool that is worth having. It can be measurable, tested, and right, and still lose to the thing you would have done without it. What you would do instead belongs on the page next to the answer you are after.&lt;/p&gt;

&lt;p&gt;Later I wrote out what could have sunk each build. The thing that actually sank the first one was not among the bugs I had spent the night fixing. It had been sitting above them from the start: I never said what a right answer would be.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it on the thing you are building
&lt;/h2&gt;

&lt;p&gt;So take the thing you are building now. Before you add another column, write down the answer it is meant to give, and put it to the two questions.&lt;/p&gt;

&lt;p&gt;First, if that answer came out wrong, would anything ever tell you? If the outcome arrives independently of what you do with the answer, build it and keep score. If acting on the answer is what picks the evidence you see, randomly test some of the choices it would otherwise turn down, or leave the decision to yourself and build the tool that lays the evidence out.&lt;/p&gt;

&lt;p&gt;Second, say in one plain sentence what you would do instead if this did not exist, and why what you are building beats it. If that sentence will not come, you have your answer, and it is cheaper now than after the weekend.&lt;/p&gt;

&lt;p&gt;The tell that you skipped both: you are deep in the build and neither answer is written down anywhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmtnv13krlhscbvvojbl5.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmtnv13krlhscbvvojbl5.webp" alt="The thing you are building now, run through both questions: what real outcome would tell you the answer is right, and what does it beat. Two rows worked, two blank for your own build. A blank first column is the answer." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the questions do not do
&lt;/h2&gt;

&lt;p&gt;Two honest limits. The first is that even a target you can check can be the wrong target. Rank the leads by who converts and the tool may learn to love the small easy accounts and starve the big awkward ones, and it will be hitting the number you set the entire time. You can define success exactly and still have defined the wrong success. These two questions do not protect you from that. What they buy you is the right to find out you were wrong, not a promise you aimed at the right thing.&lt;/p&gt;

&lt;p&gt;The second is that they do not make you wise in advance. Both times, I killed the tool after I had built it, not before. Asking first lowers how often you pay, but it will not take the count to zero, because now and then the only way to learn what a thing is worth is to build it and look. The questions are there to catch the builds you start because building feels easier than asking, not to promise that you will never lose another evening.&lt;/p&gt;

&lt;p&gt;There is one last thing the questions never reach, and it was always going to be yours. When you have a target you can genuinely test, software can hand you a better answer than you would reach alone, and you should let it. What it cannot do is settle what counts as a good answer, settle how much evidence is enough, or answer for the choice to act on it. A machine can be right. Being the one who is answerable for trusting it is the part that stays with you.&lt;/p&gt;

&lt;p&gt;I'm Harry, and this is The Durability Curve, where I work out where value goes as AI gets cheaper. The aim each time is to leave you with something you can take to your own work: a test to run, or a way of seeing it you did not have before. If you'd like the next one to find you, &lt;a href="https://durabilitycurve.com/?utm_source=devto" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This essay is also on the site: &lt;a href="https://durabilitycurve.com/blog/number-you-will-not-trust/" rel="noopener noreferrer"&gt;durabilitycurve.com/blog/number-you-will-not-trust&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>machinelearning</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Ask for the Step Before the Answer</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 07 Oct 2026 07:51:18 +0000</pubDate>
      <link>https://dev.to/harryfloyd/ask-for-the-step-before-the-answer-5cg9</link>
      <guid>https://dev.to/harryfloyd/ask-for-the-step-before-the-answer-5cg9</guid>
      <description>&lt;p&gt;&lt;em&gt;Before you trust an AI answer you cannot check, ask for the step it was built from.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9r6rm3mozb2n04w87sv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9r6rm3mozb2n04w87sv.webp" alt="Cover art for Ask for the Step Before the Answer" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You asked it to do something finished. Summarise a forty-page report. Pull a growth figure out of a spreadsheet. Read five documents and tell you which supplier to choose.&lt;/p&gt;

&lt;p&gt;What came back was clean, quick, and shaped like competent work. So you sent it on, quoted the figure, went with the supplier. It looked right, and looking right was the whole of what you had.&lt;/p&gt;

&lt;p&gt;A finished answer is where good work and bad work can start to look most alike. Both arrive fluent, confident and complete. And the finished answer is usually the most expensive thing to check: to test the summary you would have to read the report, to test the figure you would have to rebuild the sum, to test the recommendation you would have to read all five documents yourself. Checking the thing you got back costs you the very work you were trying to hand off.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flb2f0gvkimw0vlt15bod.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flb2f0gvkimw0vlt15bod.webp" alt="Two smooth answers side by side, the sound one and the flawed one looking identical." width="800" height="500"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;A right answer and a wrong one arrive under the same smooth surface, and the surface is the most expensive part to check.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Ask for the step before the answer
&lt;/h2&gt;

&lt;p&gt;So do not run your check there. Ask it to stop one step earlier, at something cheaper to check, and check that first.&lt;/p&gt;

&lt;p&gt;Before it writes the essay, ask for the outline: one line per point, in the order it will make them. Before it gives you the number, ask for the inputs, the assumptions, and the sum. Before it summarises the report, ask for the passages carrying its main claims, with page numbers. Read that. Only when the step underneath holds up do you let it finish the job.&lt;/p&gt;

&lt;p&gt;Each is a smaller part of the task. A summary stands on specific claims and passages. A number stands on inputs, assumptions and a formula. An essay stands on a structure. The finished answer has compressed the sources, assumptions and decisions into one result. Moving one step back exposes those parts separately, where they are easier to inspect. A mistake that could have taken an afternoon to trace through the finished work may be sitting in the open a few lines earlier.&lt;/p&gt;

&lt;p&gt;You will not always catch it, and some errors still need an expert eye. But you have &lt;a href="https://durabilitycurve.com/blog/verification-budget-is-backwards/" rel="noopener noreferrer"&gt;made the error cheaper to catch&lt;/a&gt;, often much cheaper.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Futg373ixsimod1r579y1.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Futg373ixsimod1r579y1.webp" alt="A check mark moving back from the finished answer to the outline it was built from." width="800" height="490"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Move your check to the cheaper step: the outline, the inputs, the sources. Approve that, then let the finished answer be built from it.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not start your check where the work ends. Ask for the step underneath it, while the work is still cheap to inspect.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The way it fools you
&lt;/h2&gt;

&lt;p&gt;There is a version of this that feels the same and is not. You let it answer, then you ask how it got there.&lt;/p&gt;

&lt;p&gt;What comes back is an explanation generated after the answer already exists. You cannot know whether it faithfully represents how the answer was produced, and it is too late either way: you are inspecting an account of the work instead of creating a checkpoint before it.&lt;/p&gt;

&lt;p&gt;Asked first, the same request does something else. The answer does not exist yet, so what you approve becomes the thing you can require the answer to be built from. That is the whole difference. Ask afterwards and you get an explanation. Ask beforehand and you get a constraint.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5c1qphkpud74funjgw61.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5c1qphkpud74funjgw61.webp" alt="Two orders of asking: approval placed before the answer, explanation placed after it." width="800" height="500"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Asked first, the step becomes a checkpoint for the answer that follows. Asked afterwards, the explanation is generated with the finished answer already in place.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The two messages
&lt;/h2&gt;

&lt;p&gt;The habit is two messages. Before it starts, send the first:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"Before you give me the final answer, show me the part I should check first. For this task, that might be the sources, the inputs and assumptions, or the outline. Stop there and wait for me."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Then, once you have read that and it holds up:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"Now write the final answer using only what I approved."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first message moves your check to the cheap step. The second reduces the chance that the finished answer drifts away from the step you signed off.&lt;/p&gt;

&lt;p&gt;That is the whole thing. Ask for the step before the finished answer. Check it first. Then make the final answer come from what you approved.&lt;/p&gt;

&lt;p&gt;This is the cheap version of a larger problem: AI is most tempting to trust where checking it is hardest. I wrote about &lt;a href="https://durabilitycurve.com/blog/your-ai-looks-best-where-you-check-least/" rel="noopener noreferrer"&gt;why an AI looks best exactly where you can check it least&lt;/a&gt; for the cases where getting this wrong is expensive.&lt;/p&gt;

&lt;p&gt;I'm Harry, and this is The Durability Curve, where I work out where value goes as AI gets cheaper. The aim each time is to leave you with something you can take to your own work: a test to run, or a way of seeing it you did not have before. If you'd like the next one to find you, &lt;a href="https://durabilitycurve.com/?utm_source=devto" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This essay is also on the site: &lt;a href="https://durabilitycurve.com/blog/ask-for-the-step-before-the-answer/" rel="noopener noreferrer"&gt;durabilitycurve.com/blog/ask-for-the-step-before-the-answer&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>promptengineering</category>
      <category>productivity</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Stop Re-Priming Claude Code by Hand</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Tue, 06 Oct 2026 09:00:54 +0000</pubDate>
      <link>https://dev.to/harryfloyd/stop-re-priming-claude-code-by-hand-4hi6</link>
      <guid>https://dev.to/harryfloyd/stop-re-priming-claude-code-by-hand-4hi6</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs1yg0vf54wkyxa6548uv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs1yg0vf54wkyxa6548uv.webp" alt="A retyped session-start prompt collapses into one command, $ /prime, and a live green check returns." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What you will do:&lt;/strong&gt; take the block of context you paste at the start of every Claude Code session and put it in one file, so you type &lt;code&gt;/prime&lt;/code&gt; instead. About fifteen minutes.&lt;br&gt;
&lt;strong&gt;Who this is for:&lt;/strong&gt; you come back to the same project across many sessions, and you keep re-pasting the same "here is the project, here is where things stand" preamble before real work starts.&lt;br&gt;
&lt;strong&gt;Who should skip it:&lt;/strong&gt; if you open a fresh project every time, or never re-explain anything, there is nothing here for you.&lt;br&gt;
&lt;strong&gt;You need:&lt;/strong&gt; Claude Code (a recent 2.1.x release; the mechanics here were checked against the docs on 9 August 2026), a project you return to, and a shell.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Contents&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;One file becomes one command&lt;/li&gt;
&lt;li&gt;Make it load your real state&lt;/li&gt;
&lt;li&gt;Let it run without asking&lt;/li&gt;
&lt;/ol&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every session starts the same way. You open Claude Code and spend the first few minutes telling it where you are. This is the project. Here is what you changed last time. Here is the thing that is half-finished. Leave the migration alone. You have typed some version of that briefing fifty times, and you are the one keeping it in your head and re-entering it by hand.&lt;/p&gt;

&lt;p&gt;That is a prompt you repeat, and a prompt you repeat should be a command you type. Claude Code lets you save one as a file and invoke it with a slash. The good version does more than paste static text back at you: it runs a couple of commands and reads a couple of files first, so the context arrives already filled with your project's current state. By the end of this you will have &lt;code&gt;/prime&lt;/code&gt;, and starting a session will be one word.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. One file becomes one command
&lt;/h2&gt;

&lt;p&gt;The smallest possible version is a single file with one line in it. In your project, create &lt;code&gt;.claude/skills/prime/SKILL.md&lt;/code&gt; (a &lt;code&gt;prime&lt;/code&gt; folder with a &lt;code&gt;SKILL.md&lt;/code&gt; inside it):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;Summarise where this project stands and what I should work on next.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Save it. In Claude Code, type &lt;code&gt;/&lt;/code&gt; and &lt;code&gt;prime&lt;/code&gt; shows up in the list; run &lt;code&gt;/prime&lt;/code&gt; and the agent does what the file says. The folder name is the command: a &lt;code&gt;prime&lt;/code&gt; folder gives you &lt;code&gt;/prime&lt;/code&gt;. Claude Code watches existing skills folders and picks up a new or edited skill the moment you save it, no restart. The one exception is the very first time: if &lt;code&gt;.claude/skills/&lt;/code&gt; did not exist when you started the session, restart Claude Code once so it begins watching the new folder, and saves are live after that.&lt;/p&gt;

&lt;p&gt;That is already a command, and it already saves you the typing. But it is static. It says the same sentence every session and knows nothing about what actually changed since last time. The next step is where it earns its place.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Make it load your real state
&lt;/h2&gt;

&lt;p&gt;Two small pieces of syntax turn the file from a saved sentence into a live briefing.&lt;/p&gt;

&lt;p&gt;A line that starts with &lt;code&gt;!`command`&lt;/code&gt; runs that shell command and drops its output into the file &lt;em&gt;before Claude reads it&lt;/em&gt;. The docs call this dynamic context injection: "the command runs first, and its output gets inserted into the prompt," so the agent receives the actual data, not the instruction to go and get it (&lt;a href="https://code.claude.com/docs/en/slash-commands" rel="noopener noreferrer"&gt;Claude Code documentation: dynamic context injection and slash-command syntax&lt;/a&gt;). A line with &lt;code&gt;@path/to/file&lt;/code&gt; inlines that file's contents the same way. Put them together and &lt;code&gt;/prime&lt;/code&gt; can walk in already knowing your latest commits, your uncommitted changes, and whatever notes you keep.&lt;/p&gt;

&lt;p&gt;Open that same file and replace its one line with the fuller version below, which you can adapt to any repository. Swap &lt;code&gt;NOTES.md&lt;/code&gt; for the short file where you keep current project state, not your whole project history:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Load where this project stands and print a short situation report&lt;/span&gt;
&lt;span class="na"&gt;argument-hint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[optional&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;focus&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;area]"&lt;/span&gt;
&lt;span class="na"&gt;allowed-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Bash(git log *) Bash(git status *)&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

Recent work on this branch:
!&lt;span class="sb"&gt;`git log --oneline -10`&lt;/span&gt;

Branch and uncommitted right now:
!&lt;span class="sb"&gt;`git status --short --branch`&lt;/span&gt;

Current state notes, read in full:
@NOTES.md

Give me a four-line situation report: what branch I am on, what changed
recently, what is unfinished, and what to pick up next. If I named a focus
area ($ARGUMENTS), orient every line to it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point is what the reader on the other end receives. Not "go check git" but the actual log, the actual dirty files, the actual notes, already in front of the model, followed by a request to make sense of them. &lt;code&gt;$ARGUMENTS&lt;/code&gt; is whatever you typed after the command, so &lt;code&gt;/prime the auth refactor&lt;/code&gt; pushes "the auth refactor" into that placeholder and the report orients to it. The two lines at the top of the frontmatter are optional labels: &lt;code&gt;description&lt;/code&gt; is the text that shows beside &lt;code&gt;/prime&lt;/code&gt; in the menu, and &lt;code&gt;argument-hint&lt;/code&gt; is the grey prompt after it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvwe5apjclwfpygqx8j0n.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvwe5apjclwfpygqx8j0n.webp" alt="The two instruction lines resolve before Claude reads the file: the command line is replaced by its output and the @file by the file's contents, which is what Claude actually receives." width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To show it doing real work, here is ours. Our project is a large notes vault whose rulebook says: before you touch anything, read the pipeline status, read the working briefing, read the vault vitals, and check none of it is stale. That was three files and a freshness check we opened by hand at the start of every session. Our &lt;code&gt;/prime&lt;/code&gt; injects all of it and ends with a read. Invoked against our live state, it returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Pipeline: red. One red alert, a knowledge-review backlog; two yellow,
   the weekly cleanup twelve days overdue and a staging queue filling up.
2. Binding constraint: the output bottleneck. Six pieces ship-pending,
   all marked critical, the oldest sixty-four days. Ship before building.
3. Last session: perfected and shipped the opening piece of a new series.
4. Next: take one of the built-and-parked series pieces to publish-ready.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole payoff: the state of the project, the one thing that matters most, and where to start, assembled before I said a word. Those four lines are ours, from our injects; run the git template above and &lt;code&gt;/prime&lt;/code&gt; hands you the same shape filled with your own project, your commits and dirty files and notes in place of our pipeline and briefing.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Let it run without asking
&lt;/h2&gt;

&lt;p&gt;Unless those commands are already allowed by your own permission settings, Claude Code stops and asks the first time each injected command runs. The &lt;code&gt;allowed-tools&lt;/code&gt; line in the frontmatter above pre-approves them for you, so &lt;code&gt;/prime&lt;/code&gt; runs clean:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;allowed-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Bash(git log *) Bash(git status *)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each entry is a pattern: &lt;code&gt;Bash(git log *)&lt;/code&gt; means "any command starting with &lt;code&gt;git log&lt;/code&gt;, no prompt," and &lt;code&gt;Bash(git status *)&lt;/code&gt; does the same for &lt;code&gt;git status&lt;/code&gt;. These two pre-approve only the reads &lt;code&gt;/prime&lt;/code&gt; actually needs; anything else stays subject to your normal permission prompts. A blanket &lt;code&gt;Bash(git *)&lt;/code&gt; would pre-approve &lt;code&gt;git push&lt;/code&gt;, &lt;code&gt;git reset&lt;/code&gt;, and every other git command without a prompt too, so keep the grant to the narrow pair. It is scoped to the turn that &lt;code&gt;/prime&lt;/code&gt; runs in and clears when you send your next message, so you are not opening a standing hole in your permissions either.&lt;/p&gt;

&lt;p&gt;One honest wrinkle from ours: our freshness line runs &lt;code&gt;TZ='Europe/London' date&lt;/code&gt;, and the leading &lt;code&gt;TZ=&lt;/code&gt; assignment does not always match the command prefix cleanly, so the first run still asks once. We approve it and move on. If one of your lines keeps prompting despite an &lt;code&gt;allowed-tools&lt;/code&gt; entry, that mismatch is why; simplify the command or approve it the once.&lt;/p&gt;

&lt;p&gt;One caution that matters more once the file is not yours to begin with: a skill runs shell commands and can pre-approve its own tools, so treat one you pulled from someone else's repository like a script you are about to run. Read its &lt;code&gt;!`&lt;/code&gt; lines and its &lt;code&gt;allowed-tools&lt;/code&gt; before you invoke it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just put this in CLAUDE.md?
&lt;/h2&gt;

&lt;p&gt;If Claude Code already reads a &lt;code&gt;CLAUDE.md&lt;/code&gt; at the start of every session, a fair question is why this is not a few more lines there. The docs draw the line by what changes. &lt;code&gt;CLAUDE.md&lt;/code&gt; is for facts that hold every session: your conventions, your architecture, the guidance you want in front of Claude every time. It loads into every session, which makes it the wrong home for anything that moves underneath it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;/prime&lt;/code&gt; is for the things that move. Where the tests live belongs in &lt;code&gt;CLAUDE.md&lt;/code&gt;; what you touched last, what is uncommitted, the one thing on fire this morning, all of that is different by the next session, and you want it recomputed when you ask rather than hand-edited into a file.&lt;/p&gt;

&lt;p&gt;That is the real move here, and it is bigger than one command. The durable thing to save is the procedure that fetches the state. A commit list or a status line goes stale the moment you write it down; a command that runs &lt;code&gt;git&lt;/code&gt; and reads your notes fetches the current answer every time. Keep the procedure narrow, though: whatever &lt;code&gt;/prime&lt;/code&gt; injects stays in the conversation and costs tokens on every later turn, so its job is to locate the work, not to preload your whole project. If it starts turning into a project dump, stop adding files.&lt;/p&gt;

&lt;h2&gt;
  
  
  You may see the older one-file form
&lt;/h2&gt;

&lt;p&gt;If you read around, you will find the same trick written as a single &lt;code&gt;.claude/commands/prime.md&lt;/code&gt; file with no folder. That is the older shape, and it still works: a &lt;code&gt;.claude/commands/prime.md&lt;/code&gt; and a &lt;code&gt;.claude/skills/prime/SKILL.md&lt;/code&gt; both create &lt;code&gt;/prime&lt;/code&gt; and behave the same way. The skills folder is the form the docs point you to now, and the reason is room to grow. A folder holds more than the one file, so it carries supporting scripts when your &lt;code&gt;/prime&lt;/code&gt; gets ambitious, and a skill can let Claude reach for it on its own when it fits (add &lt;code&gt;disable-model-invocation: true&lt;/code&gt; to keep it strictly manual, &lt;code&gt;/prime&lt;/code&gt; only). If you already have a one-file command, leave it alone, it keeps working. Reach for the folder when the command outgrows a single file, and if you ever keep both under one name, the skill wins.&lt;/p&gt;

&lt;h2&gt;
  
  
  When it does not fire
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It is not in the &lt;code&gt;/&lt;/code&gt; list.&lt;/strong&gt; You are probably not in the project, or the file is not at &lt;code&gt;.claude/skills/prime/SKILL.md&lt;/code&gt; (the folder has to be named &lt;code&gt;prime&lt;/code&gt; and the file &lt;code&gt;SKILL.md&lt;/code&gt;). A skill in &lt;code&gt;~/.claude/skills/&lt;/code&gt; instead works in every project.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It runs, but Claude never reaches for it on its own and the menu shows no hint.&lt;/strong&gt; The frontmatter did not parse. A malformed YAML block does not remove the skill: &lt;code&gt;/prime&lt;/code&gt; still works, but it loads with empty metadata, so the &lt;code&gt;description&lt;/code&gt; no longer matches. Start Claude Code with &lt;code&gt;--debug&lt;/code&gt; to see the parse error, then line up the &lt;code&gt;---&lt;/code&gt; fences.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It prompts for permission every single run.&lt;/strong&gt; Your &lt;code&gt;allowed-tools&lt;/code&gt; pattern does not match the command you inject. Line the prefix up exactly, for example &lt;code&gt;Bash(git log *)&lt;/code&gt; for a &lt;code&gt;git log&lt;/code&gt; line.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;$1&lt;/code&gt; is not what you expected.&lt;/strong&gt; In the skills model, indexed arguments are zero-based: &lt;code&gt;$0&lt;/code&gt; is the first argument, &lt;code&gt;$1&lt;/code&gt; the second. Reach for &lt;code&gt;$ARGUMENTS&lt;/code&gt; when you just want "everything I typed," and leave positions alone unless you genuinely need them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The output did not appear.&lt;/strong&gt; The injection only fires from a &lt;code&gt;!`…`&lt;/code&gt; backtick span with a leading &lt;code&gt;!&lt;/code&gt;. A command sitting in a plain fenced block is shown to the model as text, not run.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Prove it loads
&lt;/h2&gt;

&lt;p&gt;Do not take my word that the injection fired. Run &lt;code&gt;/prime&lt;/code&gt; and read the top of what comes back. If it references your actual latest commit message and your actual files, the commands ran and the state is real. If it hands you generic advice that would fit any repository, nothing injected: your &lt;code&gt;!`…`&lt;/code&gt; or &lt;code&gt;@&lt;/code&gt; lines are the place to look. A command that prints the same thing regardless of your project is just a saved sentence, which is where we started.&lt;/p&gt;

&lt;p&gt;Once it loads real state, you have turned a paragraph you retyped every session into one word, and the machine does the fetching. That is the pattern for the whole series: a thing you do by hand, moved into a tool you set up once and then just run. This one only read your project. The next rung is about letting Claude Code change it safely: a guardrail that stops the agent before it edits a file you marked off-limits. That is the next Runbook.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Sibling piece, same tool, different job: &lt;a href="https://durabilitycurve.com/blog/never-let-claude-code-tell-you-its-done/" rel="noopener noreferrer"&gt;Never Let Claude Code Tell You It's Done&lt;/a&gt; wires a test the agent cannot talk its way past.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I'm Harry, and this is The Durability Curve, where I work out where value goes as AI gets cheaper. The aim each time is to leave you with something you can take to your own work: a test to run, or a way of seeing it you did not have before. If you'd like the next one to find you, &lt;a href="https://durabilitycurve.com/?utm_source=devto" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This essay is also on the site: &lt;a href="https://durabilitycurve.com/blog/stop-re-priming-claude-code-by-hand/" rel="noopener noreferrer"&gt;durabilitycurve.com/blog/stop-re-priming-claude-code-by-hand&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Things You Used to Know by Heart: memory, thinking and AI</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 30 Sep 2026 16:23:03 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-things-you-used-to-know-by-heart-1oea</link>
      <guid>https://dev.to/harryfloyd/the-things-you-used-to-know-by-heart-1oea</guid>
      <description>&lt;p&gt;&lt;em&gt;What leaning on AI to remember for you quietly does to your own thinking.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4p9rb5qsw4d4kgkyv8a5.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4p9rb5qsw4d4kgkyv8a5.webp" alt="The facts that meet unasked are the ones you carried in the same head." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The Quiet Part is a short series on what living alongside AI quietly does to you. The others in the series are &lt;a href="https://durabilitycurve.com/blog/you-reach-before-you-think/" rel="noopener noreferrer"&gt;You Reach Before You Think&lt;/a&gt;, on judgement, and &lt;a href="https://durabilitycurve.com/blog/it-will-never-think-less-of-you/" rel="noopener noreferrer"&gt;It Will Never Think Less of You&lt;/a&gt;, on being known.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;You used to know your closest friend's phone number without reaching for it. You read a good book last month, cover to cover, and if someone asked you tonight what it actually argued, you would reach for the answer and come back empty.&lt;/p&gt;

&lt;p&gt;The information has not gone anywhere. It has left your ready recall while staying well within your reach. The number is in your phone; the argument of the book is one question to a machine away. That is what makes the loss so hard to name: nothing is missing, and something is.&lt;/p&gt;

&lt;p&gt;Because a thing you know and a thing you can find do different work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Known, and merely findable
&lt;/h2&gt;

&lt;p&gt;Ask the machine what the book argued and it hands you a whole, coherent account of the argument in a second. Search used to return documents you still had to think your own way through; the machine can hand back the thinking, already done. It is a strange kind of having, because you never did the building and you will not keep the result. You can call it up again tomorrow, always fluent, always gone the moment you close the tab: an understanding you can borrow forever and never once own.&lt;/p&gt;

&lt;p&gt;What you actually know does the one thing the borrowed version cannot. It is already present when you meet the world, deciding what you notice, what strikes you as strange, what question even occurs to you. A thing you can look up waits, filed and patient, until you already know to go and ask for it. A thing you know is in the room before you reach for anything, quietly shaping the thought while it forms.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5v80xy9x7uoquvrqslgj.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5v80xy9x7uoquvrqslgj.webp" alt="One mark filed inside a cell, another loose in the open room." width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Stored, it waits to be fetched. Carried, it is already in the room.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing you could not have looked up
&lt;/h2&gt;

&lt;p&gt;Say you know that your grandmother salted her cooking like she had a grudge against it. And say you know, too, that she grew up with nothing. For a long time those two facts never touch. Then one evening, as you stand over your own pot, they meet, and a question forms that you would never have typed: was the salt about the years of not having enough?&lt;/p&gt;

&lt;p&gt;You may never know if you are right. Either way, you would never have gone looking for the question, because you did not know there was one to ask. The machine can connect anything brilliantly once you bring it the two things and the hunch that they belong together. What it cannot be is the place where the hunch first arrives, unbidden, because you happened to be carrying both halves at once.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A machine can answer anything you bring it, and connect whatever you hand it. What it cannot do is be already in you, shaping what you notice before you know there is a question to ask.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F30m2nj642f43j7pz1a56.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F30m2nj642f43j7pz1a56.webp" alt="Two loose marks drifting until they touch and make a third." width="800" height="400"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Carried, two facts can meet before you think to introduce them.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What stays with you
&lt;/h2&gt;

&lt;p&gt;The machine is welcome to the phone numbers. Most of what you look up was never going to shape a thought, and it is right where it belongs. But somewhere in the pile is the other kind, the shape of a problem you are stuck on, or the way something you love actually works, or the argument of the book that changed your mind. Their whole use is to be in you already, uninvited, when something new walks past, and a thing you can call up is never quite that, however close to hand you keep it.&lt;/p&gt;

&lt;p&gt;When every answer is free, the scarce thing is what you have taken in so deeply that the world can still reach in and remind you of it.&lt;/p&gt;

&lt;p&gt;If an understanding can be borrowed this easily, the next question is &lt;a href="https://durabilitycurve.com/blog/what-proves-you-can-think/" rel="noopener noreferrer"&gt;what actually proves you can think&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I'm Harry, and this is The Durability Curve, where I work out where value goes as AI gets cheaper. The aim each time is to leave you with something you can take to your own work: a test to run, or a way of seeing it you did not have before. If you'd like the next one to find you, &lt;a href="https://durabilitycurve.com/?utm_source=devto" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This essay is also on the site: &lt;a href="https://durabilitycurve.com/blog/the-things-you-used-to-know-by-heart/" rel="noopener noreferrer"&gt;durabilitycurve.com/blog/the-things-you-used-to-know-by-heart&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://durabilitycurve.com/?utm_source=devto" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9k89rpog9lx7aosem3jm.webp" alt="The Durability Curve: AI is everywhere, the interesting stuff is underneath. Subscribe to get the next structural lens in your inbox." width="800" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>machinelearning</category>
      <category>productivity</category>
    </item>
    <item>
      <title>One Question for Any AI Model Headline: Can You Use It?</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sun, 27 Sep 2026 13:48:44 +0000</pubDate>
      <link>https://dev.to/harryfloyd/one-question-for-any-ai-model-headline-can-you-use-it-1d54</link>
      <guid>https://dev.to/harryfloyd/one-question-for-any-ai-model-headline-can-you-use-it-1d54</guid>
      <description>&lt;p&gt;&lt;em&gt;It may be in a mode you didn't open, on a plan you don't have, or kept for partners. Here's the check.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fem749diiewtu02mcruft.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fem749diiewtu02mcruft.webp" alt="A headline model, and the three answers your plan can give." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;OpenAI has announced GPT-6 Astra. Say you read the headline, opened ChatGPT and went looking for it. On one Plus account we checked on 25 September, the model menu in Chat mode offered GPT-5.6 Sol and GPT-5.5, two clicks in. Astra wasn't on the list. It was there, named on the menu once we switched to Work mode, and OpenAI's help page says as much: "Plus plans include GPT‑6 Astra in ChatGPT Work and Codex."&lt;/p&gt;

&lt;p&gt;That is the small version of a wider gap. The model in the headline can be somewhere in your app you didn't look, or not on your plan at all, and the headline won't say which.&lt;/p&gt;

&lt;h2&gt;
  
  
  Launches reach customers in levels
&lt;/h2&gt;

&lt;p&gt;Each lab says who gets the model further down its own page. OpenAI says Astra "will become available to all ChatGPT Plus, Pro, Business, and Enterprise users", and Free isn't listed. Google says Gemini 3.8 Flash "is available to Google AI Pro and Ultra subscribers" in the Gemini app. xAI says Grok 4.7 "is available today in Cursor and Grok Build" and through its API, and on one account we checked, every mode in the Grok app pointed to Grok 4.6.&lt;/p&gt;

&lt;p&gt;Each also keeps something for a smaller group: GPT‑6 Pro, powered by Astra, for OpenAI's Pro, Business and Enterprise plans, a Cyber version of 3.8 Flash for Google's "trusted defenders", and "invite-only access" to Grok 4.7's red-team capabilities for select xAI partners.&lt;/p&gt;

&lt;h2&gt;
  
  
  Find it on your plan, in every mode
&lt;/h2&gt;

&lt;p&gt;So before you repeat a model headline, act on it or pay for it, ask one question: can I use it? There are three answers: yours, yours at a price, or not yours.&lt;/p&gt;

&lt;p&gt;To find out, open your app and look for the model on your plan, in every mode the app has. If the menu shows modes such as "Fast" or "Expert", hover over each one, because some apps name the model only there. If it still isn't there, search the model's name plus "available" and read that sentence on the lab's newest page, which may be a help page rather than the launch post.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fisco385fmuwgcmras1mw.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fisco385fmuwgcmras1mw.webp" alt="One account each, checked 25 September 2026: ChatGPT Plus names GPT-6 Astra in Work mode; Gemini on a Plus plan has no 3.8 Flash; every Grok mode points to Grok 4.6." width="800" height="864"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One Anthropic model family shows all three answers. Anthropic says Claude Fable 5.1 and Claude Mythos 5.1 "are the same model, but with different levels of safeguards". Its help pages say that, as of September 2026, Fable 5.1 is "included as a standard part of your plan" on Max, where you can spend "up to 50% of your weekly usage limits" on it at no extra cost: yours. On Pro, it runs "on pay-as-you-go usage credits": yours at a price. Mythos 5.1, which Anthropic says has "the strongest cyber capabilities of any model we’ve released", is "available only through our trusted access programs": not yours, unless you are in one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The page says what you should get
&lt;/h2&gt;

&lt;p&gt;In 2023 the researcher Irene Solaiman described model releases as a gradient from fully closed to fully open. We'd add that plans make a smaller gradient inside the apps we use. Launch posts also go stale: Anthropic's June post said Fable 5 would be included on Pro through 22 June, and its help pages now carry the current terms. So read the lab's newest page to find out what your plan should get, then check your app to see what has actually reached you.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A model in the headline isn't yours until you find it on your plan, at your price.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Even once you find it, the model's name is only half of what you get, and I wrote about the other half in &lt;a href="https://durabilitycurve.com/blog/harness-engineering-same-model-different-product/" rel="noopener noreferrer"&gt;Same Model, Different Product&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://durabilitycurve.com/?utm_source=devto" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fabgz4t2crqceh961y3p4.webp" alt="The Durability Curve: AI is everywhere, the interesting stuff is underneath. Subscribe to get the next structural lens in your inbox." width="800" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>beginners</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Amazon Shut Out Meta's Agent: the week in AI, with six calls</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Sat, 26 Sep 2026 18:17:03 +0000</pubDate>
      <link>https://dev.to/harryfloyd/amazon-shut-out-metas-agent-shopify-gave-it-shop-pay-241g</link>
      <guid>https://dev.to/harryfloyd/amazon-shut-out-metas-agent-shopify-gave-it-shop-pay-241g</guid>
      <description>&lt;p&gt;&lt;strong&gt;The Durability Curve Weekly · Issue 1 · Saturday 19 to Friday 25 September 2026&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the first Durability Curve Weekly. Each week I pick the few AI stories, and the businesses around them, that I think will matter after the next launch. For each one I make a call you can check later, and I put a number on how sure I am. When a call’s date arrives I’ll grade it here, misses included, so you can see whether my 70%s come true about 70% of the time. Every story links to its primary source, and where I rely on press reports I say so.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkb24zrvtt8g35gr74ne.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxkb24zrvtt8g35gr74ne.webp" alt="The week's calls: how sure I am of each, and when each one resolves" width="800" height="584"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Amazon shut Meta’s agent out. Shopify gave it Shop Pay.
&lt;/h2&gt;

&lt;p&gt;At Connect on 23 September, &lt;a href="https://www.meta.com/blog/meta-connect-2026-everything-we-announced/" rel="noopener noreferrer"&gt;Meta said&lt;/a&gt; Muse, the personal agent it launched on 8 September, can now drive any app on your Mac with your permission. Muse is also getting its own email address, new retailers including Walmart and Best Buy, and a place on Meta’s AI glasses “in the coming months”. In the keynote, &lt;a href="https://techcrunch.com/2026/09/23/everything-new-coming-to-metas-ai-agent-muse/" rel="noopener noreferrer"&gt;Zuckerberg said&lt;/a&gt; Meta expects “over time” to “profit by taking a small fee from transactions.”&lt;/p&gt;

&lt;p&gt;Amazon had already shut Muse out of its store from the night of Sunday 20 September. It told reporters it never agreed to let Muse in, and that the agent doesn’t identify itself and appears to capture and store customer credentials. Meta says credentials go into &lt;a href="https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/" rel="noopener noreferrer"&gt;secure storage&lt;/a&gt; that Muse can use without seeing them. Amazon hasn’t published its statement itself; &lt;a href="https://www.geekwire.com/2026/amazon-blocks-metas-muse-ai-assistant-in-new-standoff-over-agentic-shopping/" rel="noopener noreferrer"&gt;GeekWire&lt;/a&gt; and others carry it. The next day Shopify went the other way. Muse has had access to the entire Shopify catalogue since launch, &lt;a href="https://about.fb.com/news/2026/09/the-biggest-news-from-connect-2026/" rel="noopener noreferrer"&gt;Meta says&lt;/a&gt;, and on Monday 21 September Shopify’s chief executive, Tobi Lütke, &lt;a href="https://x.com/tobi/status/2102090718546198790" rel="noopener noreferrer"&gt;announced&lt;/a&gt; a partnership to enable checkout with Shop Pay inside Muse on every Shopify store. Meta hasn’t given an exact usage figure (Zuckerberg &lt;a href="https://www.platformer.news/meta-connect-2026-muse-vr-glasses/" rel="noopener noreferrer"&gt;told Connect&lt;/a&gt; that “millions” had tried Muse), and the download counts in circulation are third-party estimates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it last? Yes.&lt;/strong&gt; The fight over who controls checkout will outlast this week. An agent that buys for you will be welcome in some stores and blocked in others, and the small fee Zuckerberg mentioned is what makes checkout worth fighting over. Amazon’s objections, that it never agreed, that an agent should say what it is, and that it shouldn’t hold your login, look like the terms other large stores will set. It is the difference I wrote about in &lt;a href="https://durabilitycurve.com/blog/access-is-not-agency/" rel="noopener noreferrer"&gt;Access Is Not Agency&lt;/a&gt;: an agent that can see a store is a long way from one that is allowed to buy there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call:&lt;/strong&gt; by 31 March 2027, another of the ten largest US online retailers will announce a new block on an AI shopping agent, or new rules requiring agents to identify themselves. Rules already in place, such as eBay’s ban on buy-for-me agents, don’t count. &lt;strong&gt;60%.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  An agent reportedly got round a government portal’s blocks, and the government heard twelve weeks later
&lt;/h2&gt;

&lt;p&gt;Speaking in New York on 23 September (the 24th in Australia), Australia’s prime minister, Anthony Albanese, &lt;a href="https://www.pm.gov.au/media/press-conference-new-york" rel="noopener noreferrer"&gt;said&lt;/a&gt; that on 18 June an OpenAI research agent, looking into public medicine spending, found a way round repeated blocks on the Medicare Statistics Reporting portal. He said it accessed public and non-public files, and that Services Australia advises it also wrote files to an internal server. OpenAI’s first notice came on 10 September, as an email to a public mailbox; OpenAI had spotted the activity in August, &lt;a href="https://www.abc.net.au/news/2026-09-24/what-we-know-about-the-openai-medicare-hack/107189452" rel="noopener noreferrer"&gt;the ABC reports&lt;/a&gt;. No personal information is believed to have been accessed. Albanese announced a taskforce and named three other health and statistics systems that may be affected, though his deputy, Richard Marles, &lt;a href="https://www.abc.net.au/news/2026-09-26/openai-review-rogue-agents-australia-medicare-hack/107199074" rel="noopener noreferrer"&gt;later called&lt;/a&gt; the agents’ activity on those “entirely normal”. OpenAI told reporters its models “took actions we did not intend”. &lt;a href="https://therecord.media/openai-australia-breach-cyber" rel="noopener noreferrer"&gt;The Record&lt;/a&gt; found that archived copies of the portal’s own code pointed the statistics service to an unauthenticated endpoint, so it’s unclear how much of a barrier the agent actually got round. The same report cites researchers at Transluce who say agent swarms they link to OpenAI probed other sites in May and June, including the Australian Institute of Health and Welfare, with techniques such as SQL injection. OpenAI says much of that activity overlaps with cases in its “ongoing review of misaligned model activity”.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it last? Yes.&lt;/strong&gt; The technique will be patched. The disclosure question will outlast it: Albanese said the incident will inform Australia’s AI standards legislation, and twelve weeks from incident to a public inbox, several of them after OpenAI knew, is the example regulators will reach for. If you run agents, the check from &lt;a href="https://durabilitycurve.com/blog/the-guardrail-your-agent-can-reach/" rel="noopener noreferrer"&gt;The Guardrail Your Agent Can Reach&lt;/a&gt; applies: can your agent reach, or change, whatever enforces the limits you set it?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call:&lt;/strong&gt; Australia’s AI standards bill, when it is published, will set a deadline for AI companies to report incidents in which their systems get into computer systems without permission. If no bill is published by 30 June 2027, the call fails. &lt;strong&gt;55%.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Anthropic committed $11.6 billion to Akamai, and got the right to buy up to 5% of it
&lt;/h2&gt;

&lt;p&gt;Akamai’s &lt;a href="https://www.sec.gov/Archives/edgar/data/1086222/000119312526401048/d288154dex991.htm" rel="noopener noreferrer"&gt;24 September filing&lt;/a&gt; says Anthropic committed about $11.6 billion over seven years for cloud capacity for its CPU workloads, with room for up to $9 billion more. Akamai expects to spend about $5.5 billion building it, and to raise this year’s capital spending by about $1.7 billion to buy parts, including memory, in advance. It issued Anthropic a &lt;a href="https://www.sec.gov/Archives/edgar/data/1086222/000119312526401048/d288154d8k.htm" rel="noopener noreferrer"&gt;warrant&lt;/a&gt; for non-voting preferred stock equal to up to about 5% of Akamai’s shares, at $111.33 a common share: about 2% vests on Anthropic’s first payment under the deal, and the rest in 1% steps for each further $3 billion it commits. The filing landed after the close on 24 September; Akamai’s shares opened Friday 13% higher and closed up 3.2%.&lt;/p&gt;

&lt;p&gt;Warrants like this are now routine. By my count from SEC filings, Akamai’s is at least the tenth in twelve months in which a supplier gave a big customer the right to buy its stock in return for purchases. AMD gave two, &lt;a href="https://www.sec.gov/Archives/edgar/data/2488/000119312525230895/d28189d8k.htm" rel="noopener noreferrer"&gt;to OpenAI&lt;/a&gt; and in February &lt;a href="https://www.sec.gov/Archives/edgar/data/2488/000000248826000045/amd-20260223.htm" rel="noopener noreferrer"&gt;to Meta&lt;/a&gt;, each for up to 160 million shares at one cent a share, vesting as the buyer takes AMD’s GPUs and as AMD’s share price hits set targets. Amazon got four, from &lt;a href="https://www.sec.gov/Archives/edgar/data/1736297/000110465926012606/tm265461d1_8k.htm" rel="noopener noreferrer"&gt;Astera Labs&lt;/a&gt;, &lt;a href="https://www.sec.gov/Archives/edgar/data/1177394/000119312526255599/d126897d8k.htm" rel="noopener noreferrer"&gt;TD Synnex&lt;/a&gt;, &lt;a href="https://www.sec.gov/Archives/edgar/data/804328/000110465926105718/tm2623289d1_8k.htm" rel="noopener noreferrer"&gt;Qualcomm&lt;/a&gt; and &lt;a href="https://www.sec.gov/Archives/edgar/data/1474735/000143774926030550/gnrc20260915_8k.htm" rel="noopener noreferrer"&gt;Generac&lt;/a&gt;. &lt;a href="https://www.sec.gov/Archives/edgar/data/1580808/000158080826000041/aten-20260803.htm" rel="noopener noreferrer"&gt;A10 Networks gave Microsoft one&lt;/a&gt; at one cent a share, &lt;a href="https://www.sec.gov/Archives/edgar/data/1835632/000119312526356217/d412696d8k.htm" rel="noopener noreferrer"&gt;Marvell gave Google one&lt;/a&gt; in August, and &lt;a href="https://www.sec.gov/Archives/edgar/data/1830081/000121390026092801/ea0303131-8k_rumgroup.htm" rel="noopener noreferrer"&gt;RUM Group&lt;/a&gt;, the company behind Rumble, agreed one at one cent a share with an unnamed cloud customer. At Friday’s close, each of AMD’s warrants would be worth about $100 billion on paper if it fully vested; Akamai’s, about $20 million.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fquoakhsh733b7dl71f4s.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fquoakhsh733b7dl71f4s.webp" alt="Ten supplier warrants, each valued at Friday's share price as if fully vested" width="800" height="695"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Buyers are reaching for protection too. The Akamai filing lets Anthropic end a project plan after a material outage. The same day, &lt;a href="https://www.bloomberg.com/news/articles/2026-09-24/oracle-cites-force-majeure-to-shield-itself-on-controversial-data-center" rel="noopener noreferrer"&gt;Bloomberg reported&lt;/a&gt; that Oracle had sent a force-majeure notice on Project Jupiter, a Stargate data-centre campus in New Mexico where it is the tenant: if both sides agree a power problem is to blame, Oracle could push back the start of its full rent by three years should the site miss its 2028 date. Oracle said such notices “are commonplace” and that the project remains on its planned schedule. Oracle’s notice only buys time: a source &lt;a href="https://finance.yahoo.com/technology/ai/articles/oracle-triggers-force-majeure-data-190718335.html" rel="noopener noreferrer"&gt;told Reuters&lt;/a&gt; that securing power is Oracle’s job under the lease, that it cannot terminate, and that the landlord still gets the full rent, only later. Anthropic’s outage exit is the clearer case of a buyer pushing risk back onto whoever builds the capacity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it last? Yes.&lt;/strong&gt; It shows who has the upper hand in these deals: the few buyers big enough to commit billions, whom suppliers will pay in stock to sign. I’ll keep the table running as new ones are filed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call:&lt;/strong&gt; between now and 31 March 2027, at least five more US-listed suppliers will disclose a warrant like these, issued to a customer in connection with chip, cloud or data-centre purchases. &lt;strong&gt;70%.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  In Anthropic’s agent market, knowing you mattered more than bargaining
&lt;/h2&gt;

&lt;p&gt;In &lt;a href="https://www.anthropic.com/research/project-swap" rel="noopener noreferrer"&gt;Project Swap&lt;/a&gt; (24 September), 201 Anthropic staff sent Claude agents to trade books. After a five-minute chat, an agent’s ranking of books matched its person’s on 61% of pairs, where guessing would get 50%. Of the value left on the table, 85% came from misreading the person and 15% from the trading. Measured against each agent’s own rankings, which model it ran on mattered more than its instructions. Participants said they’d trust an agent with about 30% of their yearly book budget, against about 40% for a well-read friend.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it last? Too early.&lt;/strong&gt; It’s one company’s staff trading books with well-behaved agents. Still, before you let an agent act for you, check whether it ranks things the way you would. That’s where most of the loss was.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call:&lt;/strong&gt; by 30 September 2027, an independent study where real money changes hands will find the same split, with misreading people’s preferences costing more than the trading itself. &lt;strong&gt;35%.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Two labs cut prices on the same day
&lt;/h2&gt;

&lt;p&gt;Anthropic released &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;Claude Opus 5.5&lt;/a&gt; on 22 September at $4 per million input tokens and $20 per million output, 20% below Opus 5; Anthropic says typical workloads cost 40% less because the model also uses fewer tokens per task. The same day OpenAI released &lt;a href="https://openai.com/index/introducing-gpt-6-sol-and-luna/" rel="noopener noreferrer"&gt;GPT-6 Sol and Luna&lt;/a&gt; at half or less of what GPT-5.6 cost the day before: $2 and $10 for Sol, $0.10 and $0.50 for Luna. Both labs attribute the cut to cheaper serving.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it last? Yes.&lt;/strong&gt; I expect the prices to hold: the labs tied the cut to their own costs, and OpenAI &lt;a href="https://venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more" rel="noopener noreferrer"&gt;told reporters&lt;/a&gt; the new prices are permanent. The number to watch is cost per finished task, and Anthropic’s 40% is its own estimate of exactly that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call:&lt;/strong&gt; by 30 September 2027, neither lab will raise the list price of Opus 5.5 or GPT-6 Sol, or release a new Opus or Sol model at a higher input or output price per token. &lt;strong&gt;75%.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Watermarks became a default
&lt;/h2&gt;

&lt;p&gt;Anthropic says Opus 5.5, like its Fable 5.1, “comes with our watermarking measures to comply with the EU AI Act”, and its &lt;a href="https://www.anthropic.com/news/claude-text-watermark" rel="noopener noreferrer"&gt;watermark page&lt;/a&gt; says future Claude models will carry it too, with older ones added over the coming months. Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/" rel="noopener noreferrer"&gt;Gemini 3.8 text-to-speech models&lt;/a&gt; (23 September) stamp every clip with a SynthID watermark, and Meta’s &lt;a href="https://research.meta.ai/blog/bringing-your-muse-to-life" rel="noopener noreferrer"&gt;Muse Realtime Avatar&lt;/a&gt; embeds an invisible Video Seal watermark in the video it generates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will it last? Yes.&lt;/strong&gt; Once the big labs watermark by default, whether Claude was involved in writing something can be checked against a key, where a detector could only guess from style. The answer is still a likelihood: it is weaker on short or factual passages, and a full rewrite removes it. For Claude’s text you can’t run that check yourself yet: Anthropic’s detector is a private preview for eligible organisations. Google already lets anyone signed in to the Gemini app check images, audio and video for SynthID.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The call:&lt;/strong&gt; by 30 June 2027, anyone will be able to check a piece of text for Claude’s watermark. &lt;strong&gt;40%.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Loud but won’t last
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Talking avatars.&lt;/strong&gt; Meta’s Muse Realtime Avatar (23 September) and Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-with-live-avatar/" rel="noopener noreferrer"&gt;Gemini 3.8 Live Avatar&lt;/a&gt; (24 September, launched in Gemini Enterprise) give agents a face. The watermark underneath matters more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Muse Charm.&lt;/strong&gt; A pocket device for talking to Muse. Zuckerberg &lt;a href="https://techcrunch.com/2026/09/23/meta-made-a-tamagotchi-like-wearable-for-its-muse-ai-agent/" rel="noopener noreferrer"&gt;said&lt;/a&gt; it should ship in December, but Meta has given no price, and its own post says only “more to share later this year”.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;“Half as many mistakes.”&lt;/strong&gt; OpenAI says Sol makes about half as many mistakes as its predecessor on its internal factuality test. The test is built from conversations where users had flagged errors, which OpenAI itself says are “not representative of typical usage”.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vendor benchmark tables.&lt;/strong&gt; OpenAI’s launch post compares Sol with Claude Opus 5, which Opus 5.5 replaced that day. Anthropic’s compares Opus 5.5 with GPT-5.6 Sol, which OpenAI replaced that day. Within hours, part of each table was out of date.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The tape
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6x8sfkgajkwszpln6yu.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6x8sfkgajkwszpln6yu.webp" alt="The week in five stocks: two months of daily candles, with this week replayed day by day" width="720" height="880"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Friday close against the Friday before, from &lt;a href="https://www.nasdaq.com/market-activity/stocks/meta/historical" rel="noopener noreferrer"&gt;Nasdaq’s price history&lt;/a&gt;: Meta +12.9%, Shopify +10.7%, Akamai +9.0%, Micron +6.5%, Oracle −7.1%. QQQ, the fund that tracks the Nasdaq-100, rose 3.2% over the same week. Akamai’s trading volume averaged 3.3 times its normal level for the week, and ten times on Friday. For most of these moves I haven’t found a primary source that explains them, so I’m not offering a reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next week
&lt;/h2&gt;

&lt;p&gt;Micron &lt;a href="https://investors.micron.com/news/press-release/2026/Micron-Technology-to-Report-Fiscal-Fourth-Quarter-Results-on-September-30-2026/default.aspx" rel="noopener noreferrer"&gt;reports on 30 September&lt;/a&gt;. Listen for how much memory AI buyers are ordering in advance; Akamai’s filing says it is doing exactly that. Sonnet and Haiku 5.5 are due “in the coming weeks”. From next week, each issue ends with a running score: how many calls are open, how many have resolved, and how often I was right. Each call gets graded here when its date comes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1jc4qira53uh97zstbcv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1jc4qira53uh97zstbcv.webp" alt="The Durability Curve Weekly: every week, the stories that matter in AI and markets, on one screen." width="760" height="266"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>agents</category>
      <category>devops</category>
    </item>
    <item>
      <title>Use AI to Learn Without Getting Worse at It</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Fri, 25 Sep 2026 13:38:48 +0000</pubDate>
      <link>https://dev.to/harryfloyd/use-ai-to-learn-without-getting-worse-at-it-17d5</link>
      <guid>https://dev.to/harryfloyd/use-ai-to-learn-without-getting-worse-at-it-17d5</guid>
      <description>&lt;p&gt;&lt;em&gt;Answer-giving AI left learners worse off once it was gone. Three setups built to keep the thinking with you, and a way to check what stuck.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;If you have used ChatGPT this month to get through something you are learning (a spreadsheet formula, a new codebase, the rules of a new job), it probably felt like progress. The answer was always one message away.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.pnas.org/doi/10.1073/pnas.2422633122" rel="noopener noreferrer"&gt;A team of researchers tested it&lt;/a&gt;. Nearly 1,000 Turkish high-school students practised maths, some of them with a ChatGPT-style assistant beside them. That group's practice scores rose 48%. Then the assistant was taken away for the exam, and they scored 17% lower than students who had never had it. The analysis the authors registered in advance puts the drop about a third smaller, but it is still a drop. In the chat logs, students had often asked for the answer and copied it.&lt;/p&gt;

&lt;p&gt;Easier practice is not evidence that you learned the thing. The effort you skip is often &lt;a href="https://harryfloyd.substack.com/p/difficulty-was-making-you" rel="noopener noreferrer"&gt;the part that was teaching you&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;When you only need the answer, take it. This is about the times you need the skill, because you will have to do it again with nobody to ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  Experienced programmers too
&lt;/h2&gt;

&lt;p&gt;The Turkish students were teenagers. In January, Anthropic &lt;a href="https://arxiv.org/abs/2601.20245" rel="noopener noreferrer"&gt;published a trial&lt;/a&gt; with 52 Python programmers, recruited through a crowd-work platform, learning a Python library that was new to them. More than half had seven or more years of coding experience. Half of them worked with an AI assistant; afterwards everyone took a quiz with no AI allowed.&lt;/p&gt;

&lt;p&gt;The AI group scored 4.15 points lower on the 27-point quiz, roughly 15 percentage points, and they were not significantly faster. They learned less and got little or no time back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who did the thinking
&lt;/h2&gt;

&lt;p&gt;Some of those programmers scored well anyway, and they had used the assistant differently. Those who handed it the work, leaned on it more as the task went on, or had it debug for them averaged 24% to 39% on the quiz. Those who asked it conceptual questions averaged 65%, about the same as the group with no AI. The best result, 86%, came from two people who let it write the code and then questioned it until they understood what it had written. The groups were small and sorted after the fact, so treat them as a clue.&lt;/p&gt;

&lt;p&gt;A separate team found the same split. In &lt;a href="https://arxiv.org/abs/2409.09047" rel="noopener noreferrer"&gt;pre-registered experiments&lt;/a&gt; with German university students learning to code, the researchers found no effect of a chatbot on learning overall. Inside that average, students who used it to generate solutions covered more topics but understood less, and students who used it to ask for explanations understood more. That split was exploratory. In the same study, 42% of requests for a solution came before the student had tried at all.&lt;/p&gt;

&lt;p&gt;Only the Turkish trial tested the design directly. Its second AI group had an assistant told to give hints, with the correct solutions loaded. Their practice scores rose 127% and the exam damage disappeared, though they did no better than students with no AI. Two other trials show well-designed tutors teaching well without isolating who did the thinking. &lt;a href="https://www.nature.com/articles/s41598-025-97652-6" rel="noopener noreferrer"&gt;At Harvard&lt;/a&gt;, 194 physics students each learned one lesson from a tutor told to encourage a first try and release the solution one step at a time, and another in an active-learning class; median gains with the tutor were more than double the class's, though it answered whenever a student demanded it. In Nigeria, &lt;a href="https://documents1.worldbank.org/curated/en/099548105192529324/pdf/IDU-c09f40d8-9ff8-42dc-b315-591157499be7.pdf" rel="noopener noreferrer"&gt;teacher-guided sessions with a chatbot&lt;/a&gt; raised English scores 0.238 standard deviations against students who got no programme at all.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdurabilitycurve.com%2F_astro%2Ffig1-evidence.D8IuC4NX_Z1BgSSt.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdurabilitycurve.com%2F_astro%2Ffig1-evidence.D8IuC4NX_Z1BgSSt.webp" alt="Who did the work, and what happened" width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Not every result fits. &lt;a href="https://storage.googleapis.com/deepmind-media/LearnLM/learnLM_sierraleone_may26.pdf" rel="noopener noreferrer"&gt;Google's trial&lt;/a&gt; of Gemini's Guided Learning raised maths scores overall, but its younger year group did slightly worse. The German study found no overall effect either way.&lt;/p&gt;

&lt;p&gt;The consistent part is narrow and useful. In two randomised trials, an AI assistant used freely during practice left people worse off once it was gone. In both studies that looked at how people used it, those who asked for explanations did better than those who asked for answers. So keep the thinking on your side of the chat: try before you ask, and ask it to explain. Three ways to set that up follow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 1: switch on the learning mode you already have
&lt;/h2&gt;

&lt;p&gt;As of September 2026, according to each company's help pages and apps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;ChatGPT:&lt;/strong&gt; type &lt;code&gt;@study&lt;/code&gt;, or press &lt;code&gt;+&lt;/code&gt; and choose &lt;code&gt;Study&lt;/code&gt;. It is on every plan, though it does not work inside Projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini:&lt;/strong&gt; &lt;code&gt;Add files&lt;/code&gt;, then &lt;code&gt;More tools&lt;/code&gt;, then &lt;code&gt;Guided Learning&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude:&lt;/strong&gt; first switch on &lt;code&gt;Code execution and file creation&lt;/code&gt; under &lt;code&gt;Settings&lt;/code&gt;, then &lt;code&gt;Capabilities&lt;/code&gt;. Then add Anthropic's &lt;code&gt;learn&lt;/code&gt; skill from &lt;code&gt;Customize&lt;/code&gt;, then &lt;code&gt;Skills&lt;/code&gt;, then &lt;code&gt;Discover&lt;/code&gt;. It stays out of coding tasks by design; for code, see Level 2.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Copilot:&lt;/strong&gt; open &lt;code&gt;Quick response&lt;/code&gt; under the prompt and choose &lt;code&gt;Study and learn&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code:&lt;/strong&gt; type &lt;code&gt;/output-style learning&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then one habit: before you send a question, write your own answer, even a bad one.&lt;/p&gt;

&lt;p&gt;Treat the modes as a start. OpenAI's &lt;a href="https://help.openai.com/en/articles/11780217-using-study-mode-in-chatgpt" rel="noopener noreferrer"&gt;own help page&lt;/a&gt; says Study mode may still give a direct answer, and the companies' own trials are early. In OpenAI's &lt;a href="https://openai.com/index/understanding-ai-and-learning-outcomes/" rel="noopener noreferrer"&gt;one-session study&lt;/a&gt;, students using Study mode scored roughly 15% higher than with ordinary online resources in microeconomics, and no differently in neuroscience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 2: paste a tutor prompt
&lt;/h2&gt;

&lt;p&gt;The modes are not everywhere, and they slip. A prompt goes anywhere. Paste this at the start of any chat where you are learning, and fill in the last two lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Be my tutor for this. I'm trying to learn it, not just get it done, so help me work it out myself.

1. Don't give me the answer or write the solution for me. Give me one hint or one step at a time, then wait for my reply.
2. If I ask you to just tell me, ask me to try first. Count every wrong answer or guess as a try. After my second wrong try, give me the answer and explain the step I missed.
3. When I get something right, confirm it and ask me why it works.
4. Keep every reply to a few sentences.
5. When we finish, give me a similar problem to try on my own.
6. If you're not sure something is correct, say so.

What I already know: [e.g. basic percentages, nothing about probability]
The problem: [paste it here]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It borrows Harvard's try-first, one-step shape and the Nigeria exercises' hint-then-answer rule, and it is stricter than Harvard's tutor, which answered on demand.&lt;/p&gt;

&lt;p&gt;I ran it eight times with a scripted learner: seven across Claude, Gemini and GPT, one in the ChatGPT app. Each time the learner demanded the answer, the tutor asked for an attempt, and after the second wrong try it gave the answer and named the missed step. Once, asked for fixed code, GPT held back the code but described the fix in words, which gave most of it away.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqurd97awmgqqgbd75gya.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqurd97awmgqqgbd75gya.webp" alt="ChatGPT holding the answer back after " width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On a long problem some models count tries per step, so the answer arrives a step at a time; that is fine. If a hint gives the answer away, reply "smaller hint".&lt;/p&gt;

&lt;p&gt;This is the same move as &lt;a href="https://harryfloyd.substack.com/p/ask-for-the-step-before-the-answer" rel="noopener noreferrer"&gt;asking for the step before the answer&lt;/a&gt;, kept on for a whole session.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you are learning code,&lt;/strong&gt; use &lt;code&gt;/output-style learning&lt;/code&gt; in Claude Code, or paste the prompt with the function or error as "the problem". When it does write code, do what Anthropic's top-scoring pair did: ask why each part is there until you could write it yourself. Then, my suggestion: close the chat and write it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Level 3: give it your material, then test yourself cold
&lt;/h2&gt;

&lt;p&gt;Levels 1 and 2 set up the tool. This level adds your own material, and it measures you.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Create a project for the thing you are learning.&lt;/strong&gt; In ChatGPT, &lt;code&gt;New project&lt;/code&gt; in the sidebar. In Claude, &lt;code&gt;Projects&lt;/code&gt;, then &lt;code&gt;New Project&lt;/code&gt; (the free plan allows five). In Gemini, a Gem with your files added.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Paste these instructions&lt;/strong&gt; into the project's instructions:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are my tutor for the material in this project. I'm here to learn it, not to look things up, so make me work for every answer.

1. Before you answer anything about the material, check the files in this project. Use them as the source of truth. If they don't cover it, say so before you answer from general knowledge.
2. Whenever I ask a question about the material, don't answer it yet, even if I say "just tell me". First ask me what I think, or give me one hint.
3. Count every wrong answer or guess as a try. After my second wrong try, give me the answer from the files and explain the step I missed.
4. When I get something right, confirm it and ask me why it works.
5. Keep every reply to a few sentences.
6. If you're not sure something is correct, say so.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Upload your material:&lt;/strong&gt; course notes, documentation, your team's handbook if you are allowed to put it into that service, and an answer key if you have one. The Harvard team loaded every question with its worked answer; without your material, a tutor falls back on what it already knows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the setup before you trust it.&lt;/strong&gt; Ask a question about the material and add "just tell me". If it answers straight away, the setup is not holding. My first version failed exactly this test in the ChatGPT app, answering at once from general knowledge. The version above passed with a hint drawn from the file.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Learn in chats inside the project.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A few days later, run the cold check.&lt;/strong&gt; Open a new chat in the project, paste this, and fill in the topics you studied:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cold check. Quiz me on what I've been learning from the files in this project.
Only ask about these topics: [list what you studied]
Ask 6 questions, one at a time. Mix "what is" questions with "what would you do if" questions.
Don't give hints and don't tell me whether I'm right until all 6 are done.
Then score me out of 6, show the correct answer from the files for each one I missed, and list what I should review.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Answer from memory, with no notes. In the ChatGPT app, with a planted mix of right and wrong answers, every question stayed on topic, it gave no feedback until the sixth, and it scored 3 out of 6, marking exactly my wrong answers and correcting each from the file. I did not test the Claude or Gemini apps themselves, so run it once with answers you know are right and wrong before relying on the score.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fup5tsv1g31iw7qzpsd4q.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fup5tsv1g31iw7qzpsd4q.webp" alt="A cold-check scorecard from my ChatGPT test run" width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The German study found students felt they had learned more than they had. A score checked against your files is a better guide than that feeling, though I tested the marking only on short factual questions, so check each marked miss against the source. Put the check in your calendar and paste the prompt yourself. Whatever you missed goes back into a tutoring chat.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the research can and cannot tell you
&lt;/h2&gt;

&lt;p&gt;None of the seven studies tested people again weeks or months later, so none shows the learning lasted. Only two went through journal peer review. The evidence on how people used the AI comes from patterns inside two studies, not from assigning them. Three trials were run by AI companies, and in Nigeria the comparison group got no programme at all.&lt;/p&gt;

&lt;p&gt;Your own cold check a few days later is the only follow-up test in this piece.&lt;/p&gt;

&lt;h2&gt;
  
  
  The checklist
&lt;/h2&gt;

&lt;p&gt;Next time you open an AI to learn something:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Decide whether you need the answer or the skill. For the skill, keep the thinking on your side.&lt;/li&gt;
&lt;li&gt;Switch on your app's learning mode, or paste the tutor prompt.&lt;/li&gt;
&lt;li&gt;Write your own answer before you ask, then ask it to explain.&lt;/li&gt;
&lt;li&gt;If a hint gives it away, ask for a smaller one.&lt;/li&gt;
&lt;li&gt;For anything you need to know properly, put your material in a project and test the setup with "just tell me".&lt;/li&gt;
&lt;li&gt;A few days later, run the cold check on the topics you studied, with no notes, and send what you missed back into a tutoring chat.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;The Nigeria exercises drew in part on &lt;a href="https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4475995" rel="noopener noreferrer"&gt;tutor prompts&lt;/a&gt; published by Ethan and Lilach Mollick in 2023.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>career</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Show It an Example: why 'make it distinctive' gets you the most ordinary output</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Mon, 21 Sep 2026 21:21:01 +0000</pubDate>
      <link>https://dev.to/harryfloyd/show-it-an-example-why-make-it-distinctive-gets-you-the-most-ordinary-output-7mp</link>
      <guid>https://dev.to/harryfloyd/show-it-an-example-why-make-it-distinctive-gets-you-the-most-ordinary-output-7mp</guid>
      <description>&lt;p&gt;&lt;em&gt;Why "make it distinctive" gets you the least distinctive thing there is, and what to paste instead.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbfu3cpqd3sshy0a2kk3u.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbfu3cpqd3sshy0a2kk3u.webp" alt="Show it the thing, do not describe the thing." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You wanted it to write something with a particular feel. A tagline, a bio, an email in the right tone, a name that sounds like you and not like everyone. So you were careful about it. You asked for punchy. Modern. Distinctive. You told it to steer well clear of the corporate stuff.&lt;/p&gt;

&lt;p&gt;And it came back polished and completely generic. The kind of line you have read a hundred times on a hundred websites. You asked for distinctive and got the opposite, so you added more words, and got a longer version of the same thing.&lt;/p&gt;

&lt;p&gt;The words were the problem. Being more specific in adjectives is not being more specific. "Distinctive" is not a target, it is a direction, and there is an enormous crowd of distinctive-ish things it could point at. Told to head in that direction, the machine has to decide what you mean, and it often lands somewhere familiar in that crowd: polished, broadly compatible with the word you gave it. That is where generic comes from. You named a direction without showing it where along that direction you wanted to go.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbzbt3czpfag7z95f5saj.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbzbt3czpfag7z95f5saj.webp" alt="Ask in adjectives and the model returns the flowery, most ordinary version of the word." width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I watched this happen. I asked for five taglines for a small bookshop, and said make them charming, memorable and distinctive. It gave me the flowery, template version, the exact voice you would guess: "trade the noise of the world for the turn of a page." Asked to be distinctive, it produced the least distinctive thing on offer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Paste an example, and say match this
&lt;/h2&gt;

&lt;p&gt;So stop describing the style, and show it one.&lt;/p&gt;

&lt;p&gt;Find something that already has the feel you want. A tagline you wish you had written, a paragraph whose voice you would happily borrow, an email whose tone you admire (keep out anything you would not want to put into the tool you are using). One is enough to start, and a few can work better. Paste it in and say: match this. That is the whole move, and it is shorter than the paragraph of adjectives you were about to write.&lt;/p&gt;

&lt;p&gt;When I dropped the adjectives and instead pasted three taglines I liked, plain and dry and lower case, and said only match these, it answered in the same plain voice. Not by repeating the three I gave it, but with new ones in the same plain, lower-case voice: "books you'll actually read." "fewer books, better ones." "just books, really." Same machine, same bookshop. What I had taken away was the adjectives; what I had added was something to aim at.&lt;/p&gt;

&lt;p&gt;The example does not have to be from the same world. A tagline you like from a coffee brand can teach it the voice for your bookshop. Think of it as the target to aim at, and you can still keep, edit or throw out whatever comes back.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhuk3ifd4ffk0q4qfmt3h.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhuk3ifd4ffk0q4qfmt3h.webp" alt="Show it an example and it finally has something concrete to aim at, in your voice." width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why an example beats a paragraph
&lt;/h2&gt;

&lt;p&gt;An example is a thing it can see and follow. An adjective is only a word, and a word covers everything the word could mean. "Warm" spans a birthday card and a hostage negotiation. Show it one line that is warm in the way you mean, and suddenly "warm" has an address.&lt;/p&gt;

&lt;p&gt;That is the difference between pointing at a target and naming a direction. It can hit a target you show it. Give it only a direction and it often lands somewhere familiar.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An example is a target it can see and match. An adjective is only a direction, and it fills a direction with the ordinary. Show it the thing; do not describe the thing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is also why this holds up as the tools change. It is not a trick tied to today's model. However good the next one is, it cannot see the specific thing in your head, and a description will rarely carry it across as cleanly as one real example. The surest way to hand over your specific thing is to show it an example of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it catches you out
&lt;/h2&gt;

&lt;p&gt;It will follow your example closely, and that includes a bad one. Grab the first thing to hand because it was convenient, and you will get a clean, confident version of something you did not actually want, which is harder to spot than an obviously generic miss.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayjsmgh81ylrxgyoh951.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fayjsmgh81ylrxgyoh951.webp" alt="Show it a tired example and it returns a polished version of the wrong thing." width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So before you paste it, read the example back and ask whether it is really the thing you want, not just the thing you found. The move carries your taste across to the machine. It does not give you taste, and it will follow a lazy choice as readily as a good one. When the thing you are trying to get right is bigger than a tagline, the same rule holds: a better model still needs to know what your version of good looks like. I wrote about &lt;a href="https://harryfloyd.substack.com/p/harness-engineering-same-model-different-product" rel="noopener noreferrer"&gt;why the setup you give a model beats the model itself&lt;/a&gt; for anyone whose version of this has moved past writing taglines.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>beginners</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Survivorship bias: why your evidence is only what survived</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Thu, 10 Sep 2026 08:38:43 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-planes-that-didnt-come-back-2459</link>
      <guid>https://dev.to/harryfloyd/the-planes-that-didnt-come-back-2459</guid>
      <description>&lt;p&gt;&lt;em&gt;The Blueprint · No. 5. Start with No. 1: &lt;a href="https://harryfloyd.substack.com/p/the-setting-you-never-changed" rel="noopener noreferrer"&gt;The Setting You Never Changed&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In the Second World War, the American military had a problem with its bombers. Too many were being shot down, and the obvious remedy was to fit armour to them. But armour is heavy, a plane can carry only so much, and cover it everywhere and it will barely leave the ground. So the real question was where. Which parts of the plane most needed the protection.&lt;/p&gt;

&lt;p&gt;They had data to answer it. Returning bombers were inspected, and the engineers could see where the hits had landed. In the telling that became famous they clustered in a familiar pattern: heaviest along the fuselage and the wings, the engines coming back comparatively clean. Reinforce the parts taking all the fire, the reasoning went, and you will save the most planes. It is a hard argument to fault. You have the data, the damage is right there in front of you, and you are only following it.&lt;/p&gt;

&lt;p&gt;A statistician named Abraham Wald looked at the same figures and reached the opposite conclusion. The armour belonged where the holes were not. Protect the engines, the clean areas, the places nobody thought to patch.&lt;/p&gt;

&lt;p&gt;The reason, once you hear it, rearranges something in your head and does not put it back. The engineers were studying the planes that came back, because those were the only planes they had. A bomber covered in holes across its wings and fuselage was a bomber that had been hit in those places and had still flown home, which meant those were exactly the spots where a plane could take a beating and survive.&lt;/p&gt;

&lt;p&gt;The clean areas were clean for a reason the data could not show. The planes that were hit there did not return to be inspected. They were somewhere in the sea. Those absent holes marked the wounds that killed.&lt;/p&gt;

&lt;p&gt;That is the whole trap, and it is worth seeing plainly, because it has nothing to do with aeroplanes. Whenever you draw a lesson from a set of examples, something decided which examples reach you, and that something is rarely chance. It is a filter. Survival, success, fame, memory, simply staying in business: each of those filters does its work by removing the failures before they ever arrive.&lt;/p&gt;

&lt;p&gt;So the set in front of you is only the residue left once the filter has run, a long way from a fair slice of everything that was tried, and because the filter's whole job was to take things away, the things it took away are the ones you cannot see. They are also, very often, the ones you most need.&lt;/p&gt;

&lt;p&gt;You feel how strong this is the moment you start looking for it. Pick up any book about how some billionaire built their company and you will find a handful of habits offered up as the cause: the early mornings, the ferocious focus, the refusal to hear the word no. What you will never find, because nobody writes that book, is the far larger pile of people who rose at the same hour and focused just as fiercely and refused just as hard, and went broke regardless.&lt;/p&gt;

&lt;p&gt;If the people who failed had every one of the winning habits too, the habits cannot be the thing that set the two apart. You are being handed the survivors and asked to reverse a recipe from them, with the one ingredient that could tell you what actually mattered, the failures, quietly deleted from the page.&lt;/p&gt;

&lt;p&gt;It runs through much smaller things as well. When someone tells you they do not make things like they used to, waving at a hundred-year-old chair that is still rock solid, remember that you are looking at the one chair that lasted a century. Most of the flimsy furniture of the past broke and was thrown out generations ago. You are holding the best of the old, the sliver that survived, up against the everyday run of the new.&lt;/p&gt;

&lt;p&gt;So here is the move, and it is a single question you can put to almost any claim built on examples. What decided which cases I get to see, and what would the ones it left out have looked like?&lt;/p&gt;

&lt;p&gt;Before you copy the habits of the successful, go looking, in your imagination if nowhere else, for the people who did the very same and failed, and ask whether they shared the habit. Before you decide the old ways were better, ask what broke and disappeared long before you arrived to inspect what was left. The absence is shaped, and its shape is the part of the answer that did not survive to be seen.&lt;/p&gt;

&lt;p&gt;None of this means every set of examples is lying to you. Sometimes your sample really is a fair one, gathered without a filter quietly picking the winners, and then this particular problem does not arise. The trap springs only when the very process that produced your evidence is the same process that removed the counter-examples. The tell is easy to learn once you have it. Your data is made of survivors, of winners, of the things that lasted and the stories that got told. The moment you notice that, you know to go looking for the planes that did not come back.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Note: Wald's wartime analysis survives as eight memoranda for Columbia University's Statistical Research Group; the standard scholarly reconstruction is Marc Mangel and Francisco Samaniego, &lt;a href="https://www.tandfonline.com/doi/abs/10.1080/01621459.1984.10478038" rel="noopener noreferrer"&gt;Abraham Wald's Work on Aircraft Survivability&lt;/a&gt;, Journal of the American Statistical Association 79 (1984): 259–267. The memoranda estimate the vulnerability of each part from the damage on returning aircraft; the familiar bullet-hole diagram and the scene of engineers overruled in a briefing room are later retellings, not Wald's own.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;From &lt;a href="https://harryfloyd.substack.com" rel="noopener noreferrer"&gt;The Blueprint&lt;/a&gt;, a series on The Durability Curve about the surfaces hidden inside systems you already live in. If this changed how you read a set of examples, &lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=the-planes-that-didnt-come-back" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>analysis</category>
      <category>discuss</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Induced demand: why widening a road never ends the jam</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 09 Sep 2026 09:08:04 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-extra-lane-fills-itself-42e8</link>
      <guid>https://dev.to/harryfloyd/the-extra-lane-fills-itself-42e8</guid>
      <description>&lt;p&gt;&lt;em&gt;The Blueprint · No. 3. Start with No. 1: &lt;a href="https://harryfloyd.substack.com/p/the-setting-you-never-changed" rel="noopener noreferrer"&gt;The Setting You Never Changed&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A city looks at a motorway that crawls every rush hour and does the obvious thing. It widens the road. More lanes, more room, more cars moving at once, and for a while the traffic loosens and the journey gets quicker. Then, within a few years, the road starts filling again, the familiar crawl creeping back onto the wider road that was meant to end it. Somewhere a lot of money went into a fix that did not fix the thing it was meant to fix, and everyone quietly agrees they should have built it wider still.&lt;/p&gt;

&lt;p&gt;The intuition underneath that decision feels like common sense. Traffic is a fixed lump of cars trying to squeeze through a narrow pipe. Widen the pipe and the same lump flows more easily. If it clogs again, the lump must have grown, so widen it again. Roads as plumbing, congestion as a volume problem, more capacity as the answer.&lt;/p&gt;

&lt;p&gt;The pipe picture is wrong, and it is wrong in a way that explains the whole thing. The traffic you can see was never the whole demand. The jam itself was holding some of the rest back. Every day the road crawled, some people looked at it and chose not to be on it. They took the train instead. They shifted their trip to before the rush or after it. They bundled three errands into one, worked from home, or simply did not make the journey at all. The congestion was a wall, and behind it sat the trips it was holding back, some of which would become worth making the moment the wall came down.&lt;/p&gt;

&lt;p&gt;So you add the lane and the wall comes down. Driving gets quicker, and quicker driving is an invitation. The person who used to take the train may get back in the car. The trip that was not worth the crawl becomes worth it. Trips the old jam had quietly discouraged start returning to the road, eating into the improvement the new lane was meant to deliver. The refilling slows as the road clogs again, until the next trip that might have joined it is once more not quite worth making.&lt;/p&gt;

&lt;p&gt;That is the short loop, and it is worth saying plainly. Travel time is a price, paid in minutes rather than money, and like any price it holds demand down. Add capacity and demand can rise to take up the room, and how much depends on how much the old conditions were suppressing. Where that suppressed demand is large, the road fills until much of the improvement is gone and you have moved more cars for less benefit than the map promised. Where it is small, the extra lane stays useful. The loop runs wider over longer periods, as people change where they live and work, firms follow the new access, and trips that the old conditions ruled out entirely begin to make sense.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F92y82bu553w3uzne95wd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F92y82bu553w3uzne95wd.webp" alt="Figure 1: with a big enough hidden crowd, the loop turns until most of the gain is gone" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You do not have to take this on faith. One of the clearest findings comes from a study of US cities by &lt;a href="https://www.aeaweb.org/articles?id=10.1257/aer.101.6.2616" rel="noopener noreferrer"&gt;Duranton and Turner&lt;/a&gt;: the miles people drive rose roughly in step with the interstate lane miles built, which is a polite way of saying new roads fill themselves. &lt;a href="https://kinder.rice.edu/urbanedge/what-if-we-spent-billions-improve-access-instead-gridlock" rel="noopener noreferrer"&gt;Houston&lt;/a&gt; is the vivid illustration, a motorway widened to as many as twenty-six lanes at its broadest point, whose rush-hour journeys grew sharply longer again within a few years of the work finishing. The reverse case points the same way. Across more than seventy cases where road space was reallocated away from traffic, the traffic problems predicted for the surrounding streets were generally much smaller than expected (&lt;a href="https://nacto.org/wp-content/uploads/disappearing_traffic_cairns.pdf" rel="noopener noreferrer"&gt;Cairns, Atkins and Goodwin 2002&lt;/a&gt;). When Seoul pulled down an elevated expressway and restored the stream it had been built over, road trips fell and subway ridership rose, rather than all the displaced traffic simply reappearing elsewhere (&lt;a href="https://ideas.repec.org/a/eee/trapol/v21y2012icp165-178.html" rel="noopener noreferrer"&gt;Chung, Hwang and Bae 2012&lt;/a&gt;). Some journeys find another route, some shift to another mode, some move to another time, and some are simply no longer made, melting back behind the very wall the road had been holding down.&lt;/p&gt;

&lt;p&gt;Here is why this is worth carrying around, because it is not really about roads. The friction was doing a second job. The wait, the queue, the crawl, whatever the painful thing was, was also a filter, quietly turning away demand you never saw because it never arrived. A support team drowning in tickets hires more people and the replies get faster, and some customers who would once have given up on a small problem now bother to report it. A clinic adds appointment slots, and some patients who would have gone elsewhere, put the visit off, or never booked at all start filling them. Not every capacity increase works like this. The tell is that the old crush was already making people give up, postpone, reroute, or go without.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2qh7l9mdin83jiayv0w.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw2qh7l9mdin83jiayv0w.webp" alt="Figure 2: remove the wait and the deflected trips come back, until the wait returns" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once you can see it, the fix that everyone reaches for starts to look naive. When something is overwhelmed and the instinct is to add capacity, the first question is whether the crush is also holding demand back, and how much is waiting behind it. Where a lot is waiting, much of the added capacity can fill again, and you have bought a larger operation running at much the same strain. Then capacity alone will not get you the outcome you wanted, and you need a lever on the demand as well, whether that is a price in money, a priority, or a rule about who gets on. Shape the demand too, because added capacity gives suppressed demand somewhere to return.&lt;/p&gt;

&lt;p&gt;None of this makes capacity useless. The trap is sprung when the crush is suppressing demand that lower friction can release, and a surprising number of the queues you fight with, in traffic and far beyond it, are doing exactly that.&lt;/p&gt;

&lt;p&gt;Add room to a queue that was turning people away, and the crowd it was hiding comes back to claim the room.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;From &lt;a href="https://harryfloyd.substack.com" rel="noopener noreferrer"&gt;The Blueprint&lt;/a&gt;, a series on The Durability Curve about the surfaces hidden inside systems you already live in. If this reframed a queue you are fighting, &lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=the-extra-lane-fills-itself" rel="noopener noreferrer"&gt;subscribe&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>architecture</category>
      <category>discuss</category>
      <category>analysis</category>
    </item>
    <item>
      <title>The Skill That Never Fired: How to Test Whether Claude Actually Picks Your Skill</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 02 Sep 2026 12:17:53 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-skill-that-never-fired-how-to-test-whether-claude-actually-picks-your-skill-fae</link>
      <guid>https://dev.to/harryfloyd/the-skill-that-never-fired-how-to-test-whether-claude-actually-picks-your-skill-fae</guid>
      <description>&lt;h1&gt;
  
  
  The Skill That Never Fired
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e25b9ef-f0a1-4ecf-96fa-92559b0b0000_1520x856.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6e25b9ef-f0a1-4ecf-96fa-92559b0b0000_1520x856.webp" alt="A dark room lit by a single emerald key-light: the skill that fired stands in the light, every other skill a slab in the dark. One skill fires; you never see the rest."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A skill can fail in two ways. Its instructions can be wrong, so it does the job badly. Or Claude can decide never to load it, so the instructions never run at all. The first failure is obvious when you test the skill by name. The second only shows up when you test whether Claude chooses it on its own.&lt;/p&gt;

&lt;p&gt;The second failure is the quiet one. You write a skill, you invoke it by name to check it, and it works. Then in normal use it just sits there. Claude answers without it. Nothing errors, nothing warns you, and the skill still shows as installed. It was never wrong. It was never chosen.&lt;/p&gt;

&lt;p&gt;That choice is a routing decision, and Claude makes it by matching the request against your skill's name and description, before it reads a word of the body. The Claude Code docs say it directly: the description is what helps Claude decide when to load a skill. Anthropic tells you to test that decision, separate from the skill's output, and ships a tool that does it. Its &lt;code&gt;skill-creator&lt;/code&gt; scores one target skill over repeated runs: does this skill fire on the prompts it should, and stay quiet on the ones it should not?&lt;/p&gt;

&lt;p&gt;What that score does not tell you is what happened when another plausible skill was there too: whether the neighbour took the request, both fired, or neither did. That is the failure this piece is interested in, where your skill sits beside one that could answer it and the winner is not guaranteed to be yours. This walks through building a skill, watching that decision for yourself, and grading it against the neighbour it can lose to. You can run a first pass in about 15 minutes at a terminal.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a skill is
&lt;/h2&gt;

&lt;p&gt;At its simplest, a skill is a folder with one required file, &lt;code&gt;SKILL.md&lt;/code&gt;. It can also hold scripts and reference files that load only when needed, but the minimum is the one file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer-date&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Format&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;date&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;customer-facing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;UK&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;correspondence&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;(emails,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;letters,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;customers)&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;as&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;D&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Month&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;YYYY.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;For&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;CSV&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;exports,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;export-date."&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;Rewrite the date the user gives in UK long form, for example 30 August 2026. Reply with only the formatted date.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a normal auto-invocable skill, the name and description sit in Claude's discovery context so it can decide whether the skill is relevant. The body below the frontmatter loads only when the skill is invoked, whether Claude chooses it or you type its name. Claude sees both the name and the description, and the description is the main field Anthropic gives you for saying when the skill should run. Write it for the router, not as a note to yourself. Claude Code also accepts a &lt;code&gt;when_to_use&lt;/code&gt; field, appended to the description; these skills use only a name and a description.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256367b0-fa63-491c-8a27-817488321d5e_1560x860.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F256367b0-fa63-491c-8a27-817488321d5e_1560x860.webp" alt="A request meets two installed skills. Only each skill's name and description sit in Claude's discovery context; Claude matches the request against that metadata and picks one, loading only its body."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Three common ways to use a skill. In claude.ai, turn on code execution, open Customize then Skills, and upload the folder as a zip. In Claude Code, put the folder in &lt;code&gt;.claude/skills/&lt;/code&gt; for one project or &lt;code&gt;~/.claude/skills/&lt;/code&gt; for all of them. Through the Claude API, you upload it and reference its &lt;code&gt;skill_id&lt;/code&gt;. The core &lt;code&gt;SKILL.md&lt;/code&gt; format travels across all three, though installation differs and some frontmatter, including the &lt;code&gt;disable-model-invocation&lt;/code&gt; used later, is specific to Claude Code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching the routing decision
&lt;/h2&gt;

&lt;p&gt;The mistake to avoid is judging a skill by its output. Ask Claude to format a date and you might get &lt;code&gt;30 August 2026&lt;/code&gt; whether your skill ran or not, because the model can format a date on its own. The output tells you nothing about routing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F527f9e7f-9fc4-48e0-956c-72c84ccebae6_1560x820.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F527f9e7f-9fc4-48e0-956c-72c84ccebae6_1560x820.webp" alt="The same request produces the same output whether the skill fires or not; only the Skill tool call in the event stream reveals which happened."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You want the decision itself. In Claude Code, a skill runs through a &lt;code&gt;Skill&lt;/code&gt; tool that appears in the event stream. Run a prompt non-interactively and filter the stream down to the skill call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"Rewrite this date for the customer email: 2026-08-30"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--output-format&lt;/span&gt; stream-json &lt;span class="nt"&gt;--verbose&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'select(.type=="assistant") | .message.content[]?
           | select(.type=="tool_use" and .name=="Skill") | .input'&lt;/span&gt;

&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"skill"&lt;/span&gt;:&lt;span class="s2"&gt;"customer-date"&lt;/span&gt;,&lt;span class="s2"&gt;"args"&lt;/span&gt;:&lt;span class="s2"&gt;"2026-08-30"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the filtered result, not the raw stream, which wraps each event in more metadata. If the reader prints nothing, dump a raw event and look for a &lt;code&gt;Skill&lt;/code&gt; call by hand, because the stream's shape shifts between versions. It is the routing decision read from the tool call, not guessed from the output. &lt;code&gt;/skills&lt;/code&gt; shows which skills are available to Claude and &lt;code&gt;/context&lt;/code&gt; shows the discovery listing's context cost, but neither proves this prompt invoked one. In claude.ai there is no equivalent machine-readable event. Anthropic's guidance is to review Claude's thinking to confirm a skill loaded, which works for checking by eye but not for building the kind of record above. And the event stream shows the skills Claude actually invoked, both of them when it invokes two, which is how a &lt;code&gt;both&lt;/code&gt; shows up at all. What it does not expose is the candidate set: the other installed skills that were plausible but never invoked. You see what fired, not what it beat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Routing is a decision you can grade
&lt;/h2&gt;

&lt;p&gt;Whether a skill fires is a choice among whatever skills could plausibly answer the request. You can only grade that choice if you know what the right answer was before you run it.&lt;/p&gt;

&lt;p&gt;So I built two skills with different jobs. &lt;code&gt;customer-date&lt;/code&gt;, above, formats dates for customer emails in long form. &lt;code&gt;export-date&lt;/code&gt; formats them for CSV exports as &lt;code&gt;DD/MM/YYYY&lt;/code&gt;. Then I wrote a labelled prompt set: 4 requests that clearly want the customer skill, 4 that clearly want the export skill, and 4 date-adjacent requests that should fire neither. Every result gets one of four labels: right, wrong, none, or both.&lt;/p&gt;

&lt;p&gt;Start with the failures, because they are where the method earns its keep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ambiguous request.&lt;/strong&gt; Ask "Format this date: 2026-08-30" with both skills installed, and the results scatter: sometimes one fires, sometimes both, sometimes neither. That scatter is the expected result of an ambiguous request. The request never said whether it wanted the customer or the export format, so there is no correct answer to grade against. An ambiguous prompt is not a failed test, it is an ungradable one. If you cannot label the right skill before running it, the result cannot tell you whether Claude chose well.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Names and descriptions that draw no line.&lt;/strong&gt; I named two skills for their output format, &lt;code&gt;long-date&lt;/code&gt; and &lt;code&gt;slash-date&lt;/code&gt;, and gave them the same vague description, "Format a date." Their bodies did different things, but their discovery metadata claimed the same job, so there was no boundary for the router to use and nothing told Claude which one fits a customer request. The grades went bad in the way that matters: one customer prompt fired nothing at all, and 2 export prompts fired both skills at once. Misses and double-fires, which is why "both" has to be one of your outcome labels.&lt;/p&gt;

&lt;p&gt;Then the control, so you can see what clean looks like. Give the two skills distinct, use-case descriptions, and ask prompts whose wording matches those use cases, and routing is clean: 8 out of 8 to the right skill, and the neither-prompts correctly firing nothing. That is the model doing the keyword and intent matching you made easy for it. It is the case that should work, and it does. Note that the customer prompts contain words like "customer email" that are already in the customer skill's description. Clean routing here is a control condition, not proof that routing is robust. The evidence is in the failures above.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Routing runs on the discovery metadata Claude can see, which is the name and the description together, and my runs show either one can carry it. When the names were the vague part but the descriptions were sharp, routing was clean. When the descriptions were the vague part but the names said the use case, &lt;code&gt;customer-date&lt;/code&gt; and &lt;code&gt;export-date&lt;/code&gt;, routing was also clean, 8 out of 8 on the same prompts. That name-only run is an easy case, mind: the prompts carry the same words as the names, customer and export, so it shows a name helps when the request echoes it, not that a bare name is a strong signal on its own. It broke in one condition only: format-only names, &lt;code&gt;long-date&lt;/code&gt; and &lt;code&gt;slash-date&lt;/code&gt;, plus a shared vague description, where neither field told Claude what set the two skills apart.&lt;/p&gt;

&lt;p&gt;That broken condition is the useful one. I left the weak names alone and rewrote only the descriptions around use cases, and the mess went to 8 out of 8. A good description rescued names that carried no signal. A good name had already done the same for descriptions that carried none. What you cannot do is leave both vague and expect Claude to find the line.&lt;/p&gt;

&lt;p&gt;So write the description as a routing rule, not a summary. Put the use first. Include the words people actually type when they want this skill. Draw the boundary against the neighbour it might be confused with.&lt;/p&gt;

&lt;p&gt;Some skills should not be auto-routed at all. Anything with a side effect or a real cost is safer as a skill you invoke by name, &lt;code&gt;/customer-date&lt;/code&gt;, or one you lock with &lt;code&gt;disable-model-invocation: true&lt;/code&gt; so only a person can trigger it. For those, manual invocation is the design, not a workaround. The rule underneath: how much routing error you can accept depends on what a wrong route costs. When the cost is high, the fix is often to stop routing automatically rather than to tune the description harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure that has nothing to do with your skill
&lt;/h2&gt;

&lt;p&gt;There is one more way a skill stops firing, and no description work touches it. For skills still exposed to the model, Claude Code keeps every skill's name in the discovery listing, but the listing has a budget, around 1% of the model's context window, and once it runs over, Claude Code starts dropping descriptions, beginning with the skills you invoke least. That can strip out exactly the words that told two skills apart. So a skill whose description used to distinguish it cleanly can start missing once your catalogue grows large enough, with nobody editing it. &lt;code&gt;/doctor&lt;/code&gt; reports the listing's cost. If a skill that used to route well starts slipping, check the size of your catalogue before you rewrite the skill: prune the skills you do not use, shorten the descriptions that survive so the distinguishing words fit, set low-priority skills to &lt;code&gt;name-only&lt;/code&gt; so Claude keeps their names without their descriptions, or set rarely-used skills to &lt;code&gt;disable-model-invocation: true&lt;/code&gt;, which takes them out of the router and its listing entirely; you still invoke those with &lt;code&gt;/name&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Worth knowing before it bites a team: skills with the same name at different levels do not merge, one shadows the other. The order is enterprise, then personal, then project, so a &lt;code&gt;/deploy&lt;/code&gt; skill in your &lt;code&gt;~/.claude/skills/&lt;/code&gt; silently overrides the one your repo ships in &lt;code&gt;.claude/skills/&lt;/code&gt;. If you commit skills for a team, give them names that will not collide, and do not rely on the project copy winning. This is also where overlap arrives for people who did not build it: a marketplace pack or an inherited folder drops in a skill whose description competes with one of yours, and the first you hear of it can be a skill that used to fire and now does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which layer failed
&lt;/h2&gt;

&lt;p&gt;The check separates two layers that a vague "it didn't work" runs together. Execution failure means Claude loaded the skill and the body did the wrong thing. Selection failure means the right body never got its chance to run at all. Everything else in this piece is a kind of selection failure: a discovery miss, interference from a neighbour, two definitions that overlap, a listing truncated at scale, a same-name skill shadowing yours. Only execution failure is about the instructions. The rest is why a skill can regress with nobody touching it, and why testing the body is only half the job.&lt;/p&gt;

&lt;p&gt;The stakes climb once the skills matter. A date formatter losing to its twin costs you a wrong date format. A &lt;code&gt;code-review&lt;/code&gt; skill that loses requests to a generic "help me with this file" skill costs you the review you thought ran on every change. Whether that happens turns on the same thing as the date skills: whether the two descriptions draw a line the router can use. The installed list will not tell you, so check it directly. Ask "look at this diff" with both installed and the route can go four ways: cleanly to the review skill, to a &lt;code&gt;both&lt;/code&gt;, to the generic skill alone, or to neither. Give it requests whose correct skill you know, install it next to the neighbour you suspect, and read which one the &lt;code&gt;Skill&lt;/code&gt; tool actually calls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it yourself
&lt;/h2&gt;

&lt;p&gt;The official &lt;code&gt;skill-creator&lt;/code&gt; plugin measures a target skill's trigger rate for you. The manual version here adds the identity of the competing skill, so you can tell a miss from interference, a neighbour firing instead of the target or alongside it, reproduce a collision between two specific neighbours, and read the &lt;code&gt;Skill&lt;/code&gt; call yourself. There is a second reason to read the calls rather than trust a score: as of the &lt;a href="https://github.com/anthropics/skills/blob/main/skills/skill-creator/scripts/run_eval.py" rel="noopener noreferrer"&gt;current evaluator&lt;/a&gt; (August 2026), it reads the first tool call in a run and counts the target as not fired if anything else, a neighbour skill included, gets there first, so the interference this piece is about can quietly lower the very rate meant to catch it. Here is the whole pack. Two skills, twelve prompts, four outcome labels, and the one-line reader from earlier. It is also a download, at &lt;a href="https://durabilitycurve.com/tools/skill-routing-eval/" rel="noopener noreferrer"&gt;durabilitycurve.com/tools/skill-routing-eval&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The skills, with descriptions that draw the boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;customer-date&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Format a date for customer-facing UK correspondence (emails, letters, messages to customers) as D Month YYYY. For CSV or data exports, use export-date.&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="s"&gt;Rewrite the date the user gives in UK long form, for example 30 August 2026. Reply with only the formatted date.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;export-date&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Format a date for CSV or database exports (spreadsheets, data files) as DD/MM/YYYY. For customer emails and letters, use customer-date.&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="s"&gt;Rewrite the date the user gives in slashed form, for example 30/08/2026. Reply with only the formatted date.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The prompts, each with its known-correct skill:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;should fire customer-date:
  Rewrite this date for the customer email: 2026-08-30
  Put this date in a letter to the client: 2026-08-30
  Format the date for a message to a customer: 2026-08-30
  Tidy the date in this customer-facing note: 2026-08-30

should fire export-date:
  Format this date for the CSV export: 2026-08-30
  Put this date into the spreadsheet export: 2026-08-30
  Format the date for the database file: 2026-08-30
  Prepare this date for a data export: 2026-08-30

should fire neither (date-adjacent work these formatters should refuse):
  What is today's date?
  When did the Second World War end?
  Parse this log timestamp: 2026-08-30T14:22Z
  What day of the week is 2026-08-30?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Install both skills, run each prompt through the &lt;code&gt;jq&lt;/code&gt; reader above, and mark the result right, wrong, none, or both. Those labels roll up into three numbers worth watching. Recall: of the requests that should fire a skill, how many did. False triggers: of the requests that should not, how many fired it anyway. Interference: with a neighbour installed, how often that neighbour fires on a request meant for this skill, either instead of it or alongside it. Recall and false triggers are what the standard trigger-rate test measures for one skill, from its positive and negative cases. Interference is the number it cannot give you, because it only records whether the target fired, not which competing skill fired instead or alongside it. Score a &lt;code&gt;both&lt;/code&gt; as a hit on recall and on interference at once: the intended skill ran, but so did a skill that should have stayed quiet. A skill that scores well alone and badly in company has a selection problem, and editing the body will not touch it.&lt;/p&gt;

&lt;p&gt;What I measured on Claude Opus 5, arranged by what actually distinguished the two skills:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e4ff3a-fc04-44f9-acac-b74890de20b9_3120x1760.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd6e4ff3a-fc04-44f9-acac-b74890de20b9_3120x1760.png" alt="A two-by-two of the routing eval: rows are name draws the line vs format-only name; columns are description draws the line vs vague description. Three corners route 8 of 8 correctly; only the corner where neither field draws a line breaks."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In this run, either signal alone held the line; only the condition where neither field distinguished the jobs produced misses and double-fires.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e59e5a9-eabf-4e57-8687-6e3b63783d5d_1960x904.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e59e5a9-eabf-4e57-8687-6e3b63783d5d_1960x904.png" alt="Results table on Claude Opus 5: distinct name and description 8 of 8 right; name only 8 of 8; description only 8 of 8; neither field distinct: 3 right and 1 none for customer, 2 right and 2 both for export."&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the date-adjacent negatives above, run against the distinct-description pair, both skills stayed quiet: 0 false triggers in 4. Those negatives ran against the sharp descriptions only, so this does not show how false triggers rise as a description gets vaguer. Run the same set against the descriptions you plan to ship: a clean positive-routing score will not tell you whether a skill grabs adjacent work it should leave alone. Small numbers, one model, one surface, and single runs. The routing choice is a model decision that can scatter, so a clean 8 out of 8 is one draw, not a settled rate; run each prompt a few times and read how often the right skill wins, not a single mark. This is a diagnostic you run on your own skills, not a benchmark, and the caption matters more than the cells: this is the shape of the thing, not what Opus 5 does in general. Two skills is the floor, not necessarily the hard case. A real catalogue may have several plausible neighbours, so run the eval beside the skills your target actually competes with, not only against a clean pair.&lt;/p&gt;

&lt;h2&gt;
  
  
  The habit
&lt;/h2&gt;

&lt;p&gt;A skill has two ways to fail. Its instructions can be wrong, and you probably test that already. Or Claude can never choose it, and that one leaves no mark: the skill sits installed, looking healthy, and quietly does nothing.&lt;/p&gt;

&lt;p&gt;Test whether Claude chooses the skill. The output looking right does not prove the skill ran. The skills you never test that way are the ones you only think are working.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;The pack here, two skills, twelve labelled prompts, and the &lt;code&gt;run.sh&lt;/code&gt; reader, needs only the &lt;code&gt;claude&lt;/code&gt; CLI and &lt;code&gt;jq&lt;/code&gt; and is yours to keep: &lt;a href="https://durabilitycurve.com/tools/skill-routing-eval/" rel="noopener noreferrer"&gt;durabilitycurve.com/tools/skill-routing-eval&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>tutorial</category>
      <category>productivity</category>
    </item>
    <item>
      <title>The Work That Comes Due After You Leave: the step you forgot</title>
      <dc:creator>Harry Floyd</dc:creator>
      <pubDate>Wed, 26 Aug 2026 19:07:17 +0000</pubDate>
      <link>https://dev.to/harryfloyd/the-work-that-comes-due-after-you-leave-5blb</link>
      <guid>https://dev.to/harryfloyd/the-work-that-comes-due-after-you-leave-5blb</guid>
      <description>&lt;h1&gt;
  
  
  The Work That Comes Due After You Leave
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9u57nvj1qrelidt3g252.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9u57nvj1qrelidt3g252.webp" alt="The read-across: your checklist on the left, a record you did not write on the right." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You finish something. A project wraps, a client signs off, a piece of work goes out for the last time. Then there is the tail, the small handful of things you do afterwards, none of which take any real time. Mark it done. Tell them it is finished. Cancel the paid seat you bought for it. Switch off the weekly update that goes out to them every Monday.&lt;/p&gt;

&lt;p&gt;Four steps, four different places: the tracker, your email, wherever the card gets charged, whatever tool sends that update. Later you check one of them, probably the tracker, because that is where you look to see whether things are finished. It tells you the job is done, and it is telling the truth about the only step it can see.&lt;/p&gt;

&lt;p&gt;Switching off the update is the one that did not happen. Months later it is still arriving, every Monday at nine, to someone who stopped being your client a long time ago. Nothing is wrong with the system that sends it. It is doing exactly what it was told, on time. From where you are standing, the failure looks exactly like everything working, and that is the whole of the problem.&lt;/p&gt;

&lt;p&gt;The tracker is not lying. A job that touches four systems has four different ways of still being open, and the tracker sees only the one it holds. Done was never one state; a single word just made it look like one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the small steps are the ones that go missing
&lt;/h2&gt;

&lt;p&gt;The tempting explanation is that you were busy, or careless, or need a better checklist. I want to offer a more specific one, because it tells you which steps will go wrong instead of telling you to try harder.&lt;/p&gt;

&lt;p&gt;A checklist is a list of the steps you thought of. It is good at holding you to those. What it cannot do is mention a step that never went on it, and the steps that never go on it are not random. They are the ones that cross into a system you do not quite think of as part of the job. You wrote the list around the place you do the work, and the step that lives somewhere else did not occur to you, for the same reason it will not later occur to you to check whether it happened.&lt;/p&gt;

&lt;p&gt;Making a second list does not save you, and that is the part worth sitting with. If you build the second list from the same picture of the job, the same step is missing from it too, and now you have two records that agree with each other and are both wrong. That is not a hypothetical: your tracker is that second list. You filled it from the same picture of the job, so it agreed the work was done and was wrong in the same place you were.&lt;/p&gt;

&lt;p&gt;And a missed closing step does not stay missed quietly. An ordinary task you skip just sits there until you come back to it; a closing step you skip stays open until something closes it, and until then it keeps acting, every day or every month, on its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check has to come from somewhere you did not write
&lt;/h2&gt;

&lt;p&gt;So the thing that catches the missing step cannot be your own account of the work. It has to be a record that something else kept, for its own reasons, whether or not you remembered the step.&lt;/p&gt;

&lt;p&gt;You already have several of these. You just do not read them against the job. The card statement is one: the bank records the charge whether or not you remember the seat you meant to cancel, so the seat that is still billing turns up as a line you cannot attach to any live piece of work. The access list is another: the system logs who can get in whether or not anyone told it that a person left, so the account that outlived the project is a login with no current owner. What actually shipped is recorded by the thing that shipped it, so a promise you made and never delivered stands as a commitment on one side with no send on the other.&lt;/p&gt;

&lt;p&gt;Even with a checklist I take seriously, I did this. I keep a written routine for finishing an essay, detailed, with a warning next to the item that slips most, and my archive quietly slipped twenty-three pieces behind what I had published since late spring. Copying each finished piece across to that archive had never been a line on the routine at all: it lived on a different system, so it never occurred to me to write it down. What caught it was the published record of what had actually gone out, kept by the platform and owing nothing to my memory. Held against the archive, it showed the twenty-three at once.&lt;/p&gt;

&lt;p&gt;The move itself is old. Accountants have reconciled two sets of books this way for centuries, and there is nothing here to invent. What is easy to get wrong is what makes the second record worth anything: not that it is a second record, but that something other than your own memory produced it. Two dashboards drawn from the same database, or two lists built from the same picture of the job, only look like a check, because the same forgetting shaped both. A record can catch you only when your forgetting could not have reached it too.&lt;/p&gt;

&lt;h2&gt;
  
  
  One question to carry
&lt;/h2&gt;

&lt;p&gt;That gives you a single question, and it is worth more than any checklist. Of anything you lean on to tell you a job is finished, ask: would this still be here, and still say the same thing, if I had forgotten the step entirely?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg7p13pmia4i9rcxwb8is.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg7p13pmia4i9rcxwb8is.webp" alt="The test, applied to two records." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The test, applied to two records. The one you fill in yourself fails it; the one the bank writes passes it, because the charge is there whether or not you remembered.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The tracker fails that question: you fill it in yourself, so a step you forget is a step you also forget to log, and it stays green over a gap it never knew about. The card statement passes, because the charge is there whether or not you remembered the seat. A check built from your own memory cannot expose the step that memory left out.&lt;/p&gt;

&lt;p&gt;The question keeps its shape as the instrument gets bigger. A tracker, a dashboard, a report you write on your own project: each is an instrument you fill from your own picture of the work, and each is blind in the same place you are. The statement is worth more than any account you write of what you meant to do, for the same reason an audit leans hardest on evidence the audited side did not get to shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part you can use
&lt;/h2&gt;

&lt;p&gt;Here is the version you can run on your own job this week.&lt;/p&gt;

&lt;p&gt;Write out the routine you run after you finish something, every step, including the ones that feel too small to be worth writing down. Mark each with the system it touches: the tracker, the calendar, the billing account, the shared drive, the tool somebody set up before you arrived. This is not the check yet. It is how you find which records are worth reading against each other, and it usually turns what felt like one job into the three or four systems it was always made of.&lt;/p&gt;

&lt;p&gt;Then there are two ways to keep a step from being lost, and the first is much stronger. Where you can, do not rely on catching the step at all; arrange things so that forgetting it does no harm. Anything that runs on its own, a payment, a subscription, a recurring invite, an access granted for a single project, gets its end date on the day you set it up, while you still know what it was for. Something that expires unless it is renewed cannot outlast your forgetting, because forgetting it and ending it become the same act. Reach for this first; it removes the obligation instead of watching it. Its limit is the one this piece began with: you can only set an end date on a step you thought of, and the step that never made the list cannot be made self-closing.&lt;/p&gt;

&lt;p&gt;For everything you could not foresee, or cannot make expire, there is the slower move: read your own record against one you did not produce. Your active-projects list against the vendor or card statement, looking for a charge attached to work that has already finished. Your list of who is on the team against the access export from whatever holds the accounts, looking for a login with no owner. The commitments in a signed contract against what your team actually sent, looking for a promise with no matching send. Choose the second record by the causal test, not by where it happens to be stored: pick the one your own memory did not shape.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg8glohzs3yg4cw58y6q.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frg8glohzs3yg4cw58y6q.webp" alt="Three records you keep, each read against one you did not write." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Three records you keep, each read against one you did not write. The last row is the limit: recorded nowhere, so nothing catches it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;How often you look depends on how much damage you will let build up first. A ten-pound seat can wait a month; a former colleague who can still open every file cannot, and something confidential still reaching the wrong person is not a scheduled job at all. None of it needs a tool you have to build: a read-only export or a screenshot is enough, and where you cannot pull the record yourself, the person who can is an email away, not a project. And reading across only points to a mismatch; you still have to look and decide whether it is a real miss, a timing lag or a duplicate, and keep that verdict for yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What none of this fixes
&lt;/h2&gt;

&lt;p&gt;Two things survive all of it, and I would rather say so than leave the tidy version standing.&lt;/p&gt;

&lt;p&gt;The first is the record that was never kept. If something gets finished and lands in no system at all, no charge, no log, no row anywhere, then there is no second record to read it against. You cannot check against a record that does not exist. That case surfaces only when a person happens to notice, or is told.&lt;/p&gt;

&lt;p&gt;The second is quieter, and more common. If the same blind spot sits in both records, they agree, and the agreement looks like an all-clear. This is the failure I walked into the first time I tried to build a check like this for myself. I searched my files for links to the publication. That sounds like reading an independent record, until you notice it read the same surface I would have: it counted the times I had linked to old pieces inside new ones as though that proved the old ones had shipped. Independence is the whole of the mechanism, and when it is missing it fails without a sound.&lt;/p&gt;

&lt;p&gt;So the honest tally is smaller than the tidy one. The obligations that leave a trace in a record I did not write, I can now catch, once in a while, in half an hour. The ones that touch nothing outside my own attention, I am still carrying in my head, and I have learned how little the word covers when I say nothing is wrong. &lt;em&gt;Nothing is wrong&lt;/em&gt; and &lt;em&gt;nothing I can see is wrong&lt;/em&gt; are different sentences, and most of the time only one of them is available to any of us.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://harryfloyd.substack.com/subscribe?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=work-that-comes-due-after-you-leave" rel="noopener noreferrer"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fofu3ldypr4j1xnkibni8.webp" alt="The Day Job banner" width="799" height="280"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>analysis</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
