<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Levelbrook Consulting</title>
    <description>The latest articles on DEV Community by Levelbrook Consulting (@levelbrook).</description>
    <link>https://dev.to/levelbrook</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4135900%2Fb51865e6-0335-434f-9b55-4a1d48ed4700.png</url>
      <title>DEV Community: Levelbrook Consulting</title>
      <link>https://dev.to/levelbrook</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/levelbrook"/>
    <language>en</language>
    <item>
      <title>Everybody has lost their minds, and the boring companies are quietly winning</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:31:43 +0000</pubDate>
      <link>https://dev.to/levelbrook/everybody-has-lost-their-minds-and-the-boring-companies-are-quietly-winning-4ejn</link>
      <guid>https://dev.to/levelbrook/everybody-has-lost-their-minds-and-the-boring-companies-are-quietly-winning-4ejn</guid>
      <description>&lt;p&gt;&lt;em&gt;A senior security engineer wrote this week that he spends three quarters of his time on AI and it has robbed him of his enjoyment of the work. He is right about the symptom and wrong about the cause. The value from these tools is landing in the least glamorous places in the company, and the people looking at the frontier are looking the wrong way.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The vent, and the sentence buried in it
&lt;/h2&gt;

&lt;p&gt;Jan Schaumann's post this week is a vent and he says so in the first line. A senior security&lt;br&gt;
engineer, decades in, writing that he spends upwards of three quarters of his time directly or&lt;br&gt;
indirectly dealing with AI and that it has robbed him of most of his enjoyment of the work. People&lt;br&gt;
with no engineering background pitching industry-changing solutions from their agent-infested&lt;br&gt;
homelab. Emails that read like influencer posts. Colleagues turned into meat proxies. It hit the&lt;br&gt;
front page because a great many people feel exactly this and most of them are not allowed to say&lt;br&gt;
it at work.&lt;/p&gt;

&lt;p&gt;You can disagree with a lot of it, and the thread did. But there is one paragraph in the middle&lt;br&gt;
that is not a vent. It is an argument, and it is the most important thing anyone wrote about AI&lt;br&gt;
this week.&lt;/p&gt;

&lt;p&gt;He describes the industry-wide effort, now many months old, to point frontier models at&lt;br&gt;
vulnerability discovery. Dozens of highly paid engineers per organisation, priorities reshuffled,&lt;br&gt;
harnesses built, pipelines built to shoehorn thousands of findings into vulnerability management.&lt;br&gt;
Thousands of new vulnerabilities found. And then: I don't think we're any safer than before. Because&lt;br&gt;
finding vulnerabilities has never been the bottleneck in information security. The bottleneck is&lt;br&gt;
still, as ever before, getting the packages updated. Patching is still hard.&lt;/p&gt;

&lt;p&gt;That is the whole essay. The models are extraordinary at the part that was never the constraint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meanwhile, at the frontier
&lt;/h2&gt;

&lt;p&gt;Consider what the rest of the industry was looking at while he wrote that.&lt;/p&gt;

&lt;p&gt;A researcher quit one of the labs and posted a warning that went, by ThePrimeagen's count on&lt;br&gt;
stream, to well over a hundred million views in a day. A colleague who stayed put the odds of&lt;br&gt;
catastrophe above ten percent in a decade. The lab's CEO published a three-point plan to pace the&lt;br&gt;
frontier. A rival CEO agreed with him, which people found unusual. AI Explained spent a video on the&lt;br&gt;
researchers' stated reason: a large gap, in one OpenAI researcher's phrase, between the internal&lt;br&gt;
and external perception of the rate of progress. Six axes of improvement, none near saturation.&lt;br&gt;
Fireship covered a 154-page threat report cataloguing eight months of misuse across seven&lt;br&gt;
categories.&lt;/p&gt;

&lt;p&gt;All of this is real and none of it is unimportant. But look at who is in the audience. Engineers,&lt;br&gt;
managers, founders, people who have to decide on Monday what to do with a budget. And the frontier&lt;br&gt;
discourse gives them precisely nothing to do on Monday, because it is about capabilities they do not&lt;br&gt;
control, on timelines they cannot influence, at companies they do not work for. It produces&lt;br&gt;
anxiety with no action attached, which is the most tiring kind, and it is why Schaumann is&lt;br&gt;
exhausted.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnnomaxywvinbg89zte85.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnnomaxywvinbg89zte85.png" alt="Where the attention went this week, and where a Monday decision can actually land. Nothing in the top row is actionable by anyone outside a lab." width="800" height="691"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Where the attention went this week, and where a Monday decision can actually land. Nothing in the top row is actionable by anyone outside a lab.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the value is landing
&lt;/h2&gt;

&lt;p&gt;We install AI systems in ordinary companies for a living, which gives us a view of this that the&lt;br&gt;
frontier discourse does not have, and the view is this. The money is being made in the bottom row of&lt;br&gt;
that figure, by people who are not on Hacker News, doing things nobody will write a threat report&lt;br&gt;
about.&lt;/p&gt;

&lt;p&gt;The shapes are always the same, and we describe them here as composites rather than clients, as&lt;br&gt;
everything on this site is. A clinic group that uses a model to reconcile years of insurance&lt;br&gt;
remittances against deposits and finds the pattern of underpayments a human was too busy to see. A&lt;br&gt;
logistics firm whose intake now reads every inbound document and writes the structured record, so&lt;br&gt;
the person who used to do that handles only the exceptions. A software team whose agent does not&lt;br&gt;
write features but does read every production error, correlate it to a deploy, and open a ticket&lt;br&gt;
with the likely cause before a human has seen the alert. And, in public this same week, Cloudflare&lt;br&gt;
saving another hundred terabytes of RAM with what their post cheerfully calls maths.&lt;/p&gt;

&lt;p&gt;None of that is at the frontier. All of it runs on models a generation or two behind the ones in&lt;br&gt;
the headlines, because the constraint was never the model. Schaumann's line about vulnerability&lt;br&gt;
discovery generalises: in almost every function of almost every company, the thing the model is&lt;br&gt;
best at was not the bottleneck, and the bottleneck is some unglamorous downstream step that nobody&lt;br&gt;
funded because it was boring. Patching. Inventory. The written specification. The approval seat.&lt;br&gt;
The person who has to say yes.&lt;/p&gt;

&lt;p&gt;The companies winning this year are the ones that noticed the bottleneck first and pointed the&lt;br&gt;
tools at it, or more often pointed the tools at everything upstream of it and then spent the&lt;br&gt;
savings on the bottleneck. They are not talking about it, partly because it is unglamorous and&lt;br&gt;
partly because it is working.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the frontier framing hurts
&lt;/h2&gt;

&lt;p&gt;There is a specific mechanism by which the frontier discourse makes ordinary organisations worse&lt;br&gt;
at this, and it is worth naming because it is avoidable.&lt;/p&gt;

&lt;p&gt;The frontier framing says the model is the variable. Wait for the next one; it will do what this&lt;br&gt;
one could not. That framing is true for the lab and false for the buyer, and a buyer who believes it&lt;br&gt;
does two bad things. They delay the boring work, because the boring work will surely be automated&lt;br&gt;
by the next release. And they evaluate every tool on ceiling, on the impressive demo, rather than&lt;br&gt;
on the floor, on what it does at two in the afternoon on a dull task with a tired reviewer.&lt;/p&gt;

&lt;p&gt;The second thing the framing does is imply that the risk is exotic. Swarms, bioweapons, the&lt;br&gt;
internet taken over. Meanwhile the actual risk in the actual company is the one CNN reported this&lt;br&gt;
week from the military: a model produced a confident report with things in it that were not true,&lt;br&gt;
and it got some distance up the chain before anyone checked. That failure does not require a&lt;br&gt;
frontier model. It requires an unreviewed one. Every company installing these tools this year is&lt;br&gt;
building that failure mode unless it builds the seat that catches it, and the seat is boring, and&lt;br&gt;
the frontier discourse never mentions it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Founftotiows4bpijg30i.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Founftotiows4bpijg30i.png" alt="The frontier framing versus the operator framing. Same tools, opposite conclusions about where to spend Monday." width="800" height="502"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The frontier framing versus the operator framing. Same tools, opposite conclusions about where to spend Monday.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How to find the boring bottleneck
&lt;/h2&gt;

&lt;p&gt;The instruction "find the least interesting problem and fix it first" is easy to nod at and hard to&lt;br&gt;
act on, so here is the exercise we run in the first week with any organisation.&lt;/p&gt;

&lt;p&gt;Pick one process that produces money or stops it. Invoices out, claims in, orders through, tickets&lt;br&gt;
closed. Walk it end to end with the people who do it, and at each step write down two numbers: how&lt;br&gt;
long the step takes when it goes well, and how long the work waits before that step starts. Almost&lt;br&gt;
nobody has the second number, and the second number is the process. A step that takes four minutes&lt;br&gt;
and waits two days is not a four-minute step.&lt;/p&gt;

&lt;p&gt;Then find the step where the wait is longest, and ask why the work is waiting. The answer is nearly&lt;br&gt;
always one of three things. A person has to decide something and is busy. A piece of information is&lt;br&gt;
missing and somebody has to go and find it. Or two systems disagree and a human has to reconcile&lt;br&gt;
them by hand. Those three are the bottleneck, and none of them is "the model is not smart enough".&lt;/p&gt;

&lt;p&gt;Now point the tools. Missing information is the easiest: a model that reads the inbound document&lt;br&gt;
and fills the record removes the wait entirely. Systems that disagree is the next: a model that&lt;br&gt;
reconciles the ninety percent that match and hands the rest to a person, with the mismatch&lt;br&gt;
highlighted, turns a day of reconciliation into an hour. The busy person deciding is the hardest and&lt;br&gt;
the most valuable, and it is the approval seat: give them the decision in a form that takes eight&lt;br&gt;
seconds and they will make forty before lunch.&lt;/p&gt;

&lt;p&gt;Nothing about this requires the frontier. It requires a whiteboard, an afternoon, and a willingness&lt;br&gt;
to be interested in a process everyone in the building considers beneath them. The companies&lt;br&gt;
winning this year did that afternoon a year ago.&lt;/p&gt;

&lt;h2&gt;
  
  
  Schaumann is also wrong, and it matters how
&lt;/h2&gt;

&lt;p&gt;The honest paragraph. He is wrong that the technology is the reason the work stopped being&lt;br&gt;
enjoyable, and the thread said so in the most upvoted reply: the models genuinely produce a lot of&lt;br&gt;
good work, can debug in an afternoon what took a team a week, and the people running the companies&lt;br&gt;
are, in that commenter's phrase, the most boring supervillains imaginable. All true at once.&lt;/p&gt;

&lt;p&gt;What robbed him of the enjoyment is the framing, not the tool. Seventy-five percent of a senior&lt;br&gt;
engineer's time spent on AI is seventy-five percent spent in the top row of the figure, on&lt;br&gt;
harnesses for vulnerability discovery that was never the bottleneck, on pipelines to process&lt;br&gt;
findings that will not be patched, on the frontier's priorities rather than the organisation's. Point&lt;br&gt;
the same engineer and the same models at the bottom row, at the inventory and the patching he&lt;br&gt;
himself names as the real work, and the time is not wasted and the work is not joyless. It is the&lt;br&gt;
job he signed up for, done faster.&lt;/p&gt;

&lt;p&gt;The frontier will keep moving. The researchers may well be right to be frightened; we are not&lt;br&gt;
qualified to say and neither is most of the audience. But the companies that come out of this decade&lt;br&gt;
ahead will not be the ones that watched it most closely. They will be the ones that found the&lt;br&gt;
boring bottleneck in their own building, wrote it down, and put the tools to work on either side of&lt;br&gt;
it while everyone else was reading the threat report.&lt;/p&gt;

&lt;p&gt;Everybody has lost their minds. The way back is to find the least interesting problem in the&lt;br&gt;
company and fix it first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.netmeister.org/blog/everybodys-lost-their-minds.html" rel="noopener noreferrer"&gt;Everybody's Lost Their Minds (Jan Schaumann)&lt;/a&gt; (HN, 365 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=J3ljHm57yU0" rel="noopener noreferrer"&gt;What AI Researchers Saw, Before Their Demand to 'Pace' AI (AI Explained, video)&lt;/a&gt; (16 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=W8IVKMGbUZE" rel="noopener noreferrer"&gt;I can't believe this is happening (ThePrimeagen, video)&lt;/a&gt; (19 Sep 2026; the viral-tweet numbers as reported on stream)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=7r4ikZHm9AI" rel="noopener noreferrer"&gt;Anthropic researchers are quitting... (Fireship, video)&lt;/a&gt; (15 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.cloudflare.com/saving-100-tb-of-ram-with-math/" rel="noopener noreferrer"&gt;Saving another 100TB of RAM (Cloudflare)&lt;/a&gt; (HN, 463 points, 18 Sep 2026)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/everybody-lost-their-minds-and-the-boring-companies-are-winning/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>If a rumour can summon ten thousand agents, your roadmap is a starting gun</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:29:12 +0000</pubDate>
      <link>https://dev.to/levelbrook/if-a-rumour-can-summon-ten-thousand-agents-your-roadmap-is-a-starting-gun-5615</link>
      <guid>https://dev.to/levelbrook/if-a-rumour-can-summon-ten-thousand-agents-your-roadmap-is-a-starting-gun-5615</guid>
      <description>&lt;p&gt;&lt;em&gt;Two mathematicians spent a year on a problem, a rumour of their approach reached a lab, and a swarm of agents reportedly reproduced it in days. Terence Tao's warning about open science applies just as well to your product. What was defensible last year is now a prompt.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The story, as told, with the caveats attached
&lt;/h2&gt;

&lt;p&gt;Here is the account that circulated last week, and it is important to say up front that it is a&lt;br&gt;
summary of public statements by parties who disagree with each other, so treat every clause as&lt;br&gt;
"reportedly".&lt;/p&gt;

&lt;p&gt;A mathematics professor and a collaborator who works at one of the labs spent about a year on the&lt;br&gt;
Navier-Stokes existence problem, one of the Millennium Prize problems. In the last month of that&lt;br&gt;
year they leaned heavily on coding agents and their progress accelerated; in mid-August they got a&lt;br&gt;
simpler related system to break, the closest anyone had come. Then, according to their statement, a&lt;br&gt;
different lab, having heard rumours of the approach, pointed a very large number of agents and a&lt;br&gt;
very large amount of compute at the problem and, within days, claimed a result on the full problem&lt;br&gt;
using what the mathematicians say was their novel approach. There was a phone call. The accounts of&lt;br&gt;
the phone call diverge sharply. Both sides published, one after the other, on the same Tuesday.&lt;/p&gt;

&lt;p&gt;You can read Fireship's telling, which is where most engineers heard it, and you can read the&lt;br&gt;
statements. We are not adjudicating it. What we want to point at is the line Terence Tao is&lt;br&gt;
reported to have written in response, because it is the most important sentence of the month for&lt;br&gt;
anyone who runs a product team, and almost nobody outside mathematics noticed it.&lt;/p&gt;

&lt;p&gt;If the mere rumour of your research can trigger a swarm of agents racing to front-run it,&lt;br&gt;
mathematicians will simply stop sharing ideas, and that will undo centuries of open science.&lt;/p&gt;

&lt;h2&gt;
  
  
  Replace "research" with "roadmap"
&lt;/h2&gt;

&lt;p&gt;Now do the substitution. If the mere rumour of your feature can trigger a swarm of agents racing to&lt;br&gt;
front-run it, what happens to your roadmap?&lt;/p&gt;

&lt;p&gt;For the entire history of the software business, the gap between having an idea and shipping it was&lt;br&gt;
the moat. Not the idea itself; ideas were always cheap and always leaked. The moat was that turning&lt;br&gt;
the idea into a working, deployed, maintained thing took a team a quarter, and a competitor who&lt;br&gt;
heard about it on a Tuesday could not have it by Friday. Product strategy, fundraising, hiring&lt;br&gt;
plans, launch timing, all of it was built on that gap being months wide.&lt;/p&gt;

&lt;p&gt;The maths story is what it looks like when the gap closes to days at the top of the market. The&lt;br&gt;
Dream RSI paper that Fireship covered the same week is what it looks like a level down: a search&lt;br&gt;
loop that, per the video's summary, wrote a solver that beat a standard library in about 300&lt;br&gt;
attempts where the fixed-policy version needed 550 and the previous record needed roughly 51,000.&lt;br&gt;
That is a search process getting two orders of magnitude cheaper at finding a solution somebody&lt;br&gt;
else already knew the shape of. Knowing the shape of the solution is what a roadmap leak gives a&lt;br&gt;
competitor.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zjj41e7rdr96mv25s0l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zjj41e7rdr96mv25s0l.png" alt="The idea-to-shipped gap was the moat. Durations illustrative; the collapse is the point." width="800" height="597"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The idea-to-shipped gap was the moat. Durations illustrative; the collapse is the point.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is not an argument for secrecy, for the same reason Tao's line is not an argument for&lt;br&gt;
mathematicians to stop publishing. Secrecy does not work either: your customers know what they&lt;br&gt;
asked for, your job postings say what you are building, and your own agents' instruction files&lt;br&gt;
describe the product in more detail than any leaked slide. It is an argument for being honest about&lt;br&gt;
what is still defensible when the build is no longer the hard part.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a swarm cannot front-run
&lt;/h2&gt;

&lt;p&gt;Go back to the maths. A swarm with a rumour and twenty million dollars of compute could,&lt;br&gt;
reportedly, reproduce a result. What it could not do is the year before the rumour: choosing that&lt;br&gt;
problem, choosing that approach out of the dozens that do not work, building the intuition that&lt;br&gt;
made the approach look promising when it looked like nothing to everyone else. The swarm needed the&lt;br&gt;
shape. The year produced the shape.&lt;/p&gt;

&lt;p&gt;The commercial version of that year is the thing we wrote about earlier this month under the&lt;br&gt;
heading that companies do not have processes, they have habits. The forty exceptions that live in&lt;br&gt;
one person's head. The knowledge of which customer needs it done differently because of the freight&lt;br&gt;
claim in 2023. The understanding of why the obvious version of the feature fails for the second&lt;br&gt;
largest account. A swarm given the rumour builds the obvious version, perfectly, in days. The&lt;br&gt;
obvious version is the one your customers already rejected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0c2w7vkoamaasshwnt7y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0c2w7vkoamaasshwnt7y.png" alt="What a swarm reproduces from a rumour, and what it cannot. The right-hand side is the whole of the defensible business." width="800" height="670"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What a swarm reproduces from a rumour, and what it cannot. The right-hand side is the whole of the defensible business.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;So the defensible assets in a world of swarms are the ones on the right of that figure, and they&lt;br&gt;
are exactly the assets most product organisations have neglected because they were not the&lt;br&gt;
bottleneck. The written specification, including the exceptions, which almost nobody has because&lt;br&gt;
the code was the specification and the code was expensive. The proprietary data about how the&lt;br&gt;
product fails and for whom. And the accountability, the fact that when the thing goes wrong there&lt;br&gt;
is a named organisation that picks up the phone and fixes it, which no swarm has ever offered&lt;br&gt;
anyone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes on the roadmap
&lt;/h2&gt;

&lt;p&gt;The practical consequences are uncomfortable for the way product teams have worked.&lt;/p&gt;

&lt;p&gt;Announcements shift from features to outcomes. Announcing "we are building X" now hands the shape&lt;br&gt;
to anyone with a swarm. Announcing "our customers in this segment no longer have this problem" hands&lt;br&gt;
them nothing they can prompt with, because the interesting part is which segment and which problem&lt;br&gt;
and why the obvious solution did not work, and that lives in your specification.&lt;/p&gt;

&lt;p&gt;The specification becomes the product. Not a slide, the actual document: what the system does,&lt;br&gt;
what it refuses to do, every exception and why. If your team's competitive advantage is knowledge&lt;br&gt;
that lives in heads, the swarm era is the strongest incentive you will ever get to write it down,&lt;br&gt;
because the written version is the only form in which it compounds and the only form in which it&lt;br&gt;
can be defended.&lt;/p&gt;

&lt;p&gt;Speed stops being a strategy and becomes table stakes. Being first to ship the obvious version was&lt;br&gt;
worth a great deal when it took a quarter. It is now worth roughly a week of attention. The teams&lt;br&gt;
that win will be the ones that were already on the third version, informed by the failures of the&lt;br&gt;
first two, when the swarm shipped the first one.&lt;/p&gt;

&lt;p&gt;And the Gowers point, from his essay on why he declined to sign the Fields medallists' letter,&lt;br&gt;
applies in full. He has argued for twenty-five years that mathematics contains two cultures,&lt;br&gt;
problem-solvers and theory-builders, and that the field needs both. The swarms are extraordinary&lt;br&gt;
problem-solvers. Theory-building, meaning the slow accumulation of a coherent understanding of a&lt;br&gt;
domain, is what makes the problems worth solving and tells you which one to solve next. Every&lt;br&gt;
product organisation is about to discover which of the two it was actually good at.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to write down this quarter
&lt;/h2&gt;

&lt;p&gt;The uncomfortable implication of the figure above is that the defensible assets are documents, and&lt;br&gt;
most organisations do not have them. So here is the writing programme, in priority order.&lt;/p&gt;

&lt;p&gt;The exceptions register. For the product's core workflow, every case where the obvious behaviour is&lt;br&gt;
wrong and why. Not the happy path, which anyone can reproduce, and not the code, which a swarm can&lt;br&gt;
regenerate. The list of forty things that live in the head of the person who has been there&lt;br&gt;
longest, with the reason attached to each. Ask that person to talk for two hours, record it, have a&lt;br&gt;
model draft the register, and have the person correct it. This is the single most valuable document&lt;br&gt;
the company can own and it usually does not exist.&lt;/p&gt;

&lt;p&gt;The refusal list. What the product deliberately does not do, and what happened when someone tried.&lt;br&gt;
A competitor building from a rumour will build the features you rejected, because they look like&lt;br&gt;
features. Knowing why they were rejected is a year of learning that cannot be prompted for.&lt;/p&gt;

&lt;p&gt;The failure data. Which customers hit which failures, how often, and what it cost. Nobody outside&lt;br&gt;
the company has this, and it is the input that decides which of the next ten things to build. Keep&lt;br&gt;
it in a form a model can read, because your own agents should be making the obvious version of the&lt;br&gt;
next feature for you, informed by it, before anyone else makes the obvious version uninformed.&lt;/p&gt;

&lt;p&gt;The accountability map. Who picks up the phone when the thing is wrong, in what timeframe, with&lt;br&gt;
what authority to fix it. Write it down, publish the shape of it to customers, and mean it. A swarm&lt;br&gt;
can ship a product. It cannot answer for one, and in the year when every product has a swarm-built&lt;br&gt;
twin, the answering is the product.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this is wrong
&lt;/h2&gt;

&lt;p&gt;The counter-argument is that most companies do not compete with labs holding twenty million dollars&lt;br&gt;
of spare compute, and that is true today. It will be less true every quarter, because the cost of&lt;br&gt;
the swarm is falling on a schedule nobody in your market controls, and because the swarm does not&lt;br&gt;
need to be a lab. It needs to be a competitor with a credit card and a rumour. The Navier-Stokes&lt;br&gt;
story is not the shape of your next year. It is the shape of your next three, arriving at the top&lt;br&gt;
of the market first, as these things do.&lt;/p&gt;

&lt;p&gt;Tao's warning was that scientists would stop sharing. The commercial equivalent is worse: companies&lt;br&gt;
will keep sharing, because they have to, and will discover that the only part they could ever have&lt;br&gt;
kept was the part they never wrote down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=aspmNhKAFMc" rel="noopener noreferrer"&gt;OpenAI's biggest math breakthrough is getting ugly (Fireship, video)&lt;/a&gt; (11 Sep 2026; the account of the Navier-Stokes dispute below is Fireship's summary of the parties' public statements, and both sides dispute the other's version)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=LoLYw--s-5w" rel="noopener noreferrer"&gt;Did Google just kickstart the intelligence explosion? (Fireship, video)&lt;/a&gt; (17 Sep 2026; the Dream RSI figures)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://gowers.wordpress.com/2026/09/17/why-i-didnt-sign-the" rel="noopener noreferrer"&gt;Why I didn't sign the Fields medallists' letter (Timothy Gowers)&lt;/a&gt; (HN, 276 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/playbook/you-do-not-have-processes/"&gt;Your company does not have processes. It has habits. (Levelbrook)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/your-roadmap-is-a-starting-gun/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>startup</category>
      <category>ai</category>
    </item>
    <item>
      <title>Readers can smell model prose in parts per trillion. So can the person reading your application.</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:26:39 +0000</pubDate>
      <link>https://dev.to/levelbrook/readers-can-smell-model-prose-in-parts-per-trillion-so-can-the-person-reading-your-application-4hj7</link>
      <guid>https://dev.to/levelbrook/readers-can-smell-model-prose-in-parts-per-trillion-so-can-the-person-reading-your-application-4hj7</guid>
      <description>&lt;p&gt;&lt;em&gt;Three of the most-read essays on Hacker News this week were about the same thing: writing with a model makes you sound like everyone else who writes with a model, and everyone can tell. Here is the mechanism, the list of tells, and the one way to use the tool that does not cost you your voice.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The week the readers pushed back
&lt;/h2&gt;

&lt;p&gt;Three essays, three very different writers, one complaint, all in the top of Hacker News inside four&lt;br&gt;
days. Thomas Ptacek, who is not a sentimental man about tools, wrote that readers can detect LLM&lt;br&gt;
words in the parts per trillion, and that however much you scuff up a model's paragraph it registers&lt;br&gt;
to your audience not as writing but as output. Martin Fowler wrote that models talk to him in a&lt;br&gt;
grating LLM-voice, an uncanny valley of talking to a real human. Erich Grunewald argued you should&lt;br&gt;
almost never use one to write at all, and the best comment under him added a rule worth keeping:&lt;br&gt;
only ever use AI to make yourself think harder, and more.&lt;/p&gt;

&lt;p&gt;And Jan Schaumann, in the angriest of the four, described people's emails now reading like&lt;br&gt;
LinkedIn-influencer posts with punchy single-sentence paragraphs, and half the people you interact&lt;br&gt;
with having turned into meat proxies for a model. That phrase went round because everyone has&lt;br&gt;
received that email.&lt;/p&gt;

&lt;p&gt;We want to take this out of the essayist's study and put it where it costs money. If you are&lt;br&gt;
applying for a job, writing a proposal, answering a customer, or posting anything under your own&lt;br&gt;
name, the person on the other end has now read several thousand model-drafted documents this year&lt;br&gt;
and has developed, without trying, a detector. When it fires, it does not think "AI". It thinks&lt;br&gt;
"this person did not bother", and then it thinks it about everything else you sent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it is detectable at all
&lt;/h2&gt;

&lt;p&gt;The mechanism is worth understanding because it explains why "humanising" prompts do not work.&lt;/p&gt;

&lt;p&gt;Ptacek's description is the most precise we have seen: frontier models are wedged in a mode where&lt;br&gt;
everything they write is a magazine headline. Every sentence is pleasing. Every phrase is the&lt;br&gt;
turn of phrase a good editor might have suggested. The problem is density. A human writer produces&lt;br&gt;
one such sentence per paragraph, if that, surrounded by ordinary load-bearing sentences that just&lt;br&gt;
carry information. A model produces nothing but the good ones, and the effect on a reader is the&lt;br&gt;
effect of a meal that is all garnish.&lt;/p&gt;

&lt;p&gt;There is a second layer, which Grunewald's piece names. Model prose is vague and wrong in&lt;br&gt;
hard-to-notice ways. It converges on the median expression of an idea. A specific claim becomes a&lt;br&gt;
general one; a number becomes "significant"; a mechanism becomes "a range of factors". The reader&lt;br&gt;
cannot always say what is missing but can feel that nothing is being risked, and prose that risks&lt;br&gt;
nothing reads as prose that knows nothing.&lt;/p&gt;

&lt;p&gt;The third layer is the one that catches people who think they have edited carefully. Models write&lt;br&gt;
in a small set of structural tics that humans almost never use spontaneously and that stick out the&lt;br&gt;
moment you know to look. We keep a list. It is the list our own linter runs against everything on&lt;br&gt;
this site before it publishes, and it is worth reproducing because most people have never seen it&lt;br&gt;
written down.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkl3nbxjfewnucfdd10rc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkl3nbxjfewnucfdd10rc.png" alt="The tells, grouped by how they get into a document. The linter that gates this site checks for every item in the first two groups." width="800" height="582"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The tells, grouped by how they get into a document. The linter that gates this site checks for every item in the first two groups.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first two columns can be caught mechanically and we catch them. The third column cannot, and it&lt;br&gt;
is the one that matters, because it is the column that describes what a document is missing rather&lt;br&gt;
than what it contains. A model-written cover letter is not detected by its vocabulary. It is&lt;br&gt;
detected by the absence of the one specific sentence that only this applicant, about this company,&lt;br&gt;
could have written.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this costs the most
&lt;/h2&gt;

&lt;p&gt;We spend a lot of our time on the receiving end of documents, and there are three places where the&lt;br&gt;
detector fires and the cost is immediate.&lt;/p&gt;

&lt;p&gt;Applications. A hiring manager reading forty cover letters for one role has, by letter fifteen,&lt;br&gt;
stopped reading for content and started reading for whether a person is present. The model-drafted&lt;br&gt;
letter is fluent, well-structured, mentions the company's mission, and could have been sent to any of&lt;br&gt;
the other thirty-nine companies with the name changed. It usually was. The reader is not offended.&lt;br&gt;
They just move on, and they do it in about four seconds, which is less than the eight seconds we&lt;br&gt;
usually talk about because there is nothing to evaluate.&lt;/p&gt;

&lt;p&gt;Proposals and outreach. The cold email that reads like a person, with one observation only that&lt;br&gt;
sender could have made and one offer only that sender could make, is answered. The one that reads&lt;br&gt;
like output is filtered, by a human or by the recipient's own model, which has been trained on the&lt;br&gt;
same tells. There is a grim symmetry in a model-drafted email being triaged into the bin by a model.&lt;/p&gt;

&lt;p&gt;Posts and essays under your name. This is the one that quietly damages careers, because it&lt;br&gt;
happens in public and it is permanent. A person who publishes model prose for a year has, at the&lt;br&gt;
end of it, a body of work that establishes nothing about how they think. Everyone who read it&lt;br&gt;
knows. Nobody says so.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the reader on the other side actually does
&lt;/h2&gt;

&lt;p&gt;It is worth describing the receiving end concretely, because most people writing applications have&lt;br&gt;
never sat on it.&lt;/p&gt;

&lt;p&gt;A hiring manager with forty applications and a Friday afternoon does not read them. They sort&lt;br&gt;
them. The first pass is under ten seconds per letter and it is looking for one thing: is there a&lt;br&gt;
sentence in here that could only have been written by this person, to us. A specific thing about&lt;br&gt;
the company that is not on the About page. A specific thing the applicant did, with a number or a&lt;br&gt;
name attached, that connects to the role in a way the applicant had to think about. One sentence is&lt;br&gt;
enough to move the letter to the second pile.&lt;/p&gt;

&lt;p&gt;The model-drafted letter almost never has that sentence, for a structural reason rather than a&lt;br&gt;
lazy one. The model does not know the specific thing, because the specific thing lives in the&lt;br&gt;
applicant's head, and the applicant did not put it in the prompt because the whole point of using&lt;br&gt;
the model was not to have to think. What comes out is fluent, complete, warm, and generic, and it&lt;br&gt;
lands in the first pile with the other thirty-one that read the same way.&lt;/p&gt;

&lt;p&gt;The second pass, on the eight that survived, is where fluency starts to count against you. The&lt;br&gt;
manager is now reading properly, and the tells in the first two columns of the figure above start&lt;br&gt;
to register: the punch paragraphs, the tricolons, the perfectly even rhythm. They do not think&lt;br&gt;
"this was written by a model". They think "I have read this letter before", which is true, and&lt;br&gt;
they start to wonder whether the specific sentence that got the letter into this pile was also&lt;br&gt;
borrowed. That doubt is expensive and the applicant never learns it was there.&lt;/p&gt;

&lt;p&gt;The way through is the method below, applied to a document that is three paragraphs long and&lt;br&gt;
therefore has nowhere to hide. Write it yourself, badly. Have the model tell you what is vague and&lt;br&gt;
what repeats. Fix it yourself. Then read it once and ask whether the person on the other end could&lt;br&gt;
tell it was written to them. If the answer is no, the model cannot help, and you should go and find&lt;br&gt;
out one more thing about the company.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one way to use the tool
&lt;/h2&gt;

&lt;p&gt;None of the four essays says do not use models, and neither do we. Ptacek's method is the right one&lt;br&gt;
and it has two rules that are easy to state and hard to follow.&lt;/p&gt;

&lt;p&gt;First, write the thing yourself. All of it. The first draft is where the thinking happens, and the&lt;br&gt;
draft you did not write contains thinking you did not do, which is exactly what the reader detects.&lt;br&gt;
This is the point Grunewald makes with the philosopher's line about the difference between nodding&lt;br&gt;
along while reading and productively generating a text. Generating is the work. Everything else is&lt;br&gt;
formatting.&lt;/p&gt;

&lt;p&gt;Second, use the model as a copyeditor, never as a ghostwriter, and never accept a single word it&lt;br&gt;
suggests. Ptacek is strict about this and he is right to be. Ask it what is wrong: where the&lt;br&gt;
passive voice piles up, which phrase you have used four times, where the argument skips a step,&lt;br&gt;
which paragraph a hostile reader would attack first, which sentences you could cut. Then fix every&lt;br&gt;
one of those yourself, in your own words. The moment you paste in its rewrite, the tell is in the&lt;br&gt;
document, and you will not see it because it is, sentence by sentence, better than what you had.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwx2ky6v8stpl0lyn9axg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwx2ky6v8stpl0lyn9axg.png" alt="The copyeditor loop. The model touches the draft only to point at it; every change is made by the writer in the writer's own words." width="800" height="1189"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The copyeditor loop. The model touches the draft only to point at it; every change is made by the writer in the writer's own words.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Ptacek's second rule, the one people skip, is to forbid the model from encouragement and then be&lt;br&gt;
vigilant about praise anyway. A model told "this is good, tighten it" will tell you it is good.&lt;br&gt;
Your first draft is not good. Its structure is wrong and it has seven hundred words you do not&lt;br&gt;
need. The rewrites you do in response to that discovery are, in his phrase, load-bearing parts of&lt;br&gt;
your voice, and a tool that talks you out of them has made your writing worse while feeling helpful.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limit of the claim
&lt;/h2&gt;

&lt;p&gt;There are documents where none of this matters. A status update, a changelog, a meeting summary, a&lt;br&gt;
form. Nobody is reading those for a person and a model should write them. The claim is narrower and&lt;br&gt;
sharper than "never use AI to write". It is that any document whose purpose is to establish that&lt;br&gt;
you, specifically, think something, cannot be delegated without defeating its purpose, and that the&lt;br&gt;
reader will know.&lt;/p&gt;

&lt;p&gt;The tools are extraordinary at finding what is wrong with your writing. They are useless at being&lt;br&gt;
you. Hire them for the first job and do the second one yourself, and the parts-per-trillion&lt;br&gt;
detector on the other end of the wire will pass you, for the plain reason that there will be&lt;br&gt;
someone there to detect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://sockpuppet.org/blog/2026/09/17/how-to-write-with-an-llm/" rel="noopener noreferrer"&gt;How To Write With An LLM (Thomas Ptacek, sockpuppet.org)&lt;/a&gt; (HN, 641 points and 380 comments, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://erichgrunewald.substack.com/p/why-you-should-almost-never-use-ai" rel="noopener noreferrer"&gt;I think you should almost never use AI to write (Erich Grunewald)&lt;/a&gt; (HN, 262 points, 19 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://martinfowler.com/articles/2026-dont-like-llms.html" rel="noopener noreferrer"&gt;I Don't Like LLMs (Martin Fowler)&lt;/a&gt; (HN, 237 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.netmeister.org/blog/everybodys-lost-their-minds.html" rel="noopener noreferrer"&gt;Everybody's Lost Their Minds (Jan Schaumann)&lt;/a&gt; (HN, 365 points, 17 Sep 2026; the "meat proxies" line)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/readers-smell-llm-prose-in-parts-per-trillion/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>writing</category>
      <category>career</category>
      <category>ai</category>
    </item>
    <item>
      <title>The free coding agent that uploaded your git history</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:24:08 +0000</pubDate>
      <link>https://dev.to/levelbrook/the-free-coding-agent-that-uploaded-your-git-history-4dlg</link>
      <guid>https://dev.to/levelbrook/the-free-coding-agent-that-uploaded-your-git-history-4dlg</guid>
      <description>&lt;p&gt;&lt;em&gt;A coding agent offered free this month was found snapshotting workspaces, git history included, to the vendor's cloud. It is the second such story this year. The tool you let touch your repository has more access than any contractor you have ever hired, and nobody ran a background check.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Second time this year
&lt;/h2&gt;

&lt;p&gt;The story, as reported this week, goes like this. A coding agent from one of the large model labs,&lt;br&gt;
promoted with a generous free tier this month, was found to be taking snapshots of the user's&lt;br&gt;
workspace, including the git history, and sending them to the vendor's cloud. The uploaded data was&lt;br&gt;
encrypted, a commenter noted, with a key the user does not hold. The reporting was careful to say it&lt;br&gt;
did not know the intent. The Hacker News thread was less careful, and its most upvoted line was that&lt;br&gt;
this is the second such story this year, after a similar discovery about another lab's agent&lt;br&gt;
earlier in the summer, and that the lesson from the first one was apparently not learned: do not&lt;br&gt;
trust harnesses, especially new ones.&lt;/p&gt;

&lt;p&gt;We are not going to relitigate the specifics, because we only know what was reported. The&lt;br&gt;
interesting part is not this vendor. It is that the discovery was made by a user reading network&lt;br&gt;
traffic, not by any process in any of the organisations that had installed the thing. Which means&lt;br&gt;
the question worth asking is not "was this one bad" but "what did you do, before you installed it,&lt;br&gt;
to find out".&lt;/p&gt;

&lt;h2&gt;
  
  
  The most privileged contractor you have ever hired
&lt;/h2&gt;

&lt;p&gt;Think about what a coding agent is, in access terms, and compare it to a human contractor.&lt;/p&gt;

&lt;p&gt;A contractor gets a laptop, a repository, a set of credentials scoped to their task, an NDA, a&lt;br&gt;
background check, and a manager who watches what they commit for the first month. They work during&lt;br&gt;
hours. They can be asked what they did. They can be fired.&lt;/p&gt;

&lt;p&gt;A coding agent gets your entire workspace, which in practice means every repository you have cloned,&lt;br&gt;
every &lt;code&gt;.env&lt;/code&gt; file you forgot was there, the shell history, the SSH keys the shell can reach, the&lt;br&gt;
cloud credentials in the credential helper, and the git history of everything, which is where the&lt;br&gt;
secret you rotated in 2024 still lives. It runs with your user's permissions. It runs while you are&lt;br&gt;
away. It makes outbound network requests you do not see, to endpoints you did not configure, and it&lt;br&gt;
was installed by a developer with a one-line command because the free tier was generous and the&lt;br&gt;
demo was good.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhfqa75dc9yznmrc2goh6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhfqa75dc9yznmrc2goh6.png" alt="What a coding agent can reach from a normal developer workstation, compared with what a contractor is given. Nothing on the left was scoped by anyone." width="799" height="718"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What a coding agent can reach from a normal developer workstation, compared with what a contractor is given. Nothing on the left was scoped by anyone.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Nobody would give a contractor the left-hand column. Every organisation with developers has given it&lt;br&gt;
to several agents this year, from several vendors, some of which did not exist in January.&lt;/p&gt;

&lt;h2&gt;
  
  
  The threat is symmetric
&lt;/h2&gt;

&lt;p&gt;There is a second half to this story and it ran the same week. Fireship's summary of Anthropic's&lt;br&gt;
recent threat report, which the video says covers eight months of misuse the company detected and&lt;br&gt;
shut down, included one detail that should land hard for anyone who ships software. A criminal group&lt;br&gt;
was reported to have mass-downloaded around 1.8 million Android application packages, decompiled&lt;br&gt;
them with the help of a model, and mined them for hard-coded secrets. The keys they wanted most, the&lt;br&gt;
video notes with some relish, were API keys for the model providers themselves. Those are the&lt;br&gt;
easiest to monetise.&lt;/p&gt;

&lt;p&gt;Read the two stories together. On one side, agents installed on developer machines with access to&lt;br&gt;
everything and an outbound connection nobody audited. On the other, agents run by attackers,&lt;br&gt;
decompiling shipped software at scale looking for exactly the kind of secret that a vibe-coded app,&lt;br&gt;
or a workspace snapshot, hands over. The same capability that makes an agent useful to you makes&lt;br&gt;
it useful to the person on the other end of the network connection, and the asymmetry that used to&lt;br&gt;
protect the small shop, that nobody would bother to decompile your app by hand, is gone.&lt;/p&gt;

&lt;p&gt;The Hacktron write-up from the same week, on chaining a heap overflow and an SSO misconfiguration&lt;br&gt;
into access to internal repositories at one of the labs, makes the point from the top of the&lt;br&gt;
market. If the people building the models can be reached through a misconfiguration, the tool on&lt;br&gt;
your laptop was not built by people who are immune to the same thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "background check" means for a tool
&lt;/h2&gt;

&lt;p&gt;The remedy is unglamorous and it is mostly the security discipline you already apply to&lt;br&gt;
dependencies, applied to a category that has been exempted because it is exciting.&lt;/p&gt;

&lt;p&gt;Read the network. Before an agent touches a real repository, run it in a scratch workspace with a&lt;br&gt;
proxy in front of it and look at every host it talks to and what it sends. This takes an hour. It is&lt;br&gt;
how this week's story was discovered, by one person, and every organisation that installed the tool&lt;br&gt;
could have done it first.&lt;/p&gt;

&lt;p&gt;Scope the workspace. An agent should see one repository, not a home directory. Run it in a&lt;br&gt;
container, a devcontainer, a VM, a separate user, anything that means "your workspace" is a&lt;br&gt;
directory you chose rather than everything the shell can reach. If the tool does not work that way,&lt;br&gt;
that is information about the tool.&lt;/p&gt;

&lt;p&gt;Scope the credentials. Short-lived tokens, per-agent, for the one system the task needs. No&lt;br&gt;
credential helper with a year-long cloud key. No SSH agent forwarding. If the agent needs to push,&lt;br&gt;
give it a deploy key for that repository and nothing else.&lt;/p&gt;

&lt;p&gt;Gate the egress. An allow-list of hosts an agent may reach is a small piece of configuration and&lt;br&gt;
it converts "we found out from a blog post" into "the request failed and we looked". This is the&lt;br&gt;
single highest-return control on the list and almost nobody has it.&lt;/p&gt;

&lt;p&gt;Treat the instruction files as code. The AGENTS.md and the skills directory an agent reads are&lt;br&gt;
executed, in the sense that matters. Review them. Own them. Pin them. Cloudflare published an&lt;br&gt;
audit skill this week for exactly this kind of surface, which is a sign the serious shops have&lt;br&gt;
started treating agent configuration as something that needs auditing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhe1xbi4urpmuqjnwkgdf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhe1xbi4urpmuqjnwkgdf.png" alt="The controls that would have turned this week's story into a failed request. None of them require trusting the vendor." width="800" height="642"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The controls that would have turned this week's story into a failed request. None of them require trusting the vendor.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The hour, in detail
&lt;/h2&gt;

&lt;p&gt;Since the whole argument rests on "this takes an hour", here is the hour.&lt;/p&gt;

&lt;p&gt;Make a scratch directory containing a small repository with a deliberately fake secret in it: a&lt;br&gt;
file called &lt;code&gt;.env&lt;/code&gt; with &lt;code&gt;API_KEY=canary-&lt;/code&gt; followed by a long random string, committed once and then&lt;br&gt;
removed in a second commit so that it exists only in history. Put a second canary in the shell&lt;br&gt;
history. These are your tracers. If either string ever appears in an outbound request, you have&lt;br&gt;
your answer without reading anything else.&lt;/p&gt;

&lt;p&gt;Run the agent inside a container or a fresh user account with a local HTTP proxy configured as the&lt;br&gt;
system proxy and its certificate trusted, so that TLS traffic is visible. Point the agent at the&lt;br&gt;
scratch repository and give it a dull, real task: rename a function, add a test, fix a typo. Let it&lt;br&gt;
finish. Do it three times, once with a fresh session each time, because some tools snapshot on&lt;br&gt;
first run and some snapshot periodically.&lt;/p&gt;

&lt;p&gt;Then read the proxy log. You are looking for four things. Every distinct host contacted, which&lt;br&gt;
should be a short list you recognise. Any request whose body is large relative to what the task&lt;br&gt;
required, which is what a workspace snapshot looks like. Any request containing either canary&lt;br&gt;
string, which is the tracer firing. And any request that happened when the agent was idle, which&lt;br&gt;
is telemetry, and which you then read with more suspicion than the rest.&lt;/p&gt;

&lt;p&gt;Write the list of hosts down. That list becomes the egress allow-list for the tool on real&lt;br&gt;
machines, and the next time the tool updates you run the hour again and diff the lists. If a new&lt;br&gt;
host appears, you find out from your own log rather than from someone else's blog post, and you&lt;br&gt;
find out before it has touched anything that matters. An hour, once per tool, once per major&lt;br&gt;
update. There is no security control on the market with a better return.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this is unfair to the vendors
&lt;/h2&gt;

&lt;p&gt;Some of this is the ordinary growing pains of a new category. Telemetry that a team thought was&lt;br&gt;
obviously fine looks very different when a user reads the packet capture. Workspace snapshots have&lt;br&gt;
legitimate uses in resumable agent sessions. The labs are, by and large, staffed by people who would&lt;br&gt;
be horrified to be described as exfiltrating anything. None of that changes the buyer's position.&lt;br&gt;
You do not get to know intent. You get to know what the tool can reach and where it sends things,&lt;br&gt;
and both of those are measurable before you install it.&lt;/p&gt;

&lt;p&gt;The reason this is the second story of its kind this year and will not be the last is that the&lt;br&gt;
category moved faster than the discipline. Every other piece of software that runs with a&lt;br&gt;
developer's full permissions and talks to the internet went through a decade of being treated as a&lt;br&gt;
risk before it was treated as a convenience. Coding agents skipped that decade because they were&lt;br&gt;
useful immediately.&lt;/p&gt;

&lt;p&gt;They are still useful. Install them the way you would hire someone who will have the keys to&lt;br&gt;
everything: find out what they can reach, decide what they may reach, and watch the door.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blog.ferstar.org/en/posts/zcode-silent-workspace-snapshot/" rel="noopener noreferrer"&gt;Inside ZCode: silently uploading your Git history to the cloud (ferstar)&lt;/a&gt; (HN, 328 points, 18 Sep 2026; details below are as reported there and in the thread)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=7r4ikZHm9AI" rel="noopener noreferrer"&gt;Anthropic researchers are quitting... (Fireship, video)&lt;/a&gt; (15 Sep 2026; Fireship's summary of Anthropic's threat report, including the decompiled-APK figure)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.hacktron.ai/blog/hacking-openai" rel="noopener noreferrer"&gt;A heap overflow and SSO misconfiguration to compromise OpenAI internal repos (Hacktron)&lt;/a&gt; (HN, 484 points, 18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/cloudflare/security-audit-skill" rel="noopener noreferrer"&gt;Cloudflare security-audit-skill&lt;/a&gt; (HN, 210 points, 17 Sep 2026)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/the-free-coding-agent-that-uploaded-your-git-history/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>ai</category>
    </item>
    <item>
      <title>AGENTS.md is now the most-read document in your company. It is also the worst-written.</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:21:37 +0000</pubDate>
      <link>https://dev.to/levelbrook/agentsmd-is-now-the-most-read-document-in-your-company-it-is-also-the-worst-written-1g83</link>
      <guid>https://dev.to/levelbrook/agentsmd-is-now-the-most-read-document-in-your-company-it-is-also-the-worst-written-1g83</guid>
      <description>&lt;p&gt;&lt;em&gt;Claude Code started reading AGENTS.md this week, which means one file is now consulted by every coding agent on every task, hundreds of times a day, more than any wiki page has been read in the history of your organisation. Most of them are forty lines of contradictions written in an afternoon.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One file, every task, all day
&lt;/h2&gt;

&lt;p&gt;The changelog entry was one line and it got seven hundred points. Claude Code now reads a repository's&lt;br&gt;
AGENTS.md when there is no CLAUDE.md. The top comments were what you would expect: relief that the&lt;br&gt;
one-line CLAUDE.md files that just said "read AGENTS.md" can go, a joke about who borrowed the idea&lt;br&gt;
from whom, and a grumble that the skills directory is still proprietary. Standards converging. Nice.&lt;/p&gt;

&lt;p&gt;Step back from the tooling and look at what just happened to your documentation.&lt;/p&gt;

&lt;p&gt;For twenty years the most-read document in a software company was, depending on the company, the&lt;br&gt;
onboarding wiki, the README, or the runbook that everyone opened during the outage. Read, in each&lt;br&gt;
case, by a handful of humans a handful of times a year, mostly in their first month. Nobody measured&lt;br&gt;
it because the number would have been embarrassing.&lt;/p&gt;

&lt;p&gt;AGENTS.md is read by every agent, on every task, before it does anything. If your team runs a few&lt;br&gt;
hundred agent sessions a day, and many do now, that file is consulted a few hundred times a day. Its&lt;br&gt;
contents shape every change that gets made. It is, by a margin that is not close, the most consequential&lt;br&gt;
piece of writing in the building. And in most of the repositories we have seen this year it was&lt;br&gt;
written in forty minutes by whoever set up the tool, has never been reviewed by anyone, and contains&lt;br&gt;
at least two instructions that contradict each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is actually in there
&lt;/h2&gt;

&lt;p&gt;We have read a lot of these files. The pattern is remarkably consistent.&lt;/p&gt;

&lt;p&gt;The first third is tooling trivia. Run the tests with this command. Use this package manager. Do not&lt;br&gt;
touch the generated directory. This is fine and it is the part that gets maintained, because when it&lt;br&gt;
is wrong the agent breaks loudly.&lt;/p&gt;

&lt;p&gt;The second third is the accumulated scar tissue of things that went wrong. Never run the migration&lt;br&gt;
directly. Always check the feature flag first. Do not use the old client library. Each line was&lt;br&gt;
added after an incident by whoever was on call, in the language of that incident. Nobody has gone&lt;br&gt;
back to check whether the flag still exists.&lt;/p&gt;

&lt;p&gt;The last third is the interesting part: the habits. Prefer small functions. We do not use that&lt;br&gt;
pattern here. Follow the existing style. Ask before adding a dependency. This is the oral tradition&lt;br&gt;
of the team, written down for the first time in its history, by one person, from memory, in the&lt;br&gt;
voice of a Slack message.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqkov2a9qpetjtdta0gd1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqkov2a9qpetjtdta0gd1.png" alt="What a typical AGENTS.md contains, and who reads each part. The bottom third is the team's undocumented process, written down once, by one person, for the first time." width="800" height="334"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What a typical AGENTS.md contains, and who reads each part. The bottom third is the team's undocumented process, written down once, by one person, for the first time.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We wrote earlier this month that companies do not have processes, they have habits, plus a document&lt;br&gt;
that describes an idealised version of some of them. AGENTS.md is that document, except that this&lt;br&gt;
time the reader is not a new hire who will paper over the gaps by watching the person at the next&lt;br&gt;
desk. The reader is a system that will follow the instruction literally, hundreds of times, and&lt;br&gt;
will resolve the contradictions by picking one, silently, differently each time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure modes are already visible
&lt;/h2&gt;

&lt;p&gt;There was a comment under the self-driving-codebases piece this week that describes the shape of&lt;br&gt;
the thing perfectly. An engineer running their own agent loops watched the first agent try to run&lt;br&gt;
an enormous dependency inspection command, run out of memory, and record a workaround in the&lt;br&gt;
agent's memory. The workaround was a way to make the enormous command succeed. Every subsequent&lt;br&gt;
agent inherited the workaround. The instruction file had learned to do the wrong thing more&lt;br&gt;
reliably. They called it cargo-cult behaviour, and the name is right.&lt;/p&gt;

&lt;p&gt;Instruction files accrete. That is what they are for. But accretion without review produces a&lt;br&gt;
document whose instructions were each correct on the day they were written and which, taken&lt;br&gt;
together, describe a codebase that no longer exists. The agent does not know that. It reads the file&lt;br&gt;
fresh every time and does its best.&lt;/p&gt;

&lt;p&gt;The second failure mode is contradiction. "Always add tests" and "do not modify files outside the&lt;br&gt;
task scope" are both reasonable lines and they conflict on roughly a third of tasks. A human resolves&lt;br&gt;
that with judgement and a quick message. An agent resolves it by weighting, and the weighting is not&lt;br&gt;
something you configured.&lt;/p&gt;

&lt;p&gt;The third is the one that should worry whoever owns the codebase. The file is unowned. It has no&lt;br&gt;
review process, no owner in the CODEOWNERS sense, no changelog and no tests. Anybody who can push&lt;br&gt;
can add a line, and the line will be obeyed by every agent from then on. If you have ever worried&lt;br&gt;
about supply-chain risk in your dependencies, the highest-privilege dependency in your repository is&lt;br&gt;
now a Markdown file.&lt;/p&gt;

&lt;h2&gt;
  
  
  A worked example, from a file we were asked to look at
&lt;/h2&gt;

&lt;p&gt;A composite, assembled from several files we have reviewed this year, with the details changed. The&lt;br&gt;
file was 61 lines long. Line 9 said to always run the full test suite before opening a pull request.&lt;br&gt;
Line 34, added after an incident in the spring, said never to run the integration tests locally&lt;br&gt;
because they hit a shared staging database. The full suite included the integration tests. Every&lt;br&gt;
agent that read the file resolved the contradiction in one of two ways: it ran the full suite and&lt;br&gt;
hit staging, or it skipped the suite and opened the pull request untested. Which one it chose&lt;br&gt;
depended on the model, the day, and how much else was in the context window. Nobody had noticed&lt;br&gt;
for four months because both outcomes looked like normal agent behaviour.&lt;/p&gt;

&lt;p&gt;Line 22 said to use the internal HTTP client wrapper rather than the raw library. The wrapper had&lt;br&gt;
been deleted in a refactor in July. Agents that obeyed line 22 searched for the wrapper, failed to&lt;br&gt;
find it, and either recreated it from the description or fell back to the raw library with a&lt;br&gt;
comment apologising. Three copies of a near-identical wrapper had accumulated in the repository, each&lt;br&gt;
written by an agent trying to comply with an instruction about a thing that no longer existed.&lt;/p&gt;

&lt;p&gt;None of this is exotic. It is what happens to any document that is executed without being&lt;br&gt;
maintained, and the fix for all of it took one person one afternoon: read the file, delete eleven&lt;br&gt;
lines, rewrite four, add the reason to each remaining rule. The next month's agent sessions were&lt;br&gt;
measurably cleaner, and the person who did it described the afternoon as the highest-return work&lt;br&gt;
they had done all quarter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat it like what it is
&lt;/h2&gt;

&lt;p&gt;The remedy follows from taking the file seriously, and it is mostly process rather than tooling.&lt;/p&gt;

&lt;p&gt;Give it an owner. A named person, in CODEOWNERS, whose approval is required to change it. Not&lt;br&gt;
because the changes are dangerous individually but because somebody has to hold the whole file in&lt;br&gt;
their head and notice when line 14 and line 31 disagree.&lt;/p&gt;

&lt;p&gt;Review it on a cadence, as a document. Once a month, read the whole thing top to bottom, with a&lt;br&gt;
model if you like, and for every instruction ask three questions. Is this still true? Is it still&lt;br&gt;
needed? Does it conflict with anything else in here? Delete generously. A shorter file that is all&lt;br&gt;
true beats a longer one that is mostly true, because the agent cannot tell which lines are the&lt;br&gt;
mostly.&lt;/p&gt;

&lt;p&gt;Separate the layers. Tooling facts, safety rules and stylistic preferences are different kinds of&lt;br&gt;
instruction with different failure costs, and they should look different on the page. A safety&lt;br&gt;
rule that the agent must never violate should not be sitting in the same list as a preference about&lt;br&gt;
function length, in the same font, with the same weight.&lt;/p&gt;

&lt;p&gt;Write the why. This is the one that turns the file from scar tissue into process. "Never run&lt;br&gt;
migrations directly" is a rule. "Never run migrations directly, because production runs them&lt;br&gt;
through the deploy pipeline with a lock, and a direct run in 2025 took the API down for forty&lt;br&gt;
minutes" is a rule an agent can reason about, including reasoning about when it does not apply. It&lt;br&gt;
is also, incidentally, the first time that piece of institutional knowledge has been written down&lt;br&gt;
anywhere a new human could find it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn4hwfs0zy35wn28tcffz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn4hwfs0zy35wn28tcffz.png" alt="The instruction file, run like the load-bearing document it has become." width="734" height="1468"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The instruction file, run like the load-bearing document it has become.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;And test it. This sounds odd for a Markdown file and it is the most useful thing on the list. Keep&lt;br&gt;
a short set of tasks that went wrong in the past, the ones that produced the scar-tissue lines. Once&lt;br&gt;
a month, run an agent against them with the current file and see whether the lines still do their&lt;br&gt;
job. If the agent makes the old mistake, the instruction has rotted or been contradicted. If it does&lt;br&gt;
not, the line is earning its place. This is the same idea as a regression test, applied to the&lt;br&gt;
document that governs the thing that writes your code.&lt;/p&gt;

&lt;h2&gt;
  
  
  The larger point
&lt;/h2&gt;

&lt;p&gt;The reason this deserves an essay rather than a checklist is what the file reveals. For the first&lt;br&gt;
time, the habits of an engineering team have been written down in a form that is executed rather&lt;br&gt;
than merely consulted. That is uncomfortable, because the writing is bad and the habits are&lt;br&gt;
inconsistent and everybody can now see both. It is also the largest documentation opportunity a&lt;br&gt;
software organisation has ever had, because for once there is an immediate, measurable cost to the&lt;br&gt;
document being wrong, and an immediate, measurable benefit to it being right.&lt;/p&gt;

&lt;p&gt;The wiki was never read, so it never mattered that it was wrong. This file is read constantly. Write&lt;br&gt;
it like something that is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://code.claude.com/docs/en/changelog" rel="noopener noreferrer"&gt;Claude Code now reads AGENTS.md if there is no CLAUDE.md (changelog)&lt;/a&gt; (HN, 715 points, 18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/playbook/you-do-not-have-processes/"&gt;Your company does not have processes. It has habits. (Levelbrook)&lt;/a&gt; (the earlier essay this one builds on)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.detail.dev/posts/towards-self-driving-codebases" rel="noopener noreferrer"&gt;Towards Self-Driving Codebases (detail.dev)&lt;/a&gt; (HN, 119 points; the agent-memory cargo-cult comment is from the thread)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/agents-md-is-the-most-read-document-in-your-company/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>documentation</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to read a coding-agent benchmark without getting sold</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:19:06 +0000</pubDate>
      <link>https://dev.to/levelbrook/how-to-read-a-coding-agent-benchmark-without-getting-sold-2pf1</link>
      <guid>https://dev.to/levelbrook/how-to-read-a-coding-agent-benchmark-without-getting-sold-2pf1</guid>
      <description>&lt;p&gt;&lt;em&gt;A nine-author study this week pulled a coding agent apart into its components and measured each one across 176 configurations. The findings are less exciting than any vendor slide and more useful than all of them, and they hand the buyer four questions no benchmark answers.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The number on the slide is a car, and you are being sold an engine
&lt;/h2&gt;

&lt;p&gt;Every coding-agent pitch you have seen this year has a number on it. SWE-Bench Verified, some&lt;br&gt;
percentage, up and to the right, usually next to a model name. The implication is that the number&lt;br&gt;
belongs to the model, and that if you buy the model you get the number.&lt;/p&gt;

&lt;p&gt;A commenter on this week's Hacker News thread about the harness study put the problem better than&lt;br&gt;
the paper's abstract does. If Car A is faster than Car B, it is not necessarily the engine. It could&lt;br&gt;
be the tyres, the gearbox, the weight, the driver. A coding agent is a car. The model is the engine.&lt;br&gt;
The harness, meaning the loop around the model that decides what it sees, what it can do and when&lt;br&gt;
it stops, is everything else. And the number on the slide is a lap time for the whole car, measured&lt;br&gt;
on a track you do not drive on.&lt;/p&gt;

&lt;p&gt;Nine researchers at Fan et al. did the thing nobody selling these tools has an incentive to do. They&lt;br&gt;
held the model fixed, held the execution loop fixed, and varied three harness components one at a&lt;br&gt;
time: planning, action space, and context management. Four models, two benchmarks (SWE-Bench Verified&lt;br&gt;
and Terminal-Bench 2.1), 176 matched configurations, five context-management strategies, four&lt;br&gt;
context-window budgets. Then they looked at the trajectories, not just the scores, to see what each&lt;br&gt;
component actually changed about how the agent behaved.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;176&lt;/strong&gt; matched harness configurations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;4&lt;/strong&gt; models held fixed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;3&lt;/strong&gt; components varied: planning, action space, context&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;5&lt;/strong&gt; context-management strategies compared&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What they found, translated out of the abstract
&lt;/h2&gt;

&lt;p&gt;The findings are almost aggressively unglamorous, which is how you know they are worth something.&lt;/p&gt;

&lt;p&gt;Context management, the machinery that decides what to throw away as the conversation fills up,&lt;br&gt;
matters more the tighter the context budget, and most of its benefit comes from one thing: not&lt;br&gt;
falling over when the window overflows. It does not make the agent smarter. It lets the agent keep&lt;br&gt;
going. The strongest strategy in their comparison was the boring one, a rule-based pass that&lt;br&gt;
deletes obviously stale material before any model-based summarisation runs. And the clever&lt;br&gt;
addition everyone builds, making elided content recoverable so the agent can go back and fetch it,&lt;br&gt;
turned out to be machinery the models rarely used and which yielded no accuracy gain.&lt;/p&gt;

&lt;p&gt;Planning, meaning an explicit plan-first step, changes role depending on the model. For weaker&lt;br&gt;
models it is an accuracy scaffold. For stronger ones it stops helping accuracy and becomes a cost&lt;br&gt;
saver, with a small decrease in success rate as the price. Another commenter connected this to&lt;br&gt;
something in the Claude Code changelog: the built-in todo and task-tracking tools were switched off&lt;br&gt;
by default on the newest model generations. The vendor, in other words, appears to have measured the&lt;br&gt;
same thing.&lt;/p&gt;

&lt;p&gt;Action space, meaning whether the agent gets a menu of predefined tools or just a shell, splits the&lt;br&gt;
same way. Predefined tools raise success rates for models that are weak at driving bash. Models that&lt;br&gt;
are good at bash do better and cheaper with bash alone, most clearly on command-line-shaped tasks.&lt;br&gt;
The paper does not define "bash-capable" crisply, which a commenter rightly flagged, but the&lt;br&gt;
direction is unambiguous: every tool you add beyond the shell is a claim that needs testing, not a&lt;br&gt;
free improvement.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffh6d0iimbng6t3rp2gk2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffh6d0iimbng6t3rp2gk2.png" alt="What each harness component does, by model strength, per the study's abstract. Read the row for the model you actually run." width="800" height="496"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What each harness component does, by model strength, per the study's abstract. Read the row for the model you actually run.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The trajectory analysis is the part a buyer should care about most. Context management extended&lt;br&gt;
how long the agent could keep working without substantially changing what it did. Planning changed&lt;br&gt;
where trajectories stopped. Action space changed the granularity at which code got written. None of&lt;br&gt;
the three made the engine better. They changed the gearbox, the tyres and the driver, and the lap&lt;br&gt;
time moved accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters more than the leaderboard
&lt;/h2&gt;

&lt;p&gt;Put the study next to what the practitioners were saying this week and a picture forms.&lt;/p&gt;

&lt;p&gt;Theo Browne's video argued, with some heat, that people who cannot feel the difference between&lt;br&gt;
model generations are prompting badly, and his most useful idea was a picture of the distribution:&lt;br&gt;
every model has a ceiling, which is what the demos show, and a floor, which is what you hit at two&lt;br&gt;
in the afternoon on a boring task. Frontier models mostly raise the floor. ThePrimeagen, the same&lt;br&gt;
week, left a current-generation model on a trivial colour bug, came back forty minutes later, and&lt;br&gt;
found it had spent 330 million tokens and 118 dollars reading the same file over and over. That is&lt;br&gt;
a floor.&lt;/p&gt;

&lt;p&gt;A benchmark number is a ceiling measurement of one car on one track. It tells you nothing about the&lt;br&gt;
floor, and the floor is where your money goes. And the study tells you the floor is mostly a&lt;br&gt;
harness property: whether the loop notices it is stuck, whether the context gets cleaned before it&lt;br&gt;
overflows, whether the agent has a shell or a menu, whether there is a plan and whether the plan is&lt;br&gt;
worth its cost for the model you actually run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four questions to ask instead
&lt;/h2&gt;

&lt;p&gt;So when the next vendor slide arrives, the number is not the question. These are.&lt;/p&gt;

&lt;p&gt;Which harness produced this number, and can I see it? If the answer is "our proprietary agent&lt;br&gt;
runtime", you are buying a car and being told the horsepower. Ask for the loop: what the model&lt;br&gt;
sees, what it can do, how context is managed, when it stops.&lt;/p&gt;

&lt;p&gt;What does it do when it is stuck? Ask for the failure trajectories, not the success ones. A good&lt;br&gt;
vendor has them and is proud of them. Ask specifically what happens at context overflow and what&lt;br&gt;
happens after the tenth identical tool call. The study says that is where the benefit of the whole&lt;br&gt;
context-management apparatus lives.&lt;/p&gt;

&lt;p&gt;What does it cost per success at the floor, not per success on the benchmark? Your workload is not&lt;br&gt;
SWE-Bench. Take twenty of your own boring tasks, run them, and divide dollars by successes. Include&lt;br&gt;
the runs you killed. The study's own finding, that stronger models do better and cheaper with fewer&lt;br&gt;
tools, is a hypothesis you can test on your codebase in an afternoon.&lt;/p&gt;

&lt;p&gt;Which components would I turn off? This is the question the paper actually equips you to ask.&lt;br&gt;
If your model is strong, the plan step might be a cost centre. If it is bash-capable, the tool menu&lt;br&gt;
might be a drag. If your context budget is generous, the elaborate recoverable-summary system might&lt;br&gt;
be doing nothing. A harness with fewer components that you understand beats one with more that you&lt;br&gt;
do not, for the same reason a car you can service beats one you cannot.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipgfidliio4q4dj6yail.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fipgfidliio4q4dj6yail.png" alt="A benchmark measures the whole car once, at its ceiling. A buyer needs the floor, per component, on their own track." width="800" height="527"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;A benchmark measures the whole car once, at its ceiling. A buyer needs the floor, per component, on their own track.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The afternoon test
&lt;/h2&gt;

&lt;p&gt;Since the third question is the one that matters and the one a vendor cannot answer for you, here is&lt;br&gt;
the protocol we use, which costs an afternoon and a modest API bill.&lt;/p&gt;

&lt;p&gt;Pick twenty tasks from your own recent history. Not the interesting ones. The ones that came in as&lt;br&gt;
tickets and got done without anyone remembering them: a null check, a copy change, a small&lt;br&gt;
migration, a flaky test, a dependency bump that broke something. Ten of them should be the kind of&lt;br&gt;
task an intern would finish before lunch. Ten should be the kind that looks trivial and turns out to&lt;br&gt;
touch four files. Write each one down as a ticket, the way it was actually written, with the same&lt;br&gt;
missing context.&lt;/p&gt;

&lt;p&gt;Run each task through the candidate agent with its default harness, on a clean checkout, with a&lt;br&gt;
fixed budget of tokens or minutes, and walk away. Do not steer. Steering is what the demo does and&lt;br&gt;
it is what you will not have time to do at volume. When the budget runs out or the agent stops,&lt;br&gt;
record three things: did it produce a change a reviewer would accept, how many tokens did it use,&lt;br&gt;
and did it at any point loop, meaning repeat an action it had already taken with the same result.&lt;/p&gt;

&lt;p&gt;Then divide the money by the accepted changes. That number, dollars per accepted change on your own&lt;br&gt;
dull work, is the only benchmark that will predict your bill. Run the same twenty through a second&lt;br&gt;
candidate and you have a comparison that no leaderboard offers. Run them again with the plan step&lt;br&gt;
disabled, or the tool menu replaced with a shell, and you have reproduced the study's method on the&lt;br&gt;
only codebase you care about. The loops column is the floor made visible; a candidate that looped&lt;br&gt;
on three of twenty will loop on fifteen percent of your work forever, and no ceiling justifies that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limit of the study, stated plainly
&lt;/h2&gt;

&lt;p&gt;It was run on a particular set of models that, as one commenter complained, did not include the&lt;br&gt;
current frontier or the strongest open-weight options. The definitions are looser than they should&lt;br&gt;
be. The benchmarks are the benchmarks, with all the training-set contamination questions those&lt;br&gt;
carry. None of that changes the shape of the result, which is that the harness is a first-class&lt;br&gt;
variable and the leaderboards treat it as noise.&lt;/p&gt;

&lt;p&gt;The engine matters. Buy a good one. But you are going to spend the next year driving the car, and&lt;br&gt;
the car is the part you were never shown.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2609.20804" rel="noopener noreferrer"&gt;An Empirical Study of Harness Design for Coding Agents (Fan et al., arXiv 2609.20804)&lt;/a&gt; (HN, 217 points, 18 Sep 2026; all study numbers below are from the abstract)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://news.ycombinator.com/item?id=49753878" rel="noopener noreferrer"&gt;HN thread on the study&lt;/a&gt; (the car analogy and the Claude Code todo-tools observation come from commenters there)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=iBrAWpjXNxs" rel="noopener noreferrer"&gt;Please stop using stupid models (t3.gg, video)&lt;/a&gt; (18 Sep 2026; the ceiling-and-floor framing)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=d3Gjq-BffuI" rel="noopener noreferrer"&gt;How are they losing so bad (ThePrimeagen, video)&lt;/a&gt; (15 Sep 2026; the 330-million-token session, as reported on stream)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/how-to-read-a-coding-agent-benchmark-without-getting-sold/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>The vibe-coding trap has a name, and the name is not "the model"</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:16:36 +0000</pubDate>
      <link>https://dev.to/levelbrook/the-vibe-coding-trap-has-a-name-and-the-name-is-not-the-model-3fbm</link>
      <guid>https://dev.to/levelbrook/the-vibe-coding-trap-has-a-name-and-the-name-is-not-the-model-3fbm</guid>
      <description>&lt;p&gt;&lt;em&gt;A programming language shipped this week with a 99 percent AI-written compiler and no mention of the forty-year-old field it reinvents. The failure was not generation quality. It was that building got cheaper than reading, and nobody put a gate between them.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A language, a proof, and a field nobody looked up
&lt;/h2&gt;

&lt;p&gt;Two stories ran side by side on Hacker News this week and they are the same story.&lt;/p&gt;

&lt;p&gt;The first is Bend 2, a language pitched for the AI coding era: humans write "laws", the AI writes&lt;br&gt;
implementations and proofs, and the compiler checks the proofs. It got six hundred points and a lot&lt;br&gt;
of admiration. Then Liam Powell wrote a response that got three hundred more, and his point was not&lt;br&gt;
that the language is bad. His point was that the demo on the home page takes 58 lines to state that&lt;br&gt;
a player can never touch the flag, 442 lines of AI-written proof to establish it, and that the&lt;br&gt;
phrase "formal verification" appears nowhere on the website or in the codebase. He then asked a&lt;br&gt;
model to redo the demo in SPARK, a language built for exactly this, with no further guidance, and&lt;br&gt;
it came back a fraction of the size. The README, a commenter noted, says the compiler is 99 percent&lt;br&gt;
AI-written and has not been fully audited.&lt;/p&gt;

&lt;p&gt;The second story is Dan Abramov's account of vibing a proof of a conjecture of Conway's with a&lt;br&gt;
model, over days, in a long transcript he published in full. It is a good post and an honest one.&lt;br&gt;
The most upvoted objection under it was a mathematician pointing to Gowers's essay from the same&lt;br&gt;
week on why he did not sign the Fields medallists' letter, and the older point Gowers has been&lt;br&gt;
making for twenty-five years: there is a difference between solving a problem and understanding a&lt;br&gt;
field, and the second is what makes the first mean anything.&lt;/p&gt;

&lt;p&gt;Powell names the mechanism precisely and we are going to steal his sentence: vibe coding makes it&lt;br&gt;
possible to build a substantial solution before learning enough about the problem to recognise that&lt;br&gt;
a much better solution exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is new
&lt;/h2&gt;

&lt;p&gt;It has always been possible to reinvent a field badly. Every senior engineer has watched a junior&lt;br&gt;
build a job queue in a spreadsheet. What is new is the ratio.&lt;/p&gt;

&lt;p&gt;For all of software's history, building was expensive relative to reading. Before you could produce&lt;br&gt;
442 lines of anything, you had spent enough hours inside the problem that you had, almost by&lt;br&gt;
accident, tripped over the prior art. You searched for the error message. You read the paper the&lt;br&gt;
library cited. You asked the person at the next desk, who said "oh, that's just a Bloom filter".&lt;br&gt;
The cost of building was a tax that paid for an education.&lt;/p&gt;

&lt;p&gt;That tax is gone. A model will produce the 442 lines in the time it takes to make coffee, and it&lt;br&gt;
will produce them competently enough that they work, and working code is the most persuasive&lt;br&gt;
argument in the world against going back to read. Nothing in the loop ever forces you to discover&lt;br&gt;
that the field exists. The model will not volunteer it unless you ask, and you do not know to ask,&lt;br&gt;
because the whole point is that you do not know the field exists.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fee39qaqtd73vu41zhu6n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fee39qaqtd73vu41zhu6n.png" alt="The old cost of building bought an education for free. The new cost does not. Time axis illustrative; the shape is the point." width="800" height="630"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The old cost of building bought an education for free. The new cost does not. Time axis illustrative; the shape is the point.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Bend story is the pure case because a language is the most expensive thing you can build and&lt;br&gt;
formal verification is one of the best-documented fields in computer science. If it can happen&lt;br&gt;
there, at that scale, with that much talent, it is happening in your codebase this week at a&lt;br&gt;
smaller scale where nobody will write a blog post about it. The agent that built your rate limiter&lt;br&gt;
from scratch instead of reading the one in your framework. The retry logic that reinvented&lt;br&gt;
exponential backoff without the jitter. The custom auth layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gate goes before the build, not after
&lt;/h2&gt;

&lt;p&gt;The instinct is to fix this with review, and review does catch some of it. But review happens after&lt;br&gt;
the 442 lines exist, when the sunk cost is already arguing for them, and the reviewer usually&lt;br&gt;
shares the author's blind spot. The place to put the gate is the fifteen minutes before anything&lt;br&gt;
is built.&lt;/p&gt;

&lt;p&gt;We run something we call the prior-art pass, and it is embarrassingly simple. Before an agent is&lt;br&gt;
allowed to build anything with a name, it has to answer four questions in writing and a person has&lt;br&gt;
to read the answers. What is this problem called by people who study it? What do they already use?&lt;br&gt;
Why does the existing thing not work here? What is the smallest version of this we could build on&lt;br&gt;
top of the existing thing instead?&lt;/p&gt;

&lt;p&gt;The model is extremely good at answering these questions. It has read the field. It will tell you&lt;br&gt;
about SPARK and Dafny and Lean and TLA+ in one paragraph if you ask it to, and it will tell you what&lt;br&gt;
each is for. The trick is that somebody has to ask before the build starts, and that somebody has to&lt;br&gt;
be willing to hear "this already exists" as good news rather than as an obstacle to the thing they&lt;br&gt;
were excited to make.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F062xjgkdwvzta44li614.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F062xjgkdwvzta44li614.png" alt="The prior-art pass: four written answers, one human read, before an agent may build anything with a name." width="800" height="727"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The prior-art pass: four written answers, one human read, before an agent may build anything with a name.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Liam Nugent's piece from the same week, on why the most important product decision is what you do&lt;br&gt;
not build, makes the organisational version of the same point. Nobody gets promoted for deleting&lt;br&gt;
things. Those who create and launch are the ones rewarded. The models have made creating and&lt;br&gt;
launching nearly free, which means the incentive that was already skewed towards building is now&lt;br&gt;
skewed by another order of magnitude, and the only counterweight is a deliberate, slightly&lt;br&gt;
unpopular gate that asks "does this need to exist" before the exciting part starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pass in practice
&lt;/h2&gt;

&lt;p&gt;A composite from our own work, because the pass sounds like a platitude until you watch it fire.&lt;/p&gt;

&lt;p&gt;A team wanted a service that deduplicated inbound customer records, which arrive from four systems&lt;br&gt;
with inconsistent formatting, so that the same person is not created four times. An agent, asked&lt;br&gt;
directly, would have built it in an afternoon: normalise the fields, hash them, compare. The&lt;br&gt;
prior-art pass asked the four questions first, and the agent's written answers were, in order: this&lt;br&gt;
is called entity resolution or record linkage; the standard approaches are probabilistic matching&lt;br&gt;
in the Fellegi-Sunter family and there are mature libraries in every major language; the naive&lt;br&gt;
hash-and-compare approach fails on exactly the inconsistent formatting the team has, because it&lt;br&gt;
treats a transposed digit as a different person; the smallest version is to run an existing&lt;br&gt;
library with blocking on postcode and hand the ambiguous pairs to a human.&lt;/p&gt;

&lt;p&gt;Fifteen minutes. The person reading the answers had never heard the phrase "record linkage". The&lt;br&gt;
team built the small version on top of the library, spent the afternoon they saved on the human&lt;br&gt;
review queue for ambiguous pairs, and did not spend the following quarter discovering, one support&lt;br&gt;
ticket at a time, every way in which the hash approach silently merges or splits real people.&lt;/p&gt;

&lt;p&gt;The point is not that the agent knew about record linkage; of course it did. The point is that&lt;br&gt;
nobody would have asked, because the task looked simple and the build was cheap, and the cost of&lt;br&gt;
the field not being known would have been paid by customers over months rather than by the team in&lt;br&gt;
one visible failure. That is the shape of the trap every time. The wrong build does not fail. It&lt;br&gt;
works, slightly worse than the right build, forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this is wrong
&lt;/h2&gt;

&lt;p&gt;The honest caveat is that the prior-art pass has a failure mode of its own: it can become an excuse&lt;br&gt;
never to build anything new, and some things genuinely are new. Bend's author may well have&lt;br&gt;
considered SPARK and rejected it for reasons that are not on the website. Abramov's proof may be&lt;br&gt;
a real contribution even if he cannot yet situate it in the field. The gate is not "never build".&lt;br&gt;
The gate is "never build without having looked", and the output of looking is sometimes "nothing&lt;br&gt;
here fits, build it, and say in the README what you looked at and why it did not fit". That&lt;br&gt;
sentence in a README is worth more than the 442 lines under it, because it tells the next reader&lt;br&gt;
that the author knew where they were standing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do on Monday
&lt;/h2&gt;

&lt;p&gt;Find the three most recent things your team or your agents built that have a name. A service, a&lt;br&gt;
library, an internal tool, a pattern with a wiki page. For each one, ask the four questions now,&lt;br&gt;
after the fact. Do it with a model; it will take ten minutes each. You will find that at least one of&lt;br&gt;
the three is a smaller, worse version of something that already existed, and you will feel the&lt;br&gt;
thing Powell's post is about, which is not embarrassment exactly. It is the realisation that the&lt;br&gt;
cost of not knowing has gone up precisely because the cost of building has gone down.&lt;/p&gt;

&lt;p&gt;Then put the pass in front of the next build. Fifteen minutes, four questions, one reader. The&lt;br&gt;
model will do most of the work. The only thing it cannot do is want to know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blog.liampwll.com/posts/bend_vibe_coding/" rel="noopener noreferrer"&gt;Bend 2 and the Vibe-Coding Trap (Liam Powell)&lt;/a&gt; (HN, 323 points, 18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://bend-lang.com/" rel="noopener noreferrer"&gt;Bend, a language that blocks AI mistakes via proof and runs on GPUs&lt;/a&gt; (HN, 603 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://overreacted.io/how-i-vibed-a-proof-of-conways-conjecture/" rel="noopener noreferrer"&gt;I vibed a proof of Conway's conjecture (Dan Abramov)&lt;/a&gt; (HN, 259 points, 18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://liamnugent.me/posts/what-you-dont-build/" rel="noopener noreferrer"&gt;The most important product decision is what you don't build (Liam Nugent)&lt;/a&gt; (HN, 156 points, 17 Sep 2026)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/the-vibe-coding-trap-is-not-the-model/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>Nobody gets fired for not using AI. People get fired for what they approved.</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:14:05 +0000</pubDate>
      <link>https://dev.to/levelbrook/nobody-gets-fired-for-not-using-ai-people-get-fired-for-what-they-approved-13p</link>
      <guid>https://dev.to/levelbrook/nobody-gets-fired-for-not-using-ai-people-get-fired-for-what-they-approved-13p</guid>
      <description>&lt;p&gt;&lt;em&gt;The anxiety of the last month has been about falling behind. The actual career risk sits in the other direction: the moment a model's output carries your signature. The military found that out this week. Law firms are about to.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The two fears, and which one is real
&lt;/h2&gt;

&lt;p&gt;There are two fears in circulation right now and they point in opposite directions.&lt;/p&gt;

&lt;p&gt;The first is the one Fireship and Theo and every LinkedIn feed are selling: you are behind. The&lt;br&gt;
models moved again, the people using them properly are producing five times what you are, and if&lt;br&gt;
you do not get on this you will be the person the organisation quietly stops giving interesting work&lt;br&gt;
to. There is a version of this fear that is accurate and we will get to it. But as a career risk it&lt;br&gt;
is slow. Nobody gets walked out on a Tuesday for having been sceptical of a tool.&lt;/p&gt;

&lt;p&gt;The second fear is the one that actually ends careers, and almost nobody is articulating it because&lt;br&gt;
it is less flattering to the tools. On Thursday CNN reported that the US military had what it called&lt;br&gt;
a close call after an AI-generated intelligence report contained hallucinated content that made it&lt;br&gt;
some distance up the chain before somebody caught it. We do not know the details beyond what was&lt;br&gt;
reported and we are not going to pretend to. But read the top comment on the Hacker News thread,&lt;br&gt;
because it is the whole essay in three sentences: AI is not responsible. People are responsible.&lt;br&gt;
As soon as people choose to remove their own accountability, that is when the bad stuff happens.&lt;/p&gt;

&lt;p&gt;The same week, OpenAI launched Astra for Law, a model tuned for legal work. The launch post, as&lt;br&gt;
several commenters noted, does not mention hallucination once. The most upvoted question under it&lt;br&gt;
was not about capability. It was: if I use this to write my contracts, who is getting sued when it&lt;br&gt;
is wrong? That question has an answer, and the answer is you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the signature lands
&lt;/h2&gt;

&lt;p&gt;Every organisation has a small number of places where a person converts information into a&lt;br&gt;
commitment. A lawyer signs the brief. An officer signs the assessment. A reviewer approves the&lt;br&gt;
change. A controller approves the payment. A doctor signs the order. These are the seats where the&lt;br&gt;
organisation has decided, in advance, that a named human stands behind the outcome.&lt;/p&gt;

&lt;p&gt;For the whole history of office work, the person in the seat also produced most of the information&lt;br&gt;
that went into the decision, or supervised the junior who did. That has quietly stopped being true.&lt;br&gt;
The brief was drafted by a model. The assessment was summarised by a model. The pull request was&lt;br&gt;
written by a model. The person in the seat is now signing for work they did not do, produced by a&lt;br&gt;
process they did not watch, at a volume that does not leave time to check.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F03uosdopi01hxhp4shjt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F03uosdopi01hxhp4shjt.png" alt="The approval seat has not moved. Everything feeding it has. The seat still carries the name, the liability and the memory." width="800" height="637"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The approval seat has not moved. Everything feeding it has. The seat still carries the name, the liability and the memory.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is the structural fact under the anxiety. The tools did not remove the seat. They removed the&lt;br&gt;
slack around the seat, the four hours of drafting during which a person used to notice things. And&lt;br&gt;
in a lot of organisations they removed it faster than anybody redesigned what the person in the&lt;br&gt;
seat is supposed to do with their eight seconds.&lt;/p&gt;

&lt;p&gt;Martin Fowler's short piece this week, the one titled simply that he does not like LLMs, is being&lt;br&gt;
read as a grumpy-elder post. It is more useful than that. His complaint is that they confidently&lt;br&gt;
bullshit him, often while giving useful answers, with the same assurance either way. That is not a&lt;br&gt;
complaint about capability. It is a description of the exact property that makes the approval seat&lt;br&gt;
dangerous: the output does not carry a signal about its own reliability. A junior who is unsure&lt;br&gt;
looks unsure. A model that is unsure looks like a model that is sure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the people who survive this will have done
&lt;/h2&gt;

&lt;p&gt;We build systems where a model is allowed to act, inside a gate a person controls, and the entire&lt;br&gt;
design problem is the seat. So here is what we have learned about occupying one well, offered to&lt;br&gt;
the lawyer and the officer and the reviewer as much as to the engineer.&lt;/p&gt;

&lt;p&gt;The first thing is to refuse to sign for what you cannot see. That sounds obvious and it is&lt;br&gt;
violated constantly, because the tools present a finished artefact and hide the process. Insist on&lt;br&gt;
the provenance: what was the model given, what did it produce, what did it check, what did it say it&lt;br&gt;
was unsure about. If the tool cannot show you that, the tool is asking you to carry liability for a&lt;br&gt;
process it will not disclose. The correct response to that is not faster reading.&lt;/p&gt;

&lt;p&gt;The second is to make the model's abstentions visible and to treat them as the most valuable line&lt;br&gt;
in the output. A model that has been built to say "I could not verify this citation" or "this&lt;br&gt;
clause is outside the examples I was given" hands the seat exactly the information it needs to spend&lt;br&gt;
its attention well. A model that has been built to always answer hands the seat nothing. Most&lt;br&gt;
vendors ship the second kind because it demos better. Buy the first kind, or add the gate yourself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frtb74fvxq760vl0eahhw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frtb74fvxq760vl0eahhw.png" alt="Two ways to build the seat. The left one is what most tools ship. The right one is the only one worth signing for." width="800" height="571"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Two ways to build the seat. The left one is what most tools ship. The right one is the only one worth signing for.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The third is the record. When something goes wrong, and it will, the difference between a career&lt;br&gt;
ending and a process improving is whether there is a log that shows what the seat was shown, what&lt;br&gt;
it decided, and why. Approvals without a record are how organisations find scapegoats. Approvals&lt;br&gt;
with one are how they find bugs. If you are the person in the seat, you want the record more than&lt;br&gt;
your employer does.&lt;/p&gt;

&lt;p&gt;The fourth is the one that answers the first fear honestly. You do need to understand the tools.&lt;br&gt;
Not to produce more, though you will, but because the person who understands how a model fails is&lt;br&gt;
the only person who can occupy the seat competently. A reviewer who has never watched an agent&lt;br&gt;
confidently rewrite a retry loop it did not understand cannot review agent-written code. A lawyer&lt;br&gt;
who has never seen a model invent a citation cannot sign a model-drafted brief. Falling behind is a&lt;br&gt;
real risk, and the reason is not productivity. The reason is that the seat is about to be occupied&lt;br&gt;
exclusively by people who know what they are looking at.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like when it works
&lt;/h2&gt;

&lt;p&gt;A composite from the systems we run, with the details changed. A finance team processes vendor&lt;br&gt;
invoices through a model that reads the document, matches it to a purchase order, and proposes a&lt;br&gt;
payment. The person in the seat sees, for each proposal, four lines: the amount, the matched order,&lt;br&gt;
the confidence of the match, and a list of anything the model could not reconcile. Most proposals&lt;br&gt;
show an empty list and take three seconds. A handful show one line, "invoice quantity 12, order&lt;br&gt;
quantity 10", and take thirty. Once or twice a day one shows "no matching order found" and gets&lt;br&gt;
routed to a human who investigates.&lt;/p&gt;

&lt;p&gt;The team approves several hundred payments a day this way, which is roughly six times what it did&lt;br&gt;
by hand, and the person in the seat is not faster at reading invoices. They never read the invoice.&lt;br&gt;
They read the abstentions. The model was built to say what it did not know, the seat was built to&lt;br&gt;
show only that, and the record of every approval, with what the seat saw, is kept for the auditor.&lt;/p&gt;

&lt;p&gt;When a duplicate payment went out in the spring, because a vendor had sent the same invoice under&lt;br&gt;
two numbers, the investigation took twenty minutes. The record showed the seat had been shown two&lt;br&gt;
proposals, each with an empty abstention list, three days apart. Nobody was blamed, because nobody&lt;br&gt;
had been shown anything that should have stopped them. The model was changed to flag matching&lt;br&gt;
amounts to the same vendor inside thirty days. That is what the seat is for: not to be infallible,&lt;br&gt;
but to make failure legible enough that the next one does not happen.&lt;/p&gt;

&lt;p&gt;Compare that with the version most organisations are building, where the same person is handed a&lt;br&gt;
finished payment run and asked to approve it by five. That seat has the same liability and none of&lt;br&gt;
the information, and when the duplicate goes out the investigation finds a name rather than a bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable version for managers
&lt;/h2&gt;

&lt;p&gt;If you run a team, the obligation runs the other way. You have almost certainly increased the volume&lt;br&gt;
flowing into your approval seats this year without increasing the time, the tooling, or the&lt;br&gt;
authority of the people in them. You may have done it with a slide that said something about&lt;br&gt;
productivity. When one of those seats signs something that goes wrong, the organisation will look at&lt;br&gt;
the name on the signature, not at the slide.&lt;/p&gt;

&lt;p&gt;The remedy is not to slow down. It is to treat the seat as the product. Decide, explicitly, which&lt;br&gt;
decisions a model may take alone, which need a person, and which need a senior person. Write it&lt;br&gt;
down. Give the seats the provenance, the abstentions and the record described above. Cap the volume&lt;br&gt;
per seat per day, because a person who approves three hundred things approves none of them. Then,&lt;br&gt;
and only then, let the models run.&lt;/p&gt;

&lt;p&gt;The last month of discourse has been about who is behind. The next year is going to be about who&lt;br&gt;
signed. Make sure that when it is your name, you were shown enough to deserve it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.cnn.com/2026/09/18/politics/us-military-ai-false-intelligence" rel="noopener noreferrer"&gt;US military had close call after using AI for hallucinated intelligence report (CNN)&lt;/a&gt; (HN, 499 points, 18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://openai.com/index/astra-for-law/" rel="noopener noreferrer"&gt;Astra for Law (OpenAI)&lt;/a&gt; (HN, 578 points and 678 comments, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://martinfowler.com/articles/2026-dont-like-llms.html" rel="noopener noreferrer"&gt;I Don't Like LLMs (Martin Fowler)&lt;/a&gt; (HN, 237 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/tools/approval-queue/"&gt;Levelbrook: the approval queue demo&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/nobody-gets-fired-for-not-using-ai/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>management</category>
      <category>ai</category>
    </item>
    <item>
      <title>Your engineers are not slow. Your review queue is.</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:11:34 +0000</pubDate>
      <link>https://dev.to/levelbrook/your-engineers-are-not-slow-your-review-queue-is-948</link>
      <guid>https://dev.to/levelbrook/your-engineers-are-not-slow-your-review-queue-is-948</guid>
      <description>&lt;p&gt;&lt;em&gt;Every team that bought coding agents this year got the same result: the generation side of the shop sped up ten times and the shipping rate barely moved. The constraint moved to the one seat nobody re-engineered, and it is the seat with a person in it.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The generation side won. Nobody told the review side.
&lt;/h2&gt;

&lt;p&gt;Start with the concession, because it is true and it is the whole reason this matters.&lt;/p&gt;

&lt;p&gt;Coding agents work. The team at detail.dev put it as plainly as anyone this week: agents can&lt;br&gt;
oneshot games that are actually fun, and with the right guardrails they execute migrations and&lt;br&gt;
language rewrites in complex codebases that would have been a quarter's work two years ago. Theo&lt;br&gt;
Browne spent an entire video this week arguing that if you cannot tell the difference between this&lt;br&gt;
year's frontier models and last year's, the problem is your prompting, and he is mostly right. The&lt;br&gt;
ceiling on what a single engineer can emit in a day has gone up by an amount that is genuinely hard&lt;br&gt;
to describe to someone who has not sat in front of it.&lt;/p&gt;

&lt;p&gt;And yet the same detail.dev post opens with the sentence every engineering manager has been&lt;br&gt;
avoiding saying out loud: a lot of orgs spent the first half of this year offloading as much work as&lt;br&gt;
possible to armies of agents, and the results have been disappointing. Mountains of dubious code, no&lt;br&gt;
tsunami of incredible software. They call it the trough of disillusionment. We would call it&lt;br&gt;
something more boring. The factory got a faster machine at one station and nothing else changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the work actually went
&lt;/h2&gt;

&lt;p&gt;Here is the shape of a change to production software in 2024, roughly, in the units that matter:&lt;br&gt;
minutes of a human's attention.&lt;/p&gt;

&lt;p&gt;Somebody spends four hours writing it. Somebody else spends twenty minutes reviewing it. A machine&lt;br&gt;
spends six minutes testing it. Somebody spends five minutes deploying it. The human write step&lt;br&gt;
dominates, so every tool of the last fifteen years attacked the write step: better editors, better&lt;br&gt;
languages, better frameworks, and now agents.&lt;/p&gt;

&lt;p&gt;Now the write step takes twelve minutes. Not four hours. The agent drafts, the engineer steers, and a&lt;br&gt;
change that used to be an afternoon is a coffee. Which means the engineer produces, on a good day,&lt;br&gt;
somewhere between five and fifteen times as many changes as before. Every one of them still needs&lt;br&gt;
twenty minutes of somebody else's attention before it ships.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl26r4wf47ssx2q06cugj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl26r4wf47ssx2q06cugj.png" alt="The write step shrank by an order of magnitude; the review step did not move. Minutes are illustrative for a mid-sized change; the ratio is what every team we have talked to describes." width="800" height="766"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The write step shrank by an order of magnitude; the review step did not move. Minutes are illustrative for a mid-sized change; the ratio is what every team we have talked to describes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Do the arithmetic once and it stops being a vibe. If the review capacity of a six-person team is&lt;br&gt;
fixed at, say, forty reviews a day, then the team ships forty changes a day whether the engineers&lt;br&gt;
produce forty or four hundred. The other three hundred and sixty sit in a queue. Engineers notice the&lt;br&gt;
queue, so they stop producing, or they start rubber-stamping each other, or they merge their own&lt;br&gt;
work at six in the evening when nobody is looking. Every one of those behaviours shows up in the&lt;br&gt;
incident log within a month.&lt;/p&gt;

&lt;p&gt;This is why the productivity numbers are so confusing. The individual is faster. The system is&lt;br&gt;
producing the same amount of finished software as before, with worse review. Both things are true.&lt;/p&gt;

&lt;h2&gt;
  
  
  The seat nobody re-engineered
&lt;/h2&gt;

&lt;p&gt;Look at what the tools industry did in response. It built more agents. Agents to write the tests,&lt;br&gt;
agents to review the pull request, agents to fix the review comments, agents to review the fix.&lt;br&gt;
Some of it is good. Fireship's sponsor this week, Macroscope, says its review tool now auto-approves&lt;br&gt;
forty percent of pull requests across its customers, which is a vendor claim from a sponsor segment&lt;br&gt;
and we would treat it as exactly that, but it tells you where the market thinks the money is. The&lt;br&gt;
market thinks the answer to a review bottleneck is to remove the reviewer.&lt;/p&gt;

&lt;p&gt;We think that is the wrong objective function, for the same reason it was wrong in the back office.&lt;br&gt;
The human in the review seat is the accountability. When the change breaks production, somebody&lt;br&gt;
approved it, and that somebody has a name and a manager and a memory of what they were told. An&lt;br&gt;
auto-approval does not have a memory. It has a log line. You can automate the reading of a diff;&lt;br&gt;
you cannot automate the standing-behind of one.&lt;/p&gt;

&lt;p&gt;So the question is the one detail.dev asks at the end of their piece and then does not quite answer:&lt;br&gt;
when the software mostly drives itself, what do the engineers do? Our answer is not glamorous. They&lt;br&gt;
review. And the entire engineering problem of the next two years is making that review seat fast&lt;br&gt;
enough to keep up with the machines feeding it, without turning it into a stamp.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a fast review seat looks like
&lt;/h2&gt;

&lt;p&gt;We have written before about the eight-second decision, and the number is not rhetorical. It is&lt;br&gt;
roughly the time a competent reviewer needs to accept or reject a change when everything they need&lt;br&gt;
to know is in front of them and nothing they do not need is. Almost no review tool is built for that&lt;br&gt;
number. They are built for the twenty-minute review, which was designed for human-written code where&lt;br&gt;
the intent had to be reverse-engineered from the diff.&lt;/p&gt;

&lt;p&gt;Agent-written code has a property human-written code never had: the intent already exists in&lt;br&gt;
writing, because somebody typed it into a prompt. The specification, the plan, the reason for every&lt;br&gt;
choice, the tests it ran, the things it decided not to do. All of that is sitting in a transcript&lt;br&gt;
that the review tool throws away. The review seat is slow because it is being asked to rediscover&lt;br&gt;
information the system already had.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyo4uqmgspc43sl7hql6k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyo4uqmgspc43sl7hql6k.png" alt="What the reviewer needs in front of them for an eight-second decision, and where each piece already exists today (nowhere in the pull request)." width="800" height="270"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What the reviewer needs in front of them for an eight-second decision, and where each piece already exists today (nowhere in the pull request).&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Five things. The request, the plan, the verification, the blast radius, and the list of things the&lt;br&gt;
agent decided it was not sure about. That last one is the one nobody surfaces and the one that&lt;br&gt;
makes the decision fast, because a reviewer who can see "the agent was unsure about the retry logic&lt;br&gt;
and left it as before" knows exactly where to spend their eight seconds.&lt;/p&gt;

&lt;p&gt;The harness study that hit Hacker News this week is interesting in this light. Nine researchers ran&lt;br&gt;
176 matched configurations across four models on SWE-Bench Verified and Terminal-Bench, varying&lt;br&gt;
planning, action space and context management. One of their findings is that for stronger models,&lt;br&gt;
explicit planning stopped improving accuracy and became mainly a cost saver. Read as a review&lt;br&gt;
problem rather than a benchmark problem, that says the plan is cheap to produce and does not hurt&lt;br&gt;
the agent. Which means there is no excuse for it not being attached to the pull request, because it&lt;br&gt;
is the single most useful artefact a reviewer could have and the model will write it for nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do on Monday
&lt;/h2&gt;

&lt;p&gt;Measure the queue. Not the cycle time, which averages away the problem, but the number of changes&lt;br&gt;
waiting for a human and how long the oldest one has waited. If that number is growing week over week,&lt;br&gt;
you have the disease and no amount of model upgrades will treat it.&lt;/p&gt;

&lt;p&gt;Then re-engineer the seat rather than removing it. Attach the prompt and the plan to the pull&lt;br&gt;
request, automatically, as the description. Make the agent state what it verified and how, in a&lt;br&gt;
fixed format, at the bottom. Make it list what it was unsure about. Route by blast radius: a change&lt;br&gt;
that touches one file and no data can go to a fast lane with a fast reviewer; a change that touches&lt;br&gt;
billing goes to a slow lane with a senior one. Give reviewers a budget of decisions per day rather&lt;br&gt;
than a queue of infinite length, and watch what happens to the quality of the decisions.&lt;/p&gt;

&lt;p&gt;One more thing worth stealing from manufacturing, since the factory metaphor is doing so much work&lt;br&gt;
here. A line with a bottleneck station is run at the pace of the bottleneck, on purpose, and the&lt;br&gt;
upstream stations are told to stop rather than pile up inventory. Software teams do the opposite:&lt;br&gt;
they celebrate the pile. A hundred open pull requests is not throughput. It is inventory, and&lt;br&gt;
inventory decays, because the codebase underneath it keeps moving and every day a change waits is a&lt;br&gt;
day closer to a merge conflict and a re-review. Cap the queue. Let the agents idle. It feels wrong&lt;br&gt;
for about a week.&lt;/p&gt;

&lt;p&gt;And resist the auto-approve until you have done all of that. Not because the tools are bad but&lt;br&gt;
because auto-approving forty percent of changes into a system with no fast human lane for the other&lt;br&gt;
sixty is how you get a queue of the hardest sixty percent with the least context, reviewed by the&lt;br&gt;
most tired people. The machines will keep getting faster on a schedule you do not control. The seat&lt;br&gt;
is the only part of the pipeline you actually own.&lt;/p&gt;

&lt;p&gt;The engineers were never the slow part. The place where a person has to say yes is the slow part,&lt;br&gt;
and it was always going to be, and the job now is to make yes cheap without making it meaningless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://blog.detail.dev/posts/towards-self-driving-codebases" rel="noopener noreferrer"&gt;Towards Self-Driving Codebases (detail.dev)&lt;/a&gt; (HN, 119 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=iBrAWpjXNxs" rel="noopener noreferrer"&gt;Please stop using stupid models (t3.gg, video)&lt;/a&gt; (18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=7r4ikZHm9AI" rel="noopener noreferrer"&gt;Anthropic researchers are quitting... and now we know why (Fireship, video)&lt;/a&gt; (15 Sep 2026; the Macroscope figure is from the sponsor segment, so a vendor claim)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://arxiv.org/abs/2609.20804" rel="noopener noreferrer"&gt;An Empirical Study of Harness Design for Coding Agents (arXiv 2609.20804)&lt;/a&gt; (HN, 217 points, 18 Sep 2026)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://dev.to/playbook/"&gt;Levelbrook: the eight-second review (earlier playbook piece)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/your-engineers-are-not-slow-your-review-queue-is/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>ai</category>
      <category>management</category>
    </item>
    <item>
      <title>You are not falling behind. You are watching the wrong scoreboard.</title>
      <dc:creator>Levelbrook Consulting</dc:creator>
      <pubDate>Mon, 21 Sep 2026 14:05:45 +0000</pubDate>
      <link>https://dev.to/levelbrook/you-are-not-falling-behind-you-are-watching-the-wrong-scoreboard-46ke</link>
      <guid>https://dev.to/levelbrook/you-are-not-falling-behind-you-are-watching-the-wrong-scoreboard-46ke</guid>
      <description>&lt;p&gt;&lt;em&gt;The model leaderboard reshuffles every six weeks and everything you learn about a specific model has a half-life of months. Five things do not decay, and they are the same five that the people running the swarms say they still need humans for. This is a list for the engineer who is tired.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The scoreboard that changes every six weeks
&lt;/h2&gt;

&lt;p&gt;Here is what the last ten days sounded like if you are an engineer with a job and a family and a&lt;br&gt;
finite amount of evening.&lt;/p&gt;

&lt;p&gt;Theo Browne released a video saying that if you cannot feel the difference between this&lt;br&gt;
generation's frontier models and the last one, you suck at prompting, and he meant it kindly but he&lt;br&gt;
meant it. ThePrimeagen spent an episode on how a company that invented the transformer and the&lt;br&gt;
tensor processing unit is, on the coding leaderboards he showed, being beaten by a lab with a few&lt;br&gt;
hundred employees, and then left one of its models on a trivial bug for forty minutes and watched&lt;br&gt;
it spend 330 million tokens reading the same file. AI Explained walked through six axes on which&lt;br&gt;
the labs say capability will keep improving, none of them near saturation, and quoted a researcher&lt;br&gt;
saying there is a large gap between how fast progress looks from inside and from outside. A model&lt;br&gt;
called Astra that could do things nothing before it could. Another called Fable that was the safe&lt;br&gt;
choice three weeks ago. A Chinese open-weight model that is months old and beating both on some&lt;br&gt;
table.&lt;/p&gt;

&lt;p&gt;If you tried to keep up with that, you did not sleep and you learned nothing durable, because&lt;br&gt;
almost every specific fact in the paragraph above will be wrong by November. That is not a&lt;br&gt;
prediction about any particular model. It is a description of the scoreboard. It reshuffles every&lt;br&gt;
six to eight weeks and it has done so for two years and every lab in the AI Explained video says it&lt;br&gt;
will keep doing so.&lt;/p&gt;

&lt;p&gt;So the anxiety is real and the scoreboard is real and the two have almost nothing to do with each&lt;br&gt;
other. The question worth an evening is which things you can learn now that will still be true when&lt;br&gt;
the scoreboard has reshuffled four more times.&lt;/p&gt;

&lt;h2&gt;
  
  
  The half-life of what you know
&lt;/h2&gt;

&lt;p&gt;Sort what an engineer learns about these tools by how long it stays true.&lt;/p&gt;

&lt;p&gt;At the bottom, decaying in weeks: which model is best at which task, which one reads the same file&lt;br&gt;
twenty-six times, which one is over-eager and which one is lazy, the prompt phrasing that makes a&lt;br&gt;
particular model stop apologising. This is what most of the content is about, because it changes&lt;br&gt;
constantly and change is what content is made of. Learning it is not worthless. Learning it is&lt;br&gt;
maintenance, like knowing this quarter's prices.&lt;/p&gt;

&lt;p&gt;In the middle, decaying in months: the harness. Which tool, which agent runtime, which instruction&lt;br&gt;
file format, which permission model. The study we wrote about earlier this week found that the&lt;br&gt;
harness is a first-class variable, more important than the leaderboards admit, and that is true. It&lt;br&gt;
is also true that the harnesses are being rewritten as fast as the models, and that the changelog&lt;br&gt;
entry that got seven hundred points this week was about one tool starting to read another tool's&lt;br&gt;
config file.&lt;/p&gt;

&lt;p&gt;At the top, not decaying at all as far as anyone can tell: five things. We will name them and then&lt;br&gt;
argue for each one, because the argument is where the reassurance actually lives.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F71opqu7bh97ihcouu2xd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F71opqu7bh97ihcouu2xd.png" alt="What an engineer learns about AI tools, sorted by how long it stays true. Almost all the content is about the bottom row. Durations illustrative." width="800" height="486"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;What an engineer learns about AI tools, sorted by how long it stays true. Almost all the content is about the bottom row. Durations illustrative.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The five things
&lt;/h2&gt;

&lt;p&gt;Specification. The single skill that has appreciated most in two years is the ability to say,&lt;br&gt;
precisely and in writing, what a system should do, including what it should refuse to do and every&lt;br&gt;
exception. This used to be a skill people apologised for having, because the code was the&lt;br&gt;
specification and writing it down twice was waste. Now the specification is the input to the&lt;br&gt;
machine that writes the code, and the quality of the output is bounded by the quality of that&lt;br&gt;
input in a way that no model upgrade changes. Every team that got disappointing results from agents&lt;br&gt;
this year got them from underspecified tasks. The detail.dev post says the most valuable engineering&lt;br&gt;
work is going to be having good ideas, and an idea that cannot be specified is not yet an idea.&lt;/p&gt;

&lt;p&gt;Verification. Knowing whether a thing works, and being able to prove it to someone else, is the&lt;br&gt;
skill the machines are worst at and the one they generate the most demand for. Every agent-written&lt;br&gt;
change needs someone who can say what test would catch the failure, whether the test that was&lt;br&gt;
written is that test, and whether the green result means what it appears to mean. Dan Luu's essay&lt;br&gt;
this week is on the front page for a reason: there is no point at which turning your brain off&lt;br&gt;
works, and verification is the name for the brain being on. This is also, not coincidentally, the&lt;br&gt;
skill Theo was actually describing when he said people prompt badly. The people who get good&lt;br&gt;
results from strong models are the people who can tell when the result is bad.&lt;/p&gt;

&lt;p&gt;Judgement. Which of the forty things the model could do is the one worth doing. Which of the&lt;br&gt;
three approaches it offered is the one that will still be fine in a year. When to stop. When the&lt;br&gt;
obvious solution is obvious because it is wrong. Every essay on this site comes back to this because&lt;br&gt;
every failure we have watched comes back to it: the model did something competent that nobody&lt;br&gt;
should have asked for. Judgement is slow to build, does not transfer from a video, and is the entire&lt;br&gt;
reason a senior engineer is paid more than a junior one with the same tools.&lt;/p&gt;

&lt;p&gt;Domain intimacy. Knowing the business, the customers, the forty exceptions and why the freight&lt;br&gt;
claim in 2023 changed how one account is handled. We have argued at length that companies do not&lt;br&gt;
have processes, they have habits, and that the habits live in heads. The heads are the moat. A&lt;br&gt;
model with the rumour builds the obvious version of your product; the person who knows why the&lt;br&gt;
obvious version fails for the second-largest customer is the person the swarm cannot replace, and&lt;br&gt;
that person is more valuable this year than last, not less.&lt;/p&gt;

&lt;p&gt;The seat. The ability to occupy the approval seat well: to look at a model's output, know where to&lt;br&gt;
spend eight seconds, sign for it, and stand behind the signature. This is the one that combines the&lt;br&gt;
other four and it is the one the organisation will pay for when the volume of model output has&lt;br&gt;
outrun everyone's ability to read it. It is also the one nobody teaches, because until eighteen&lt;br&gt;
months ago the person in the seat had also done the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this means for the evening
&lt;/h2&gt;

&lt;p&gt;Spend it differently.&lt;/p&gt;

&lt;p&gt;Stop trying to keep up with the scoreboard and start reading it the way you read exchange rates:&lt;br&gt;
glance, note the direction, move on. Pick one strong model and one harness and learn them properly&lt;br&gt;
for a quarter rather than four of each badly. Theo's floor-and-ceiling point is correct and it is&lt;br&gt;
also an argument for depth: you learn where a model's floor is by using it on dull tasks for weeks,&lt;br&gt;
not by watching someone else's demo of its ceiling.&lt;/p&gt;

&lt;p&gt;Then spend the real time on the five. Write the specification for the next thing your team builds&lt;br&gt;
before anyone, human or model, writes a line, and notice how much you did not know. Write the test&lt;br&gt;
that would catch the failure before you look at the agent's tests. When an agent offers three&lt;br&gt;
approaches, write down why you picked one, and read your reasons back in a month. Learn the part of&lt;br&gt;
the business your team pretends is simple. Volunteer for the review seat that everyone else is&lt;br&gt;
avoiding because the queue is long, and get fast at it, because that queue is the most important&lt;br&gt;
place in the company and almost nobody wants to sit there.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft4buynuctadvhg58uw8m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft4buynuctadvhg58uw8m.png" alt="The trade an engineer can make this quarter. The left column is what the content wants you to do; the right is what compounds." width="799" height="529"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The trade an engineer can make this quarter. The left column is what the content wants you to do; the right is what compounds.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest caveat
&lt;/h2&gt;

&lt;p&gt;It is possible to hide from the tools behind this list, and some people will. Specification and&lt;br&gt;
verification and judgement are worth nothing if you refuse to use the machines that make them&lt;br&gt;
valuable, and the engineer who has not sat in front of a frontier model for a hundred hours does not&lt;br&gt;
actually know where its floor is and cannot occupy the seat. The scoreboard is not the point, but&lt;br&gt;
the tools are, and the argument here is for depth with them rather than distance from them.&lt;/p&gt;

&lt;p&gt;The researchers in the AI Explained video may be right that the gap between inside and outside is&lt;br&gt;
large and that things will move faster than the outside expects. If so, the scoreboard will&lt;br&gt;
reshuffle faster, not slower, and the half-life of everything in the bottom two rows gets shorter.&lt;br&gt;
The five things at the top do not get shorter. They get more expensive, because the volume of&lt;br&gt;
machine output that needs specifying, verifying, judging, situating and signing for goes up with&lt;br&gt;
every release, and the number of people who can do those things well does not.&lt;/p&gt;

&lt;p&gt;You are not behind. You have been watching the part of the screen that changes. Look at the part&lt;br&gt;
that does not, and get good at it while everyone else is refreshing the leaderboard.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=iBrAWpjXNxs" rel="noopener noreferrer"&gt;Please stop using stupid models (t3.gg, video)&lt;/a&gt; (18 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=d3Gjq-BffuI" rel="noopener noreferrer"&gt;How are they losing so bad (ThePrimeagen, video)&lt;/a&gt; (15 Sep 2026; the token count and the model timeline are as reported on stream)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.youtube.com/watch?v=J3ljHm57yU0" rel="noopener noreferrer"&gt;What AI Researchers Saw, Before Their Demand to 'Pace' AI (AI Explained, video)&lt;/a&gt; (16 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://blog.detail.dev/posts/towards-self-driving-codebases" rel="noopener noreferrer"&gt;Towards Self-Driving Codebases (detail.dev)&lt;/a&gt; (HN, 119 points, 17 Sep 2026)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://danluu.com/brain-off/" rel="noopener noreferrer"&gt;There's no point at which turning your brain off will work (Dan Luu)&lt;/a&gt; (HN, 196 points, 18 Sep 2026)&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://ai.levelbrook.com/playbook/you-are-not-falling-behind/" rel="noopener noreferrer"&gt;Levelbrook playbook&lt;/a&gt;. Levelbrook is a principal-led Rails and AI-systems consultancy; the playbook is where we write down what we see.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>career</category>
      <category>ai</category>
    </item>
  </channel>
</rss>
