<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Michael Tuszynski</title>
    <description>The latest articles on DEV Community by Michael Tuszynski (@michaeltuszynski).</description>
    <link>https://dev.to/michaeltuszynski</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1447774%2Fa99eea93-7845-4764-9fce-b1755bcfa456.png</url>
      <title>DEV Community: Michael Tuszynski</title>
      <link>https://dev.to/michaeltuszynski</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/michaeltuszynski"/>
    <language>en</language>
    <item>
      <title>The 26% Who Never Rolled Back an Agent Aren't Winning</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Wed, 29 Jul 2026 00:00:48 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/the-26-who-never-rolled-back-an-agent-arent-winning-2dj1</link>
      <guid>https://dev.to/michaeltuszynski/the-26-who-never-rolled-back-an-agent-arent-winning-2dj1</guid>
      <description>&lt;p&gt;Somewhere right now there's a slide in a board deck claiming a clean record: zero agent rollbacks since launch. It reads like a win. It's closer to a smoke detector with the battery pulled out.&lt;/p&gt;

&lt;p&gt;New survey data from Sinch puts the number at &lt;a href="https://sinch.com/news/sinch-releases-ai-production-paradox/" rel="noopener noreferrer"&gt;74% of enterprises that have rolled back a deployed AI agent&lt;/a&gt; after it went live — customer-communications agents specifically, pulled over governance failures. That figure got picked up everywhere as an indictment — three out of four agent deployments blowing up in production. But the interesting number is the one underneath it. Among organizations with the most mature guardrails, the rollback rate goes &lt;strong&gt;up&lt;/strong&gt;, to 81%.&lt;/p&gt;

&lt;p&gt;Read that again. Better safety infrastructure correlates with &lt;em&gt;more&lt;/em&gt; rollbacks, not fewer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rollback rate measures your eyes, not your agent
&lt;/h2&gt;

&lt;p&gt;There's only one clean way to explain that inversion. The teams with better instrumentation aren't shipping worse agents. They're catching things the other teams are shipping past.&lt;/p&gt;

&lt;p&gt;Every agent in production is doing something wrong at some rate. Hallucinated policy details, tool calls against stale records, a tone that drifts on the eighth turn of an angry conversation, a refusal loop that pushes a customer to a human queue that's already 40 deep. The question was never whether the failure exists. It's whether anything in your stack notices it before your customer does.&lt;/p&gt;

&lt;p&gt;So the 26% split into two very different groups. A small number genuinely nailed scope — narrow task, tight tool surface, a human approving anything consequential. The rest have no detector. Their agent is failing at whatever the base rate is, silently, and the absence of a rollback is being reported upward as quality.&lt;/p&gt;

&lt;p&gt;SRE teams learned this lesson two decades ago and it stuck: an incident count is a function of your monitoring, not your reliability. A team that goes from 3 incidents a quarter to 30 after installing real alerting did not get 10x worse. They got 10x more honest. It's the same principle the write-ups on this data keep circling — &lt;a href="https://medium.com/@ripenapps-technologies/why-74-of-enterprises-are-rolling-back-ai-agents-after-launch-738b15a213a3" rel="noopener noreferrer"&gt;the model rarely breaks, the infrastructure around it buckles&lt;/a&gt;. Which means model choice isn't what separates the 26% from everyone else.&lt;/p&gt;

&lt;p&gt;Treating rollback rate as a failure metric creates the exact incentive you don't want. A VP who gets graded on rollbacks will not build better agents. They'll build quieter ones — loosen the eval thresholds, downgrade the alert to a weekly digest, route the complaint channel to a shared inbox nobody owns. The metric improves. The agent doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  The guardrail tax nobody budgeted for
&lt;/h2&gt;

&lt;p&gt;The second number in that survey is the one that should reset your staffing plan. &lt;a href="https://sinch.com/news/sinch-releases-ai-production-paradox/" rel="noopener noreferrer"&gt;84% of AI engineering teams spend at least half their time on safety infrastructure&lt;/a&gt; rather than on the agent itself. And enterprise investment now skews toward trust, security, and compliance (76%) over AI development proper (63%).&lt;/p&gt;

&lt;p&gt;That's not overhead creeping in at the edges. That's the majority of the work.&lt;/p&gt;

&lt;p&gt;It also matches what production looks like. The agent is a prompt, a model call, and a tool list — a week of work for a competent engineer. The other eleven weeks go to the eval set, the PII redaction layer, the escalation path, the audit log that survives a compliance review, the shadow-mode comparison, the per-version rollback switch, and the on-call rotation that owns all of it at 2 a.m.&lt;/p&gt;

&lt;p&gt;Microsoft's 2026 Work Trend Index frames the organizational side of this as a capacity question — &lt;a href="https://www.microsoft.com/en-us/worklab/work-trend-index/agents-human-agency-and-the-opportunity-for-every-organization" rel="noopener noreferrer"&gt;whether organizations are built to capture&lt;/a&gt; the agency that agents free up. Most aren't, and the guardrail tax is a good part of why. The headcount that was supposed to move up the value chain is instead maintaining the machinery that keeps the agent honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who paid for this, and what it actually covers
&lt;/h2&gt;

&lt;p&gt;Now the part the LinkedIn reposts skipped.&lt;/p&gt;

&lt;p&gt;Sinch is a CPaaS vendor. They sell customer-communications infrastructure. A survey they sponsored concluding that &lt;em&gt;infrastructure quality predicts agent success&lt;/em&gt; is a survey concluding that you should buy more of what Sinch sells. That doesn't make the data wrong, but it does mean the framing was chosen before the responses came in.&lt;/p&gt;

&lt;p&gt;The scope is narrower than the headline suggests, too. This is customer-communications agents — support, messaging, conversational flows. Not coding agents, not internal RAG assistants, not the finance-ops bot reconciling invoices. A customer-facing agent has an unusually loud failure mode: the customer complains, and the complaint is logged in a system somebody already reads. Detection is close to free. In a domain where the agent's mistakes land quietly in a document nobody re-reads for six weeks, the 26% number would almost certainly be higher — and mean even less.&lt;/p&gt;

&lt;p&gt;Hold the core finding anyway. The correlation between guardrail maturity and rollback frequency is directionally strong enough to survive the sponsorship discount, because it points the &lt;em&gt;opposite&lt;/em&gt; direction from the sponsor's simplest sales pitch. "Buy our stuff and you'll roll back more often" isn't a slogan anyone reverse-engineers into a survey.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument the rollback, then go trigger one
&lt;/h2&gt;

&lt;p&gt;Stop reporting rollback count. Start reporting rollback &lt;em&gt;provenance&lt;/em&gt;. One row per deployed agent version, and the field that matters is who noticed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;release_id&lt;/span&gt;      &lt;span class="err"&gt;agent-support-v14&lt;/span&gt;
&lt;span class="err"&gt;model&lt;/span&gt;           &lt;span class="err"&gt;claude-sonnet-5&lt;/span&gt;
&lt;span class="err"&gt;prompt_sha&lt;/span&gt;      &lt;span class="err"&gt;a91f3c2&lt;/span&gt;
&lt;span class="err"&gt;tools_sha&lt;/span&gt;       &lt;span class="err"&gt;7de0b41&lt;/span&gt;
&lt;span class="err"&gt;deployed_at&lt;/span&gt;     &lt;span class="py"&gt;2026-07-14T09&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s"&gt;12Z&lt;/span&gt;
&lt;span class="err"&gt;rolled_back_at&lt;/span&gt;  &lt;span class="py"&gt;2026-07-16T23&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s"&gt;41Z&lt;/span&gt;
&lt;span class="err"&gt;detector&lt;/span&gt;        &lt;span class="err"&gt;eval_regression&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;slo_breach&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;abuse_signal&lt;/span&gt;
                &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;agent_self_report&lt;/span&gt; &lt;span class="err"&gt;|&lt;/span&gt; &lt;span class="err"&gt;human_complaint&lt;/span&gt;
&lt;span class="err"&gt;time_to_detect&lt;/span&gt;  &lt;span class="err"&gt;3410&lt;/span&gt; &lt;span class="err"&gt;minutes&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now run the distribution on &lt;code&gt;detector&lt;/code&gt;. If 80% of your rollbacks say &lt;code&gt;human_complaint&lt;/code&gt;, your detection layer is your customers, and your &lt;code&gt;time_to_detect&lt;/code&gt; is however long it takes an annoyed person to find the contact form. Whatever that interval turns out to be in your own data, multiply it by your conversation volume — that's how many exchanges you can't un-send.&lt;/p&gt;

&lt;p&gt;Three things follow from that ledger.&lt;/p&gt;

&lt;p&gt;Set a target on time-to-detect, not on rollback count. An hour is achievable with automated eval replay on a sampled slice of live traffic. Days is what you get when complaint volume is the detector.&lt;/p&gt;

&lt;p&gt;If you've never rolled back, go cause one. Inject a deliberately degraded prompt into a canary slice — drop a rule from the system prompt, or point a tool at a stale index — and time how long your stack takes to flag it. A rollback path that has never been exercised is a hypothesis, not a control. Same logic as restoring from backup: nobody has a backup, they have a restore they've tested or a file they hope is fine.&lt;/p&gt;

&lt;p&gt;And when a vendor pitches you an agent, ask how many times they've rolled back their own. A confident zero means they aren't looking. The right answer sounds like "four times last quarter, median detection 22 minutes, here's the ledger."&lt;/p&gt;

&lt;p&gt;The 81% cohort isn't failing more often. They're the only ones who can prove what their agent did last Tuesday at 3 a.m.&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>aiengineering</category>
      <category>observability</category>
      <category>enterpriseai</category>
    </item>
    <item>
      <title>AI Use Is a 50/50 Coin Flip. Coding Already Tipped 61% Autonomous.</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Thu, 23 Jul 2026 20:08:27 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/ai-use-is-a-5050-coin-flip-coding-already-tipped-61-autonomous-525h</link>
      <guid>https://dev.to/michaeltuszynski/ai-use-is-a-5050-coin-flip-coding-already-tipped-61-autonomous-525h</guid>
      <description>&lt;h2&gt;
  
  
  The Copilot Framing Is Already the Minority Report
&lt;/h2&gt;

&lt;p&gt;Every vendor deck for a coding tool sells you the same picture: a developer in the driver's seat, the AI riding shotgun, a human hand on every merge. Copilot. Human-in-the-loop. Pair programming with a machine. It's the comfortable frame because it keeps the person central and the tool subordinate.&lt;/p&gt;

&lt;p&gt;For software development, that frame already describes the smaller half of what's actually happening.&lt;/p&gt;

&lt;p&gt;Anthropic's &lt;a href="https://www.anthropic.com/economic-index" rel="noopener noreferrer"&gt;Economic Index&lt;/a&gt; publishes real data on how people use Claude, and it tags every conversation one of two ways. Augmentation means the person stays actively in the loop — back and forth, iterating, learning, steering each step. Automation means the person hands off a task and directs Claude to complete it. Across everything, it's nearly a coin flip: 51 percent augmentation, 49 percent automation globally, and a clean 50/50 in the US.&lt;/p&gt;

&lt;p&gt;Narrow to software development tasks and the coin lands differently. 39 percent augmentation, 61 percent automation. The pattern everyone quotes in their pitch is the one coding has already moved past.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Split Actually Measures
&lt;/h2&gt;

&lt;p&gt;Read the label carefully, because it's easy to turn this into something it isn't. Augmentation versus automation is a &lt;strong&gt;conversation style&lt;/strong&gt;, not a jobs claim. It says nothing about displacement, headcount, or whether a role survives. It measures one thing: in a given exchange, is the human collaborating turn by turn, or delegating the whole task?&lt;/p&gt;

&lt;p&gt;That distinction matters because the two styles want different products underneath them. An augmentation session is a dialogue — the tool's job is to be responsive, explain itself, and stay legible while a person drives. An automation session is a delegation — the tool's job is to run to completion and come back with something you can check.&lt;/p&gt;

&lt;p&gt;Most coding tools are still built for the first job. The interface assumes you're watching, approving diffs one at a time, keeping a hand on the wheel. When 61 percent of the actual work is delegation, that assumption is backwards for the majority of what your developers are doing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Number That Keeps This Honest
&lt;/h2&gt;

&lt;p&gt;Here's the counter-weight, and it's in the same dataset. Software development is only about 11.5 percent of global Claude requests, and 8.1 percent in the US. Coding is not the center of gravity for how people use these models. It's a slice.&lt;/p&gt;

&lt;p&gt;That cuts against the temptation to read "61 percent automation in coding" as "AI is taking over programming." It isn't taking over anything. It's one task category, and inside that category the delegation style leads. Both things are true, and the small share is the part that keeps the big number from getting oversold.&lt;/p&gt;

&lt;p&gt;One more caveat worth stating plainly: this is a single snapshot from May 2026. No trend line, no trajectory, no "up 12 points from last quarter." One reading. The 61/39 split is a photograph, not a movie. I'd bet the direction of travel is toward more automation as tools get better at running unattended — but that's my read, not something the data shows. Treat it as a fixed point, not a slope.&lt;/p&gt;

&lt;p&gt;And note the seam in the categories: Anthropic reports "Computer and Mathematical" as the number one task category overall, yet software development specifically sits at 11.5 percent. Those aren't in conflict — the top category is broad and covers a lot more than writing application code. But if you conflate them, you'll overstate how much of AI usage is engineers shipping features. Most of it isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Buyer Should Do With This
&lt;/h2&gt;

&lt;p&gt;If you're choosing or building a coding agent, design for the automation majority. That's the whole recommendation, and it changes what "good" looks like.&lt;/p&gt;

&lt;p&gt;The copilot default optimizes for a person watching every step. The automation pattern optimizes for a person checking the result. Those need different architecture. I've written before that a &lt;a href="https://www.mpt.solutions/the-coding-agent-stack-has-two-layers/" rel="noopener noreferrer"&gt;coding agent stack has two layers&lt;/a&gt; — the model and the scaffolding around it — and that &lt;a href="https://www.mpt.solutions/the-model-doesnt-matter-the-harness-does/" rel="noopener noreferrer"&gt;the scaffolding decides the outcome&lt;/a&gt; more than the model choice does. This data is why. When 61 percent of real coding work is delegation, the layer that runs the agent unattended and catches its mistakes is doing the heavy lifting, not the chat window.&lt;/p&gt;

&lt;p&gt;Concretely, designing for automation means:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Autonomous execution with verification gates, not approval-per-diff.&lt;/strong&gt; The agent should run the task end to end, then hand you a result gated by checks it had to pass — tests green, types clean, lint quiet, a diff that touches only what it claimed. You review the outcome and the evidence, not every keystroke. A human approving line-by-line is the augmentation product, and it doesn't scale to 61 percent of your work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real gates, not vibes.&lt;/strong&gt; The gates are the whole safety story once the human steps back from each step. That means a test suite the agent can't skip, a build it can't merge past when red, and provenance on what changed. If you can't articulate what has to be true before an agent's output reaches main, you haven't built for automation — you've built a faster way to generate unreviewed code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Legibility at the boundary, not throughout.&lt;/strong&gt; Augmentation needs the tool legible at every turn. Automation needs it legible at one point: the handoff. Invest your explanation budget there — what did it do, what did it check, what should you look at first.&lt;/p&gt;

&lt;p&gt;The tools that ship copilot-by-default are optimizing for the 39 percent. That's a real 39 percent, and for exploratory work, learning a new codebase, or high-stakes changes where you genuinely want to drive, augmentation is the right mode. Don't rip it out. But if your default interaction model assumes a human hand on every commit, you've built the minority product and called it the platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Frame Is Behind the Practice
&lt;/h2&gt;

&lt;p&gt;The interesting gap here isn't between humans and machines. It's between how the industry talks about coding agents and how people actually use them. The marketing still runs on the copilot metaphor — reassuring, human-centered, a decade old. The usage data says developers crossed over to delegation a while ago and mostly stopped narrating it.&lt;/p&gt;

&lt;p&gt;61 percent. That's the number to bring to your next tool evaluation. Ask the vendor what happens after the human stops watching — what runs, what gets checked, what stops a bad change cold. If the answer is "well, you approve each diff," they built for the half that's shrinking.&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>softwaredevelopment</category>
      <category>aiautomation</category>
      <category>developertools</category>
    </item>
    <item>
      <title>The Meter Is Always Running</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Tue, 21 Jul 2026 14:02:11 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/the-meter-is-always-running-1poi</link>
      <guid>https://dev.to/michaeltuszynski/the-meter-is-always-running-1poi</guid>
      <description>&lt;h1&gt;
  
  
  The Meter Is Always Running
&lt;/h1&gt;

&lt;p&gt;For forty years, enterprise software rested on a comfortable assumption: that it was a fixed cost. You paid a license, amortized it over a depreciation schedule, and every unit of value you squeezed out afterward felt like it was approaching free. Seats, not usage. Capex, not cost of goods. A CFO could draw a straight line from spend to return and sleep at night.&lt;/p&gt;

&lt;p&gt;AI breaks that line. Most organizations adopted it without noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing that actually changed
&lt;/h2&gt;

&lt;p&gt;The interesting shift in enterprise AI is not cloud versus on-premises. That is a deployment question, and deployment questions are the kind of thing infrastructure teams have solved a hundred times. The deeper shift is quieter: software stopped being a fixed cost and became a variable one.&lt;/p&gt;

&lt;p&gt;A large language model does not bill like a license. It bills like a taxi. The meter runs on every token, upstream and downstream, and &lt;a href="https://www.silicondata.com/blog/llm-cost-per-token" rel="noopener noreferrer"&gt;input and output are priced per million tokens&lt;/a&gt; at rates that differ by model and by direction. The fare depends on how long the conversation ran, how much reasoning it demanded, how many times a user hit retry. Two customers doing the "same" task can cost 10x different amounts because one of them pasted a novel into the prompt. There is no seat count that caps it. There is no version you buy once and freeze. The cost scales with success: the more people use the feature, the more it costs you to have built it.&lt;/p&gt;

&lt;p&gt;That is a new financial object, and the mental model most companies applied to it, SaaS licensing or depreciable capital, is the wrong shape. You cannot amortize a variable cost. You can only manage it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The on-prem mirage
&lt;/h2&gt;

&lt;p&gt;The popular rebuttal goes like this: just run it yourself. Buy the hardware once, own it forever, and your marginal cost per query approaches zero. Convert the scary variable cost back into a familiar fixed one.&lt;/p&gt;

&lt;p&gt;It is a seductive move, and it is mostly wrong.&lt;/p&gt;

&lt;p&gt;The "own it forever" asset is not the GPU. It is the model, and the model is spoiling in real time. The cluster you justified this year runs a frontier model that will be mid-tier in twelve months. The hardware ages too: how long an AI accelerator stays economically useful &lt;a href="https://www.cnbc.com/2025/11/14/ai-gpu-depreciation-coreweave-nvidia-michael-burry.html" rel="noopener noreferrer"&gt;is now an open argument on Wall Street&lt;/a&gt;, with investors questioning whether the six-year depreciation schedules on the books match a reality where a new generation lands every eighteen months. You did not buy a depreciable asset amortized over seven years. You bought a seat on a refresh treadmill, plus a standing bill for power, cooling, and the engineers who keep it fed.&lt;/p&gt;

&lt;p&gt;And unlike the cloud meter, that silicon costs the same whether it runs at 90% load or 5%. Idle capacity is the norm, not the exception: it is common for &lt;a href="https://www.devzero.io/blog/why-your-gpu-cluster-is-idle" rel="noopener noreferrer"&gt;GPU clusters to sit 70 to 80% idle&lt;/a&gt; between bursts of real work. You did not eliminate the variable cost. You traded a usage-based meter for a utilization risk, and for most enterprises, whose demand is spiky rather than steady, that is the worse trade.&lt;/p&gt;

&lt;p&gt;Marginal-cost-to-zero is real, but only at high, sustained load on a model you are content to keep running. That describes a narrow band of workloads. It does not describe "our AI strategy."&lt;/p&gt;

&lt;h2&gt;
  
  
  Inference is cost of goods, not capex
&lt;/h2&gt;

&lt;p&gt;Here is the reframe that survives contact with a spreadsheet: inference is cost of goods sold.&lt;/p&gt;

&lt;p&gt;Not a license you buy. Not a machine you depreciate. It is a per-transaction input cost, like the cloud compute behind a web request or the fee on a card payment. Companies already know how to run a business where serving each customer costs money. They have gross-margin discipline, unit economics, dashboards that flag when a product line goes underwater. They simply never had to apply that discipline to &lt;em&gt;software features&lt;/em&gt; before, because software features used to be free to serve once written.&lt;/p&gt;

&lt;p&gt;File inference under cost of goods instead of "IT capex," and the right behaviors fall out on their own. You start asking the questions a margin-conscious operator always asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that discipline looks like
&lt;/h2&gt;

&lt;p&gt;Know your unit cost per feature. If you cannot say what a single run costs, you are blind on the one number that decides whether the feature is a business or a liability. Instrument tokens the way you already instrument latency.&lt;/p&gt;

&lt;p&gt;Then route by need. Most requests do not warrant a frontier model; send the cheap, frequent, predictable work to smaller or local models and reserve the expensive reasoning for the calls that actually earn it. A good router is worth more than a bigger model.&lt;/p&gt;

&lt;p&gt;Cache aggressively. Repeated system prompts, retrieved context, common answers: every hit is a fare you skip. Anthropic reports that &lt;a href="https://www.anthropic.com/news/prompt-caching" rel="noopener noreferrer"&gt;prompt caching can cut input costs by up to 90%&lt;/a&gt; and latency along with it. Caching is not a nice-to-have optimization here. It is margin.&lt;/p&gt;

&lt;p&gt;Put a ceiling on it. Variable cost without guardrails is how one looping agent turns into a five-figure surprise. Budget caps, rate limits, and per-tenant quotas are the circuit breaker. Build them before the invoice, not after.&lt;/p&gt;

&lt;p&gt;And decide placement per workload, not per company. Some things belong on a local model you own; some belong on a metered API. "Cloud-first" and "on-prem-first" are both the wrong altitude. The unit is the workload, and the answer is a portfolio.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reckoning is real. It is not the one on the thumbnail.
&lt;/h2&gt;

&lt;p&gt;A reckoning is coming for a lot of AI initiatives, but it will not arrive as the dramatic collapse of anyone's cloud strategy. It will arrive as a quiet quarterly review where someone finally divides the AI line item by the number of customers it served and does not like the answer. The projects that survive that meeting will not be the ones that picked the right deployment location. They will be the ones that treated inference as what it is, a variable operating cost with no natural ceiling, and built the discipline to manage it from the first day.&lt;/p&gt;

&lt;p&gt;The meter was always running. The only question is whether you were watching it.&lt;/p&gt;

</description>
      <category>aieconomics</category>
      <category>enterpriseai</category>
      <category>cloudstrategy</category>
      <category>aiinfrastructure</category>
    </item>
    <item>
      <title>Your AI Coding Pilot Cost $80K and Shipped 9% Faster. Here's the Line Item Finance Missed.</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Mon, 20 Jul 2026 14:07:50 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/your-ai-coding-pilot-cost-80k-and-shipped-9-faster-heres-the-line-item-finance-missed-4hlb</link>
      <guid>https://dev.to/michaeltuszynski/your-ai-coding-pilot-cost-80k-and-shipped-9-faster-heres-the-line-item-finance-missed-4hlb</guid>
      <description>&lt;p&gt;The slide always looks the same. Pilot spend: $80K. Velocity: up 9%. Recommendation: expand to the full org. I have sat through enough of these readouts, on both sides of the table, to recite the speaker notes from memory. The room approves it, because $80K buying a 9% faster engineering org of any size is obviously a deal.&lt;/p&gt;

&lt;p&gt;The math on the slide is fine. The problem is a line item that never made it onto the slide, and it is the most expensive input in the building.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the $80K Went
&lt;/h2&gt;

&lt;p&gt;Run the composite pilot. Fifty seats of an AI coding assistant at roughly $20 a head is about $12K a year. Agent and API tokens for the teams that went past autocomplete, call it $3K a month once real workloads land — &lt;a href="https://www.mpt.solutions/what-my-ai-workflow-actually-costs-per-month/" rel="noopener noreferrer"&gt;I've published my own bill, so I know what this curve looks like&lt;/a&gt;. Add the integration sprint, the eval work, and the enablement time nobody bills to the pilot, and $80K is a fair all-in number for a mid-size org's first serious year.&lt;/p&gt;

&lt;p&gt;Every dollar of that is metered, invoiced, and visible. Seats show up on a purchase order. Tokens show up on a usage dashboard. Finance can read the pilot's cost to the penny, which is exactly why the readout feels rigorous.&lt;/p&gt;

&lt;p&gt;Now the return side. A 9% velocity gain across a 12-engineer group is roughly one extra engineer's worth of throughput, and &lt;a href="https://www.mpt.solutions/an-engineer-costs-250k-their-tokens-cost-20k-that-math-is-a-trap/" rel="noopener noreferrer"&gt;a loaded engineer runs $250K, which is the number that dwarfs every token bill in sight&lt;/a&gt;. So the slide implies $80K bought about $270K of capacity. A 3.4x return, approved before the coffee gets cold.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Line Item That Never Shows Up
&lt;/h2&gt;

&lt;p&gt;Here is what the slide does not carry: the review and verification attention the pilot consumed, priced at senior-engineer rates.&lt;/p&gt;

&lt;p&gt;AI-assisted pull requests change shape. Diffs get bigger. Volume goes up. And the failure mode shifts from obviously-broken to plausible-but-wrong, which is the expensive kind, because &lt;a href="https://www.mpt.solutions/babysitter-auditor-prayer-or-tests/" rel="noopener noreferrer"&gt;plausible-but-wrong has to be caught by a human who understands the system&lt;/a&gt;. Suppose each engineer ships five AI-assisted PRs a week and each one takes a reviewer just twenty extra minutes. Across 12 engineers that is 20 hours a week of review time, half a full-time engineer, roughly $125K a year. And it does not land evenly: it lands on the two or three senior people qualified to catch the plausible-but-wrong category, &lt;a href="https://www.mpt.solutions/attention-is-the-new-bottleneck-engineer-it-like-one/" rel="noopener noreferrer"&gt;the same scarce attention that was already the bottleneck&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Take the $270K gain, subtract $80K in tool spend and $125K in unpriced senior attention, and the triumphant 3.4x collapses to about 1.2x, before counting rework. That might still be worth doing. But nobody in the room approved &lt;em&gt;that&lt;/em&gt; deal, because nobody saw it.&lt;/p&gt;

&lt;p&gt;The reason nobody saw it is structural, and it is worth saying plainly: &lt;strong&gt;finance manages what is metered.&lt;/strong&gt; Seats and tokens arrive as invoices. Attention arrives as nothing at all. It is smeared invisibly across a salary line that looks identical whether your seniors spent the quarter building or spent it correcting a machine's confident guesses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evidence Says the Slide Flatters Itself
&lt;/h2&gt;

&lt;p&gt;The 9% is usually self-reported or lightly measured, and self-report runs hot. In &lt;a href="https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/" rel="noopener noreferrer"&gt;METR's randomized study of experienced open-source developers&lt;/a&gt;, participants estimated AI made them 20% faster while the measured result was 19% &lt;em&gt;slower&lt;/em&gt;. The perception gap is the mechanism that hides the missing line item: the tool feels fast because typing is fast, while the slow part — verifying, correcting, re-prompting — doesn't register as "using the tool."&lt;/p&gt;

&lt;p&gt;The vendor numbers run hotter still. &lt;a href="https://github.blog/news-insights/research/research-quantifying-github-copilots-impact-on-developer-productivity-and-happiness/" rel="noopener noreferrer"&gt;GitHub's own study reported 55% faster task completion&lt;/a&gt;, a first-party figure measured on a greenfield toy task, and it should be discounted accordingly. Meanwhile &lt;a href="https://dora.dev/ai/gen-ai-report/" rel="noopener noreferrer"&gt;DORA's research keeps finding the bill on the other side of the ledger: as AI adoption rises, software delivery stability dips&lt;/a&gt;. Faster generation, wobblier delivery. That wobble is rework, and rework is more senior attention, unpriced.&lt;/p&gt;

&lt;p&gt;To be fair to the other side of the argument: METR's result is one population on one kind of task, not a universal law, and pilots built on scoped tasks with an explicit verification budget genuinely do clear the honest bar. The point is narrower and harder to dodge. A pilot that measured only what was invoiced has not measured its return. It has measured its receipts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five-Line Version
&lt;/h2&gt;

&lt;p&gt;The fix costs nothing but honesty. Here is the pilot P&amp;amp;L with all the lines on it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI coding pilot P&amp;amp;L

+ Velocity delta
    measured (cycle time,
    merged scope), never
    surveyed
- Seats + tokens
    the invoice (the only
    line the readout had)
- Review-hours delta
    PR review time before
    vs. during, x reviewer
    cost
- Rework delta
    change-failure/rollback
    rate, before vs. during
- Comprehension debt
    merged code nobody on
    the team can explain, %
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first two lines exist in every pilot. The last three exist in almost none, and they are all measurable with tooling you already run: your git host timestamps every review, your incident tracker already counts rollbacks, and the comprehension question takes one uncomfortable team meeting.&lt;/p&gt;

&lt;p&gt;If the pilot still clears the bar with five lines on the page, expand it with confidence — mine does, and that's precisely why &lt;a href="https://www.mpt.solutions/where-your-20k-in-tokens-actually-goes/" rel="noopener noreferrer"&gt;I keep publishing the full bill instead of the flattering half&lt;/a&gt;. A tool that survives honest accounting is a tool worth scaling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ask for the Line
&lt;/h2&gt;

&lt;p&gt;The 9% on the slide is neither a lie nor a result. It is an unfinished calculation, presented with the confidence of a finished one, to a room that is being asked to multiply it by the whole org.&lt;/p&gt;

&lt;p&gt;So ask the one question that finishes it: show me the review-hours delta. If the answer is a number, you are looking at a rare, well-run pilot, and you should probably fund it. If the answer is a pause, then the 9% was never a measurement. It was a survey wearing one, and the most expensive people in the building are quietly paying the difference.&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>engineeringleadership</category>
      <category>developerproductivity</category>
      <category>softwaredelivery</category>
    </item>
    <item>
      <title>The Bubble Popper and the Payoff Are the Same Thing</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Thu, 16 Jul 2026 23:59:00 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/the-bubble-popper-and-the-payoff-are-the-same-thing-fk5</link>
      <guid>https://dev.to/michaeltuszynski/the-bubble-popper-and-the-payoff-are-the-same-thing-fk5</guid>
      <description>&lt;p&gt;Cory Doctorow thinks AI is a bubble, and he argues it better than almost anyone who cheers him on. In &lt;a href="https://pluralistic.net/2025/12/05/pop-that-bubble/" rel="noopener noreferrer"&gt;his December speech at the University of Washington&lt;/a&gt;, the case runs like this: the tech giants stopped growing years ago, and a company priced as a growth stock faces catastrophe the moment the market stops believing. So they pump whatever keeps the multiple alive. Pivot to video, crypto, NFTs, Metaverse, now AI, each pitched with the same conviction, each abandoned when the next vehicle arrives. "Superintelligence" is the pitch this cycle because science fiction makes a better P/E story than enterprise software ever did.&lt;/p&gt;

&lt;p&gt;His sharpest tool is the reverse centaur. A centaur is a human assisted by a machine: you drive the car, the car amplifies you. A reverse centaur is a human bolted onto a machine as its peripheral, hired to absorb the failures the machine can't, at machine pace, under machine supervision. Doctorow's thesis is that the money behind AI is betting on reverse centaurs: workers demoted into error-cleanup for systems sold to their bosses as replacements.&lt;/p&gt;

&lt;p&gt;Take him seriously. The losses are real, the incentives he names are real, and I have sat through enough board-mandated AI pilots to know the reverse centaur is not a hypothetical. And then notice that his argument, stated carefully, is about two different balance sheets that the doom discourse keeps smearing into one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Balance Sheets, One Shouting Match
&lt;/h2&gt;

&lt;p&gt;Seller economics asks whether the labs make money. Company-level losses are enormous, but look at where they live: training runs, data-center capex, sales, and the free tier. The serving unit underneath is a different story. &lt;a href="https://newsletter.semianalysis.com/p/anthropic-3q26-profit-over-1b-the" rel="noopener noreferrer"&gt;SemiAnalysis puts Anthropic's gross margins in the mid-60s&lt;/a&gt; — an independent analyst number, though &lt;a href="https://www.theinformation.com/articles/anthropic-lowers-profit-margin-projection-revenue-skyrockets" rel="noopener noreferrer"&gt;The Information reported the same company projecting 40% when inference costs spiked&lt;/a&gt;, so treat the exact figure as contested. The direction is not. &lt;a href="https://www.seangoedecke.com/ai-inference-is-obviously-profitable/" rel="noopener noreferrer"&gt;Serving paid tokens is gross-margin-positive&lt;/a&gt;; what bleeds is everything wrapped around it.&lt;/p&gt;

&lt;p&gt;Even the famous counterexample proves the shape. When Sam Altman admitted &lt;a href="https://x.com/sama/status/1876104315296968813" rel="noopener noreferrer"&gt;OpenAI loses money on the $200 Pro plan&lt;/a&gt;, the reason was that heavy users outran the flat price. Sell unbounded agent workloads at a fixed monthly fee and the top of the usage tail eats you. That is a pricing-model problem, and pricing models get fixed. Physics problems don't. This one is being fixed right now, mostly at the expense of people like me: usage caps, tier splits, metered agents.&lt;/p&gt;

&lt;p&gt;Buyer economics asks a different question: does a scoped deployment pay for itself? That is where operators live, and none of the training-run losses show up on this side of the table. &lt;a href="https://www.mpt.solutions/what-my-ai-workflow-actually-costs-per-month/" rel="noopener noreferrer"&gt;I've written down what my own stack costs per month&lt;/a&gt;, and &lt;a href="https://www.mpt.solutions/an-engineer-costs-250k-their-tokens-cost-20k-that-math-is-a-trap/" rel="noopener noreferrer"&gt;why "$20K in tokens against a $250K engineer" is the wrong math even when it flatters the tools&lt;/a&gt;. A deployment either produces measured value over measured cost or it doesn't. The lab's income statement has nothing to do with it.&lt;/p&gt;

&lt;p&gt;Doctorow is right about the first balance sheet. The mistake is letting the first one answer for the second.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Question That Does All the Work
&lt;/h2&gt;

&lt;p&gt;Here is my actual thesis, earned from reps rather than from a valuation model: the economics work themselves out, but only for operators disciplined enough to ask "what problem am I solving" before pointing the tool at anything, and rigorous enough to measure the answer afterward.&lt;/p&gt;

&lt;p&gt;An efficiency you don't understand is not an efficiency. It is an unmeasured cost wearing a demo. The team that deploys a coding agent with no baseline, no success metric, and no owner has not automated anything; it has hired a very fast intern nobody supervises and booked the salary as savings. That is blind deployment, and blind deployment is exactly the reverse-centaur failure: the humans end up serving the machine's output because nobody defined what the machine was for.&lt;/p&gt;

&lt;p&gt;Which means my thesis and Doctorow's frame are the same observation read from opposite ends. Scoping rigor is what makes you a centaur instead of a reverse one. The human who decides what the machine is for stays the head. The human who cleans up after an unscoped machine becomes the peripheral.&lt;/p&gt;

&lt;p&gt;The scoping questions I use, written down (writing them down matters, and the last section says why):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. What is the problem, stated without naming a tool?
2. What does it cost today, in hours or dollars I can point at?
3. What does "working" look like, measured how, by whom?
4. What is the failure mode, and who catches it?
5. What is the kill criterion, decided before the pilot starts?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Payoff and the Pin Are the Same Force
&lt;/h2&gt;

&lt;p&gt;Now the part I have not seen anyone say plainly: the rigor that makes AI pay off is the same force that pops the bubble.&lt;/p&gt;

&lt;p&gt;Walk through what happens if scoping discipline actually spreads. Every FOMO seat license gets audited against question 2, and half of them fail. Every board-mandated pilot meets question 3 and dies for lack of a metric. The all-you-can-eat experimentation budgets get metered. A real slice of today's AI revenue is hype-spend, and disciplined buyers stop paying for hype. Revenue evaporates precisely as deployments get better.&lt;/p&gt;

&lt;p&gt;That is not bearish on AI. It is bearish on the bubble, and those are different positions. It is Doctorow's own &lt;a href="https://pluralistic.net/2025/12/05/pop-that-bubble/" rel="noopener noreferrer"&gt;"some bubbles leave behind something productive"&lt;/a&gt; scenario made concrete. He reaches for Worldcom, a fraud whose CEO died in prison while the dark fiber it buried still carries Doctorow's 2-gigabit home connection. Swap the fiber for scoped deployments quietly compounding inside businesses while the FOMO spend burns off. The economics working out and the bubble deflating are the same event, narrated by an optimist and a pessimist.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Craftsman Clause
&lt;/h2&gt;

&lt;p&gt;There is an investor-side implication buried in this that is more uncomfortable than any crash prediction.&lt;/p&gt;

&lt;p&gt;If every deployment that pays needs a human who understands both the domain and the tool's failure modes, then AI is not the frictionless labor replacement the trillion-dollar valuations are priced on. It is a power tool that still needs a craftsman. A nail gun does not fire the carpenter; it makes the carpenter's judgment the binding constraint on every house. &lt;a href="https://www.mpt.solutions/attention-is-the-new-bottleneck-engineer-it-like-one/" rel="noopener noreferrer"&gt;Attention was always the bottleneck&lt;/a&gt;; the tools just moved it.&lt;/p&gt;

&lt;p&gt;Great news for buyer ROI. Quietly fatal for the "human exits the loop" story being sold upstream, because the wholesale-replacement bet only pays if the craftsman requirement goes away, and every honest deployment I have run or reviewed says it doesn't. The augmentation bet keeps cashing small checks. The replacement bet keeps pre-spending checks nobody has figured out how to write.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Half-Lives, One Fork
&lt;/h2&gt;

&lt;p&gt;So if the scarce input is the human who can scope, the personal question is what that scarcity is worth and how long it lasts.&lt;/p&gt;

&lt;p&gt;The scoping skill is teachable in principle and bottlenecked in practice, and the bottleneck is structural, not a fad. The skill is tacit. It lives on the seam between a domain and a tool, where few people sit. And organizations promote people for shipping, not for the scoping that made the shipping cheap; nobody's OKRs reward the pilot that didn't happen.&lt;/p&gt;

&lt;p&gt;But the edge has two halves with opposite half-lives, and confusing them is the trap. Tool mechanics — prompt-craft, knowing this month's model's failure modes, the tricks — depreciates fast, because every model release is the labs folding exactly that knowledge into the product. &lt;a href="https://www.mpt.solutions/how-to-run-an-agent-loop-without-burning-your-token-budget/" rel="noopener noreferrer"&gt;The agent-loop mechanics I wrote up&lt;/a&gt; are already aging out from under that post. Problem formulation — translating a business's real problem out from under its stated one — depreciates slowly, because it is organizational translation, and organizations stay human. The trap is that the fast-depreciating half is the legible, demoable half. It feels like the edge because you can show it off. The durable half looks like "just asking questions."&lt;/p&gt;

&lt;p&gt;Which leaves the fork, and I am going to leave it as a fork rather than resolve it for you. Stay the indispensable bottleneck and you extract rent: real money, now, capped at your personal throughput, gone the day the scarcity ends. Or codify the tacit half into artifacts a colleague can run without you — a written scoping method like the checklist above, reference patterns from the deployments that worked — and you dissolve your own bottleneck while capturing the value of having systematized it. Rent pays this quarter. The moat is owning the codified version when everyone else finally needs it.&lt;/p&gt;

&lt;p&gt;The bubble being real is what creates the opening; nobody pays a premium for discipline in a calm market. The economics work out for whoever is on the right side of the rent-to-moat conversion. And the thing to be dismantling, deliberately, starting now, is the bottleneck being you personally — before the labs dismantle the cheap half of your edge for free and leave you holding only the part you never bothered to write down.&lt;/p&gt;

</description>
      <category>aieconomics</category>
      <category>enterpriseai</category>
      <category>aistrategy</category>
      <category>aiagents</category>
    </item>
    <item>
      <title>Torvalds' Best Review Trick Just Stopped Working</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Thu, 16 Jul 2026 18:11:41 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/torvalds-best-review-trick-just-stopped-working-37gg</link>
      <guid>https://dev.to/michaeltuszynski/torvalds-best-review-trick-just-stopped-working-37gg</guid>
      <description>&lt;p&gt;Linus Torvalds has spent years reviewing the most consequential codebase on Earth without reading much of the code. He said so himself in &lt;a href="https://www.tag1consulting.com/blog/interview-linus-torvalds-linux-and-git" rel="noopener noreferrer"&gt;a 2020 interview with Tag1 Consulting&lt;/a&gt;: "While I still look at patches, I actually tend to look more at the explanations, and the history of how the patch came to me."&lt;/p&gt;

&lt;p&gt;That habit was never laziness. It was the sharpest quality filter in software, and it rested on an economic fact: a coherent explanation of why a change belongs in the kernel was expensive to produce. You had to understand the subsystem, the failure mode, and the design decisions that shaped the current code. Faking the explanation cost more than having the understanding. So the explanation worked as proof of understanding, and the man at the top of the review pyramid could read intent instead of implementation.&lt;/p&gt;

&lt;p&gt;At &lt;a href="https://www.zdnet.com/article/open-source-summit-linus-torvalds/" rel="noopener noreferrer"&gt;Open Source Summit India 2026 in Mumbai&lt;/a&gt; this month, Torvalds described the thing that broke that filter. He never framed it as broken. The quotes do it for him.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Explanation Was Proof of Work
&lt;/h2&gt;

&lt;p&gt;Tests can be gamed. Diffs can be pattern-matched from similar fixes without understanding either one. But for decades, a paragraph that correctly situated a change in the design history of a kernel subsystem could not be written by someone who lacked the mental model. Reviewers up and down the kernel hierarchy leaned on that correlation. Torvalds built his entire post-programming job on it.&lt;/p&gt;

&lt;p&gt;The correlation was always a proxy, though. Nobody cares about the explanation itself. They care about the understanding it demonstrates, and &lt;a href="https://www.mpt.solutions/goodharts-law-just-got-a-slash-command/" rel="noopener noreferrer"&gt;Goodhart's law&lt;/a&gt; says any proxy works right up until it becomes a target. A proxy survives on the cost of faking it. This one survived for thirty years because the fake was more work than the real thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  LLMs Made the Proxy Free
&lt;/h2&gt;

&lt;p&gt;Large language models generate code that is sometimes wrong. They generate prose about code that is nearly always fluent. The explanation, the one artifact that used to be hardest to fake, is now the cheapest part of the submission.&lt;/p&gt;

&lt;p&gt;Torvalds described what that looks like from the receiving end: bug reports that read as entirely valid and turn out to be fabricated. &lt;a href="https://linux.slashdot.org/story/26/07/12/2053201/linus-torvalds-on-ai-junk-patches-humans-and-godzilla" rel="noopener noreferrer"&gt;"It can actually be a huge drain on resources when it takes humans a lot of effort to figure out that, hey, this machine-generated report was not true,"&lt;/a&gt; he said. Note the asymmetry: the report costs nothing to generate and real investigative work to disprove. That is a filter running backwards, taxing the reviewers it used to protect.&lt;/p&gt;

&lt;p&gt;The patches carry the same signature. He called many of them "mindless band-aid kind of patches... they may fix the immediate problem, but the kind of bug remains, and it just is waiting in the hallway to hit you in another place." A band-aid patch in 2019 usually arrived with a thin, awkward description, and the description gave it away. The same patch in 2026 arrives wrapped in a confident, well-structured explanation of a root cause the author never found. The tell is gone. Prose quality no longer predicts anything about the understanding behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fix He Announced Is a Provenance Requirement
&lt;/h2&gt;

&lt;p&gt;Buried in the Mumbai remarks is the actual news, and it got almost no coverage: "If you find a bug with an LLM, it's not enough to just ask the LLM to make a bug report and then throw it over the fence to us. We want to see a suggested patch; we want to see the human who ran the LLM act as a kind of back-and-forth."&lt;/p&gt;

&lt;p&gt;Read that as a submission requirement, not a complaint. He is no longer asking the explanation to prove anything. He is asking for evidence of the process: show me the dialogue, show me where you pushed back, show me what the model got wrong before you fixed it. The unit of review is shifting from the artifact to its provenance. For decades the kernel asked whether the description demonstrated understanding. The new question is whether the human can show their working relationship with the tool that produced it.&lt;/p&gt;

&lt;p&gt;That is a bigger change than it sounds. The transcript of the back-and-forth, not the polished description sitting on top of it, is becoming the thing a maintainer actually wants to see.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody's Tooling Captures This Yet
&lt;/h2&gt;

&lt;p&gt;PR templates ask what changed and why. CI gates check tests, coverage, lint. None of them ask the question Torvalds is now asking: how do you know? The description field is where the fluent fake lives. The evidence of a genuine back-and-forth lives in session logs most teams throw away.&lt;/p&gt;

&lt;p&gt;Closing that gap does not need a platform. A template section does most of it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Provenance&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Generated with: &lt;span class="nt"&gt;&amp;lt;tool&lt;/span&gt;&lt;span class="err"&gt;/&lt;/span&gt;&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="na"&gt;or&lt;/span&gt; &lt;span class="err"&gt;"&lt;/span&gt;&lt;span class="na"&gt;by&lt;/span&gt; &lt;span class="na"&gt;hand&lt;/span&gt;&lt;span class="err"&gt;"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; First-pass mistake I caught: &lt;span class="nt"&gt;&amp;lt;what&lt;/span&gt; &lt;span class="na"&gt;you&lt;/span&gt; &lt;span class="na"&gt;pushed&lt;/span&gt; &lt;span class="na"&gt;back&lt;/span&gt; &lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; What I checked myself: &lt;span class="nt"&gt;&amp;lt;command&lt;/span&gt; &lt;span class="na"&gt;you&lt;/span&gt; &lt;span class="na"&gt;ran&lt;/span&gt;&lt;span class="err"&gt;,&lt;/span&gt; &lt;span class="na"&gt;output&lt;/span&gt; &lt;span class="na"&gt;you&lt;/span&gt; &lt;span class="na"&gt;read&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The middle field is the tell. A developer who worked the problem with a model fills it in ten seconds, because every real session has one. A developer who piped output over the fence has nothing to put there, and an empty field is a louder signal than a beautiful description.&lt;/p&gt;

&lt;p&gt;I started keeping this kind of evidence before I had a name for it. When I published &lt;a href="https://www.mpt.solutions/build-a-self-improving-agent-harness-in-an-afternoon/" rel="noopener noreferrer"&gt;a self-improving agent demo&lt;/a&gt; earlier this month, the repo shipped with the actual run logs, a 13-of-15 test pass climbing to 15-of-15 across iterations, because without the logs the claim is indistinguishable from every other AI demo on the internet. The logs are the proof the loop ran. The README is just the description.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Breaks
&lt;/h2&gt;

&lt;p&gt;The obvious objection: transcripts can be faked too. Ask the model that wrote your patch to also write a plausible back-and-forth and it will produce one, complete with staged pushback. If teams start grading transcripts, Goodhart eats the transcript next. Grading them with &lt;a href="https://www.mpt.solutions/your-llm-judge-needs-a-test-suite/" rel="noopener noreferrer"&gt;an LLM judge&lt;/a&gt; inherits the same problem one level up.&lt;/p&gt;

&lt;p&gt;The kernel's answer is already visible, and it is not paperwork. Torvalds has been blunt that &lt;a href="https://www.phoronix.com/forums/forum/software/programming-compilers/1604826-linus-torvalds-the-ai-slop-issue-is-not-going-to-be-solved-with-documentation" rel="noopener noreferrer"&gt;the AI slop problem will not be solved with documentation&lt;/a&gt;, and kernel maintainers already &lt;a href="https://techstrong.ai/articles/open-source-makes-bugs-shallow-linus-torvalds-says-ai-makes-them-public/" rel="noopener noreferrer"&gt;deprioritize drive-by reports whose submitters won't answer questions&lt;/a&gt;. The gate is not the artifact. The gate is whether you can continue the conversation. A fabricated transcript buys exactly one round; the first follow-up question a maintainer asks lands on the human, in real time, with no model in the loop to sound fluent for them.&lt;/p&gt;

&lt;p&gt;So Torvalds' test did not die in Mumbai. It moved. Reading the explanation stopped working, so he replaced it with watching you respond.&lt;/p&gt;

&lt;p&gt;The next time a pull request lands in your queue with a suspiciously fluent description, run the kernel's version of the test. Ask one follow-up question the description does not answer, and watch the clock. The answer that comes back in five minutes was always the author's. The one that takes a day was written by whatever wrote the description.&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>opensource</category>
      <category>codereview</category>
      <category>softwarequality</category>
    </item>
    <item>
      <title>Attention Is the New Bottleneck. Engineer It Like One.</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Tue, 14 Jul 2026 16:43:39 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/attention-is-the-new-bottleneck-engineer-it-like-one-6fo</link>
      <guid>https://dev.to/michaeltuszynski/attention-is-the-new-bottleneck-engineer-it-like-one-6fo</guid>
      <description>&lt;h2&gt;
  
  
  The Diagnosis Is Right
&lt;/h2&gt;

&lt;p&gt;Atomic Object published a piece last month arguing that &lt;a href="https://spin.atomicobject.com/ai-agents-attention-bottleneck" rel="noopener noreferrer"&gt;with AI agents, attention is the new bottleneck&lt;/a&gt;. Agents made execution cheap. You can fire off four parallel coding agents before your coffee cools. What you can't do is review four streams of output at once, hold the context of each in your head, and catch the one that quietly wrote a migration that drops a column.&lt;/p&gt;

&lt;p&gt;The diagnosis is correct. Simon Willison described &lt;a href="https://www.youtube.com/watch?v=so9l_MwS2yg" rel="noopener noreferrer"&gt;running four parallel agents and being wiped out by 11am&lt;/a&gt; — not because the tools failed, but because he became the rate limiter. Michael Novati made the same point from a different angle: AI removed the production bottleneck and &lt;a href="https://michaelnovati.substack.com/p/the-real-bottleneck-in-the-ai-era" rel="noopener noreferrer"&gt;revealed the real one underneath — the human system that surrounds production&lt;/a&gt;. Everyone circling this problem is seeing the same thing. Execution went to zero and human attention became the scarce resource.&lt;/p&gt;

&lt;p&gt;So the diagnosis holds. The prescription is where it falls apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  The "Type Faster" Answer
&lt;/h2&gt;

&lt;p&gt;Atomic Object's advice is all personal discipline. Automate trivial decisions. Plan for cognitive load. Take deliberate breaks. Budget your attention like a finite resource. Good advice for a person. Useless as a system.&lt;/p&gt;

&lt;p&gt;We've seen this exact move before. When execution was the bottleneck — when developers spent their day typing, compiling, catching their own syntax errors — nobody's answer was "type faster." Nobody wrote a productivity blog telling engineers to be more disciplined about their keystrokes. We built compilers so you didn't hand-check types. We built linters so nobody argued about a missing semicolon. We built CI so the machine ran the test suite at 2am and told you which commit broke it.&lt;/p&gt;

&lt;p&gt;The bottleneck moved and we answered with infrastructure, not willpower. Telling a developer in 1998 to concentrate harder on avoiding null-pointer bugs would have been absurd. Telling a developer in 2026 to budget their attention better is the same absurdity wearing new clothes.&lt;/p&gt;

&lt;p&gt;Discipline doesn't scale. Infrastructure does. If your plan for the attention bottleneck is "I'll be more focused," you've already lost, because the failure mode of human attention isn't insufficient effort — it's that there's a finite amount of it and agents produce work faster than any amount of focus can absorb.&lt;/p&gt;

&lt;h2&gt;
  
  
  Engineer the Bottleneck Instead
&lt;/h2&gt;

&lt;p&gt;Here's the reframe. Attention is a resource your system spends on your behalf. Most systems spend it wastefully — they make a human look at everything. The job is to build the parts that spend it carefully, the same way a good compiler spends your debugging time carefully by pointing at line 47 instead of making you read the whole file.&lt;/p&gt;

&lt;p&gt;Four pieces do most of the work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verification gates so humans review exceptions, not everything.&lt;/strong&gt; The default agent workflow asks you to eyeball every output. That's the waste. Instead, have the agent verify its own work and only surface what it can't confirm. My content pipeline ships blog posts across five channels. After it publishes, a &lt;code&gt;verifyPublishedPost()&lt;/code&gt; step re-fetches the committed markdown and checks for the two failures that actually happened to me — a missing feature image, and a leading H1 that double-rendered the title. If it finds one, it fires a 🚨 line into my notification channel. If everything's clean, I hear nothing. I went from reading every published post to reading only the broken ones. That's not discipline. That's a gate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structured escalation so one line beats a log stream.&lt;/strong&gt; Uptime monitoring is the canonical trap. You can watch a dashboard, or you can read logs, or you can have every alert route to a single channel with a single flagged line: what broke, where, since when. I don't read my nightly job logs. They write to disk. If a job fails, one message arrives. The difference between "check the logs" and "here is the one thing that needs you" is the difference between spending attention and saving it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Trust tiers per task class.&lt;/strong&gt; Not all work deserves the same scrutiny, and pretending it does is how you burn attention on things that don't need it. Low-stakes work ships on green checks — my content pipeline's Tier 1 curated shares go out on a passing lint gate, no human in the loop. High-stakes work queues for review. The on-demand blog publisher treats my explicit request as the approval and ships without a gate, because I asked for it. The cron pipeline that drafts on its own routes through an approval message first. Same infrastructure, different trust tier, matched to the blast radius of being wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decision logs so the same call never reaches a human twice.&lt;/strong&gt; This is the one people skip, and it's the one that compounds. Every time you make a judgment call for an agent, write it down where the agent reads it. My project instructions carry a running list of hard-won lessons — never hand-author markdown into the blog repo, opt into Instagram explicitly because the publishing API returns false-success statuses, strip any leading body H1 unconditionally. Each of those is a decision I made exactly once. The agent now makes it every time without me. A decision log is a cache for judgment. Without it, you re-answer the same question forever, which is the most expensive way there is to spend attention.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Objection Worth Taking Seriously
&lt;/h2&gt;

&lt;p&gt;The honest counterargument: infrastructure has a build cost, and discipline is free today. Writing a self-verification step, wiring escalation, defining trust tiers — that's real work, and for a solo developer running one agent occasionally, personal focus genuinely is cheaper. Atomic Object isn't wrong that a human should also automate their own trivial decisions and take breaks.&lt;/p&gt;

&lt;p&gt;But that's a bet on the problem staying small, and it won't. The whole premise is that agents multiply output. The developer running one agent this month runs six next quarter. Discipline that works at one stream collapses at six — that's the entire bottleneck being described. You pay the infrastructure cost once and it holds as you scale. You pay the discipline cost every single day, and it fails exactly when the load gets heavy enough to matter. Build the gate before you need it, because the moment you need it you won't have attention left to build it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Separates the Winners
&lt;/h2&gt;

&lt;p&gt;The framing that treats attention as a personal-productivity problem is going to age like every other "work smarter" answer to a structural constraint. It puts the burden on the operator to be superhuman, when the whole point of the last forty years of tooling was to stop requiring humans to be superhuman.&lt;/p&gt;

&lt;p&gt;The people who win with agents won't be the ones with the most discipline about their attention. Discipline is a fixed, small, human quantity, and the workload is about to be neither fixed nor small. The winners will be the ones whose systems spend their attention the least — whose agents verify themselves, escalate the one thing that matters, ship low-stakes work without asking, and never bring the same decision back a second time.&lt;/p&gt;

&lt;p&gt;Attention is the new bottleneck. So engineer it like one. We didn't beat the execution bottleneck by typing faster, and we won't beat this one by concentrating harder.&lt;/p&gt;

</description>
      <category>aiagents</category>
      <category>platformengineering</category>
      <category>aiengineering</category>
      <category>developerproductivity</category>
    </item>
    <item>
      <title>Try It: A Working Assessment-First Course</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Mon, 13 Jul 2026 11:01:23 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/try-it-a-working-assessment-first-course-203b</link>
      <guid>https://dev.to/michaeltuszynski/try-it-a-working-assessment-first-course-203b</guid>
      <description>&lt;p&gt;Eight posts ago the claim was that the AI-education industry is building the wrong product — chatbots students ignore, while the thing that actually moves exam scores is an LLM grading written answers against a rubric, wrapped in spaced cumulative review. Now there's a running system to argue with instead of a claim to nod at. This is the capstone of the &lt;a href="https://www.mpt.solutions/the-ai-tutor-everyone-builds-is-the-one-students-ignore/" rel="noopener noreferrer"&gt;assessment-first series&lt;/a&gt;: what got built, how to run it, and where the bet breaks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it in five minutes
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/michaeltuszynski/doerkit" rel="noopener noreferrer"&gt;doerkit&lt;/a&gt; is a full course — six statistics lessons from OpenStax OER, quizzes, cumulative review, a dosage dashboard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/michaeltuszynski/doerkit &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;doerkit
npm &lt;span class="nb"&gt;install
export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-ant-...
npm run dev          &lt;span class="c"&gt;# http://localhost:8734&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pick a name, read a lesson, take its quiz. Write a real answer to a constructed-response question and watch it get graded against the rubric with feedback in about a second; write "the median because reasons" and watch it get partial credit with a specific note on what's missing. Fail the 90% review bar, get nudged to come back tomorrow instead of cramming. Open &lt;code&gt;/dashboard&lt;/code&gt; and see your own dosage. The &lt;a href="https://github.com/michaeltuszynski/rubric-bench" rel="noopener noreferrer"&gt;grader is regression-tested&lt;/a&gt; by the sibling repo, including against the prompt-injection answers a real student would try.&lt;/p&gt;

&lt;p&gt;That's the whole thesis, executable. The LLM never chats, never does the student's work, never assigns a grade directly — it judges rubric criteria as booleans and code computes the rest.&lt;/p&gt;

&lt;h2&gt;
  
  
  What eight posts actually shipped
&lt;/h2&gt;

&lt;p&gt;Two repositories, both MIT, both green in CI, both tagged v1.0:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://github.com/michaeltuszynski/rubric-bench" rel="noopener noreferrer"&gt;rubric-bench&lt;/a&gt;&lt;/strong&gt; — regression testing for any LLM judge. Golden sets, run scoring, drift diffs, an adversarial suite, tone metrics. The general-purpose one; useful well beyond education.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;doerkit&lt;/strong&gt; — the platform: grading engine, lessons, mixed-format quizzes, interleaved spaced review, telemetry.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The findings that surprised me, collected: a frontier model shrugged off first-generation prompt injections that a cheaper model fell for, so grader security lives in the &lt;em&gt;model-prompt pair&lt;/em&gt; and moves when you swap either. Grader severity and grader warmth are separable knobs: you can be kind without inflating grades, which means a cold grader is a defect, not rigor. And the boring cumulative-review feature carried the biggest effect size in the source study, beating both the AI grader and the chatbot everyone demos.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a real deployment would still need
&lt;/h2&gt;

&lt;p&gt;The honest gap between "runs on my laptop" and "runs a gateway course," so nobody mistakes this for the second thing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LMS integration&lt;/strong&gt;: LTI 1.3, roster sync, gradebook. Unglamorous, mandatory, and deliberately absent here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Auth and multi-tenancy&lt;/strong&gt;: the demo trusts a self-typed name. A real one needs SSO, real accounts, and per-institution isolation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A FERPA data agreement&lt;/strong&gt;: the moment student-keyed telemetry leaves a laptop it's regulated education data, with all the procurement that implies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-rater validation&lt;/strong&gt;: &lt;a href="https://www.mpt.solutions/your-llm-judge-needs-a-test-suite/" rel="noopener noreferrer"&gt;post 3&lt;/a&gt; regression-tests grading consistency, not agreement with instructors. A pilot needs an inter-rater study against real graded work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An RCT&lt;/strong&gt;: everything here rests on one observational pilot at one selective school. The design is a hypothesis with strong priors, not proof.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are hard research problems. They're the difference between a portfolio and a product, and pretending otherwise is how edtech demos oversell.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the whole bet breaks
&lt;/h2&gt;

&lt;p&gt;The strongest counterargument to this series is selection. The students who complete more lessons and pass all three reviews are the ones who were going to ace the final anyway; the &lt;a href="https://intextbooks.science.uu.nl/workshop2026/files/itb26_s1s2.pdf" rel="noopener noreferrer"&gt;Dartmouth data&lt;/a&gt; brackets the effect between 0.71 SD (over-adjusted) and 1.30 SD (selection-inflated) precisely because it can't fully separate the platform from the motivation. I believe the effect is real and meaningful — the cross-format contrast, where constructed-response dosage tracked scores and multiple-choice didn't within the same students, is hard to explain by motivation alone, but "real and meaningful" is a defensible position, not a settled one. Anyone who tells you AI tutoring has proven 1.3-SD gains is selling.&lt;/p&gt;

&lt;p&gt;And there's a tension the series surfaced without resolving: disabling constructed response in the pilot &lt;em&gt;raised&lt;/em&gt; completion rates, because writing answers is more work than clicking. The highest-efficacy format may carry an engagement tax. The whole bet is that the tax is worth paying and that better grader tone shrinks it, but that's the open question a real study exists to answer, not one this code settles.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual takeaway
&lt;/h2&gt;

&lt;p&gt;If you build one thing from these eight posts, don't make it an education product. Make it the &lt;a href="https://github.com/michaeltuszynski/rubric-bench" rel="noopener noreferrer"&gt;eval suite&lt;/a&gt;. Every team putting an LLM judge into production — grading, triage, moderation, ranking — has the exact problem post 3 solved and mostly doesn't know it yet: their judge's behavior is an untested production dependency that changes when the model updates. Golden sets, drift diffs, adversarial cases, tone guards. That pattern outlives statistics, outlives edtech, and outlives whatever model you're calling this quarter.&lt;/p&gt;

&lt;p&gt;The chatbot got two years of the industry's attention. The quiz engine moved the exam scores. Both repos are public, both are yours to fork, and the code is the argument.&lt;/p&gt;

</description>
      <category>aieducation</category>
      <category>llmevaluation</category>
      <category>opensource</category>
      <category>developertools</category>
    </item>
    <item>
      <title>Where Your $20K in Tokens Actually Goes</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Mon, 13 Jul 2026 02:03:54 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/where-your-20k-in-tokens-actually-goes-408f</link>
      <guid>https://dev.to/michaeltuszynski/where-your-20k-in-tokens-actually-goes-408f</guid>
      <description>&lt;p&gt;The last piece argued that comparing a $250K engineer to a $20K token bill is a trap, because the two numbers measure different things and the cheap one hides its real cost. Fine. But "the token bill is not the whole story" leaves a follow-up hanging: what is actually in that $20K? Where does the money go once the invoice clears?&lt;/p&gt;

&lt;p&gt;It goes to waste, mostly. Not fraud, not overpriced models. Ordinary, invisible waste that nobody instruments because the bill arrives as one number and one number tells you nothing. A $20K monthly spend on production AI is not a cost. It is a pipeline with leaks at four specific joints, and every one of them is measurable.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bill is not one line, it's four leaks
&lt;/h2&gt;

&lt;p&gt;Start with the model. On Claude, Opus 4.8 runs &lt;a href="https://www.anthropic.com/pricing" rel="noopener noreferrer"&gt;$15 per million input tokens and $75 per million output&lt;/a&gt;, and Sonnet lands at roughly a fifth of that. Those are the posted numbers, and they are the part of the bill you can't argue with. What you can argue with is how many tokens you send, how many times you send them, and how many of those sends did no useful work.&lt;/p&gt;

&lt;p&gt;Break a real production bill apart and the same four categories show up every time: retries and stalls, context and prompt bloat, tool-schema overhead from MCP servers, and redundant eval or judge passes. Each one has a fix. None of the fixes require a better model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Leak one: retries and stalls
&lt;/h2&gt;

&lt;p&gt;An agent loop that hits a rate limit, a malformed tool call, or a timeout doesn't fail cleanly. It retries. And the retry re-sends the entire context that got you to the failure point, so a stall at turn nine costs you turns one through nine again, in full, at output prices if the model already started generating.&lt;/p&gt;

&lt;p&gt;This is the quiet killer in agentic workloads. A single user request that should cost one round trip can spawn six because the agent got confused, called a tool wrong, read the error, and tried again. In production agent loops, error-recovery and re-planning routinely consume more tokens than the productive reasoning does. The model isn't thinking harder. It's redoing work.&lt;/p&gt;

&lt;p&gt;The fix is boring and it works: cap retries explicitly, log every retry with the reason, and treat a high retry rate as a bug in your tool definitions, not a cost of doing business. If your agent retries a tool 30% of the time, the tool's description is unclear or its schema is wrong. Fix the schema and the retries disappear. Instrument this first, because it's usually the biggest single line and the one teams never look at.&lt;/p&gt;

&lt;h2&gt;
  
  
  Leak two: context and prompt bloat
&lt;/h2&gt;

&lt;p&gt;Every request carries a system prompt, a set of instructions, skill definitions, and whatever conversation history you've accumulated. Most teams have no idea how big that payload is. They wrote the system prompt eight months ago, bolted three more skills onto it, and never measured the total.&lt;/p&gt;

&lt;p&gt;Measure it. A tool like &lt;a href="https://github.com/michaeltuszynski/token-baseline" rel="noopener noreferrer"&gt;token-baseline&lt;/a&gt; run across your prompt, skill, and command corpus gives you a real number per component, so you can see that the 4,000-token "helpful preamble" nobody has read since launch is riding along on every single call. At Opus input rates, 4,000 wasted tokens on 100,000 daily calls is real money, and it buys you nothing.&lt;/p&gt;

&lt;p&gt;The structural fix is prompt caching. Anthropic's &lt;a href="https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching" rel="noopener noreferrer"&gt;prompt caching&lt;/a&gt; charges a 25% premium to write a cache entry and then serves cache reads at 10% of the base input price. If your system prompt is stable across calls, and it should be, caching turns a repeated 4,000-token tax into a one-time write plus pennies per read. Teams that cache their stable prefix routinely cut input spend by half or more. Teams that don't are paying full freight to re-send the same unchanging text thousands of times a day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Leak three: MCP tool-schema overhead
&lt;/h2&gt;

&lt;p&gt;This one is new enough that most teams haven't caught it yet. Connect an agent to a handful of MCP servers and every request now carries the full JSON schema for every tool those servers expose, before the model reads a word of the actual task. GitHub's server alone ships around 35 tools. Stack a few servers and you can burn 50,000-plus tokens on schemas the model will not use on this particular request.&lt;/p&gt;

&lt;p&gt;You can audit this directly. A tool like &lt;a href="https://github.com/michaeltuszynski/mcp-token-audit" rel="noopener noreferrer"&gt;mcp-token-audit&lt;/a&gt; measures the token cost of each connected server's schema payload, and the results are usually ugly. Half your context window can be tool definitions for capabilities the current task doesn't touch.&lt;/p&gt;

&lt;p&gt;The fix is on-demand tool loading: expose tool names to the model, and load the full schema only when the model decides to call the tool. &lt;a href="https://dev.to/loading-tool-schemas-on-demand-is-how-agents-scale/"&gt;I wrote about the mechanics of this separately&lt;/a&gt;, but the short version is that you should never pay to describe a tool the request won't use. Audit your schema payload, then defer everything you can.&lt;/p&gt;

&lt;h2&gt;
  
  
  Leak four: redundant eval and judge passes
&lt;/h2&gt;

&lt;p&gt;The last leak comes from good intentions. You added an LLM judge to grade outputs, then an eval pass to check the judge, then a second judge for confidence. Now every production response triggers three extra model calls, and two of them are asking nearly the same question.&lt;/p&gt;

&lt;p&gt;Judge calls are output-heavy and they compound. If your judge re-reads the full input plus the candidate answer plus a rubric on every call, you're paying to re-process the same context three times to answer one quality question. Sample instead of grading everything. Grade 5% of production traffic continuously and the full set only when you ship a prompt change. Collapse redundant judges into one call with a structured multi-field output. The goal is confidence in your quality, not a receipt for every token you can spend proving it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Instrument the pipeline, then fix the joints
&lt;/h2&gt;

&lt;p&gt;Here's the position: stop treating the token bill as a price and start treating it as telemetry. The number on the invoice is the sum of four measurable subsystems, and every one of them leaks in a way you can see the moment you point a tool at it.&lt;/p&gt;

&lt;p&gt;Run token-baseline against your prompt and skill corpus this week. Run mcp-token-audit against your connected servers. Turn on prompt caching for your stable prefix. Put a counter on retries and set an alarm when the rate climbs. None of this is exotic, and none of it needs budget approval. It needs someone to accept that "$20K in tokens" is a diagnosis waiting to happen, not a fact to shrug at.&lt;/p&gt;

&lt;p&gt;The trap in the first piece was believing the token number was small. The trap in this one is believing it's a single number at all. It isn't. It's a pipeline. Instrument it, and half of it turns out to be leak.&lt;/p&gt;

</description>
      <category>aiengineering</category>
      <category>costoptimization</category>
      <category>aiagents</category>
      <category>llmops</category>
    </item>
    <item>
      <title>Instrument Like a Learning Scientist</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Sun, 12 Jul 2026 11:01:42 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/instrument-like-a-learning-scientist-4o47</link>
      <guid>https://dev.to/michaeltuszynski/instrument-like-a-learning-scientist-4o47</guid>
      <description>&lt;p&gt;The most valuable thing the Dartmouth team built wasn't the grader. It was the fact that they could answer "did completing this lesson's quiz correlate with doing better on the exam?" — per module, per format. That question is why they discovered multiple-choice quizzing produced no measurable learning while constructed-response did. Without per-lesson dosage logged against exam outcomes, that finding is invisible, and the platform ships the useless format forever because everyone &lt;em&gt;felt&lt;/em&gt; engaged.&lt;/p&gt;

&lt;p&gt;This is post 7 of the &lt;a href="https://www.mpt.solutions/the-ai-tutor-everyone-builds-is-the-one-students-ignore/" rel="noopener noreferrer"&gt;assessment-first series&lt;/a&gt;. It's about the least glamorous and most compounding part of &lt;a href="https://github.com/michaeltuszynski/doerkit" rel="noopener noreferrer"&gt;doerkit&lt;/a&gt;: the telemetry that lets the platform measure itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dosage is the variable that matters
&lt;/h2&gt;

&lt;p&gt;Most edtech analytics report engagement — logins, time-on-page, questions attempted. Those are vanity metrics; they measure whether people showed up, not whether showing up did anything. The variable the &lt;a href="https://intextbooks.science.uu.nl/workshop2026/files/itb26_s1s2.pdf" rel="noopener noreferrer"&gt;Dartmouth study&lt;/a&gt; built its whole argument on is &lt;em&gt;dosage&lt;/em&gt;: how many lessons a student actually completed, regressed against exam performance. The distinction is the entire finding: engagement was comparable-or-higher under multiple-choice, but dosage only tracked exam scores under constructed response. If you log engagement you learn nothing; if you log dosage you learn which features work.&lt;/p&gt;

&lt;p&gt;So doerkit logs two things from day one: an append-only &lt;code&gt;events&lt;/code&gt; table (lesson views, quiz starts, submissions with score and pass/fail) and an &lt;code&gt;attempts&lt;/code&gt; table (every quiz and review attempt with its score). Both carry a student key and a timestamp. That's the minimal schema, and it's enough to reconstruct dosage-versus-outcome for any cohort you later attach exam scores to.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;attempts&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;student&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;-- 'lesson' | 'review'&lt;/span&gt;
  &lt;span class="n"&gt;lesson_id&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="nb"&gt;REAL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;passed&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The dashboard is the instrument
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/michaeltuszynski/doerkit/blob/main/src/app/dashboard.ts" rel="noopener noreferrer"&gt;dosage dashboard&lt;/a&gt; rolls this up per student: lessons passed (the dosage number), quiz attempts, average score, reviews passed versus attempted, total events. It's a plain SQL rollup rendered as an HTML table, with no charting library and no analytics vendor. The point isn't the visualization; it's that the raw material for an efficacy analysis exists the moment the first student touches the platform, instead of being a data-collection project you scramble to start after someone asks whether the thing works.&lt;/p&gt;

&lt;p&gt;That framing matters for what this platform is &lt;em&gt;for&lt;/em&gt;. Efficacy evidence is the currency of institutional edtech sales and the thing every rigorous claim in this space is missing. A platform instrumented for dosage-outcome analysis generates its own evidence base as a byproduct of being used — every cohort makes the next efficacy claim stronger. The data asset compounds; the code doesn't. That's the actual moat in this category, and it costs two database tables to start accruing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The minimal event schema for any learning product
&lt;/h2&gt;

&lt;p&gt;If you're building anything with practice and outcomes, log these from commit one, before you think you need them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The dose&lt;/strong&gt; — the countable unit of work (lessons completed, problems solved), per user, timestamped. Not time-on-page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The verdict&lt;/strong&gt; — pass/fail and score on each attempt, so you can separate "attempted a lot" from "attempted well."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The retry structure&lt;/strong&gt; — every attempt, not just the last, with its timestamp. The &lt;a href="https://doi.org/10.1177/1529100612453266" rel="noopener noreferrer"&gt;~1.5-day spacing finding&lt;/a&gt; only existed because retries were individually logged.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stable subject key&lt;/strong&gt; — so you can join to outcomes later without re-identifying anyone.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Retrofitting this after launch means the first cohort is unmeasurable, and the first cohort is exactly the one a skeptical instructor asks about. Instrument before you need it, because the need arrives as a question you can't answer retroactively.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks
&lt;/h2&gt;

&lt;p&gt;Dosage-outcome correlation is not causation, and this is the load-bearing caveat for the whole series: motivated students both complete more lessons and score higher, so raw dosage regressions are selection-inflated. The Dartmouth authors handled it by controlling for prior midterm performance, which brackets the true effect between an over-adjusted 0.71 SD and a selection-inflated 1.30 SD, and doerkit's telemetry can produce the same bracketing only if you feed it exam scores, which it doesn't collect on its own. There's a privacy surface too: a student-keyed event log is FERPA-relevant data the moment this leaves a laptop, so the demo uses a self-chosen name and no real roster, and a genuine deployment needs a data agreement this scope deliberately avoids. Telemetry that measures learning is also telemetry that surveils learners; build it, and own that both are true.&lt;/p&gt;

&lt;p&gt;Next post is the capstone: run the whole thing yourself, what an actual institutional deployment would still need, and an honest accounting of where the assessment-first bet holds and where it doesn't.&lt;/p&gt;

</description>
      <category>edtech</category>
      <category>learningscience</category>
      <category>productanalytics</category>
      <category>assessment</category>
    </item>
    <item>
      <title>The Biggest Effect Size Was the Boring Feature</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Sat, 11 Jul 2026 11:01:35 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/the-biggest-effect-size-was-the-boring-feature-2bpo</link>
      <guid>https://dev.to/michaeltuszynski/the-biggest-effect-size-was-the-boring-feature-2bpo</guid>
      <description>&lt;p&gt;The feature with the largest effect size in the &lt;a href="https://intextbooks.science.uu.nl/workshop2026/files/itb26_s1s2.pdf" rel="noopener noreferrer"&gt;Dartmouth pilot&lt;/a&gt; is the one no startup would put on a landing page. Not the AI grader. Not the chatbot. Cumulative module reviews: a big quiz covering every lesson, questions interleaved across topics, a 90% bar, unlimited retries. Students who passed all three scored 7.1 points higher on the final (d = 0.66), the strongest signal in the study. It looks like a quiz from 2005.&lt;/p&gt;

&lt;p&gt;This is post 6 of the &lt;a href="https://www.mpt.solutions/the-ai-tutor-everyone-builds-is-the-one-students-ignore/" rel="noopener noreferrer"&gt;assessment-first series&lt;/a&gt;, and it assembles the pieces from posts 2 through 5 into a running web app — &lt;a href="https://github.com/michaeltuszynski/doerkit" rel="noopener noreferrer"&gt;doerkit&lt;/a&gt;. Lessons, quizzes, and that unglamorous review engine, which is the part worth dwelling on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why interleaving beats blocked practice
&lt;/h2&gt;

&lt;p&gt;A lesson quiz tests one topic while it's fresh — you just read about the median, then answer about the median. That's "blocked" practice, and it produces a specific illusion: fluency inside the current topic that evaporates when topics are mixed. The student feels like they know it because the context is doing half the retrieval.&lt;/p&gt;

&lt;p&gt;A cumulative review mixes topics. Question 3 is about outliers, question 4 is about z-scores, question 5 is back to sampling. Now each question forces the harder move (&lt;em&gt;which&lt;/em&gt; concept does this even call for) before you can answer it. &lt;a href="https://doi.org/10.1177/1529100612453266" rel="noopener noreferrer"&gt;Interleaved retrieval practice&lt;/a&gt; is one of the most replicated findings in learning science, and it consistently loses on the in-session feeling of mastery while winning on the exam weeks later. That gap between how it feels and how it works is exactly why it doesn't sell, and exactly why it's the feature that moved scores.&lt;/p&gt;

&lt;p&gt;doerkit's review builds a 10-question quiz by taking at most two questions per lesson across the whole module, so no student sees a topic-blocked run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;pickReviewQuiz&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;banks&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;QuestionBank&lt;/span&gt;&lt;span class="p"&gt;[]):&lt;/span&gt; &lt;span class="nx"&gt;QuizQuestion&lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;perLesson&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;banks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;pick&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;questions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;pick&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;perLesson&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flat&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nx"&gt;REVIEW_QUIZ_SIZE&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// interleaved by construction&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Spacing, and the nudge
&lt;/h2&gt;

&lt;p&gt;The other half of the effect was timing. The Dartmouth retry logs showed students returning to reviews a median of ~1.5 days apart, spaced retrieval rather than cramming, and that spacing wasn't designed, it emerged from a high pass bar plus unlimited retries. A 90% threshold you'll rarely clear on the first try, with no penalty for coming back, quietly manufactures the exact study schedule the research recommends.&lt;/p&gt;

&lt;p&gt;doerkit makes the nudge explicit. Retry within 20 hours of your last review attempt and it says so: &lt;em&gt;"Retrieval sticks better with a gap; coming back in about 14h beats retrying now (you can retry anyway)."&lt;/em&gt; It never blocks the retry. The whole design principle from post 1 holds: the platform succeeds only when students choose to come back, so it persuades rather than gates. Manufacturing a good habit out of a pass threshold and a soft nudge is cheaper and more durable than any streak mechanic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack, deliberately small
&lt;/h2&gt;

&lt;p&gt;The whole app is &lt;a href="https://github.com/michaeltuszynski/doerkit/tree/main/src/app" rel="noopener noreferrer"&gt;Hono plus SQLite&lt;/a&gt;, server-rendered HTML, no client framework. Lessons are markdown authored from OpenStax OER. Quizzes draw from per-lesson banks of ten. The &lt;a href="https://www.mpt.solutions/grading-written-answers-with-an-llm-properly/" rel="noopener noreferrer"&gt;grading engine from post 2&lt;/a&gt; grades written answers concurrently; multiple choice is a pure function. Content is never gated — you can read any lesson and take any quiz in any order, because the Dartmouth platform wasn't gated and hit 90% voluntary adoption.&lt;/p&gt;

&lt;p&gt;What's deliberately absent is the tell. No LMS integration, no SSO, no multi-tenancy, no student roster, no RAG chatbot. Those are the features that make edtech an 18-month enterprise sale, and none of them touch the thing that moves exam scores. The scope is frozen to exactly what the evidence supports: read, write an answer, get judged, come back spaced. A pilot instructor can run this on a laptop the afternoon they find it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Applying the pattern past statistics
&lt;/h2&gt;

&lt;p&gt;The review engine is subject-agnostic. Any course with lessons and question banks gets interleaving for free — the scheduler just needs topics to mix and a pass bar high enough to invite a second visit. Language vocabulary, medical board prep, onboarding curricula: the same two knobs (interleave across units, set a threshold that manufactures spacing) transfer directly. The AI grader is what makes the written half economical; the review engine is what makes any of it stick. You need both, and only one of them is exciting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks
&lt;/h2&gt;

&lt;p&gt;The spacing nudge is honest but toothless by design; a determined crammer ignores it, and without course credit attached (the Dartmouth quizzes were ungraded) the population that most needs spacing is the least likely to self-impose it. The interleaving is within-module only; true long-horizon spacing across a whole term needs scheduling state this version doesn't keep. And the biggest honest caveat carries over from the study itself: the all-reviews-passed group was also the most self-selected, so some of that d = 0.66 is motivated students being motivated. The within-module comparison (passing the review predicted a 6.1-point midterm gain holding cohort fixed) is the cleaner evidence, and it's smaller. Real, but smaller.&lt;/p&gt;

&lt;p&gt;Next post — the last one — is the capstone: run it yourself, what a real institutional deployment would still need, and where this whole assessment-first bet does and doesn't hold.&lt;/p&gt;

</description>
      <category>edtech</category>
      <category>learningscience</category>
      <category>aiengineering</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Grader Was Right and the Students Quit Anyway</title>
      <dc:creator>Michael Tuszynski</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:34:31 +0000</pubDate>
      <link>https://dev.to/michaeltuszynski/the-grader-was-right-and-the-students-quit-anyway-38ba</link>
      <guid>https://dev.to/michaeltuszynski/the-grader-was-right-and-the-students-quit-anyway-38ba</guid>
      <description>&lt;p&gt;The most dangerous failure in the &lt;a href="https://intextbooks.science.uu.nl/workshop2026/files/itb26_s1s2.pdf" rel="noopener noreferrer"&gt;Dartmouth Phosphor pilot&lt;/a&gt; wasn't a wrong grade. Students found the constructed-response grader "rigid and discouraging," complained loudly enough that the team removed those questions from an entire module. And the module without them turned out to produce no measurable learning. The feature that worked got pulled because of how it &lt;em&gt;felt&lt;/em&gt;. Accuracy survived contact with students; tone didn't.&lt;/p&gt;

&lt;p&gt;This is post 5 of the &lt;a href="https://www.mpt.solutions/the-ai-tutor-everyone-builds-is-the-one-students-ignore/" rel="noopener noreferrer"&gt;assessment-first series&lt;/a&gt;, and it treats grader tone the way post 3 treated grader accuracy: as a measurable property under regression test, not a vibe you hope survives the next prompt edit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Severity and tone are different knobs
&lt;/h2&gt;

&lt;p&gt;The intuition says a "strict" grader gives lower grades and a "warm" grader inflates them — that you buy kindness with rigor. I tested that. Same 72-case golden set, same model (claude-sonnet-5), two system prompts: the default from &lt;a href="https://www.mpt.solutions/grading-written-answers-with-an-llm-properly/" rel="noopener noreferrer"&gt;post 2&lt;/a&gt; ("precise and warm: a good TA, not a gatekeeper... name what the answer got right first") and a strict-examiner variant ("do not give benefit of the doubt... do not praise, do not soften").&lt;/p&gt;

&lt;p&gt;Verdicts barely moved. Expected-partial answers graded down to incorrect: &lt;strong&gt;3 of 24 under both prompts&lt;/strong&gt;, the same three cases. Overall accuracy within one case (69 vs 68 of 72). The persona change did not make the grading harsher.&lt;/p&gt;

&lt;p&gt;The &lt;em&gt;experience&lt;/em&gt; changed completely. Across the 60 genuine cases:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;metric&lt;/th&gt;
&lt;th&gt;default&lt;/th&gt;
&lt;th&gt;strict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;feedback opening with something the student got right&lt;/td&gt;
&lt;td&gt;54/60&lt;/td&gt;
&lt;td&gt;38/60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;feedback containing scolding phrases ("fails to", "unfortunately"...)&lt;/td&gt;
&lt;td&gt;0/60&lt;/td&gt;
&lt;td&gt;5/60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;average feedback length&lt;/td&gt;
&lt;td&gt;37.8 words&lt;/td&gt;
&lt;td&gt;32.8 words&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same case, both prompts. Default: &lt;em&gt;"You correctly state that the mean exceeds the median, but you don't describe the shape of the distribution... so the reasoning credit isn't earned."&lt;/em&gt; Strict: &lt;em&gt;"No description of the distribution shape is given... only vaguely references 'mean follows the skew,' which does not meet the required attribution."&lt;/em&gt; Identical verdict, identical partial credit. One reads like a TA who wants you to pass; the other reads like a rejection letter. A student on their third retry at 11pm reads the second one and closes the tab.&lt;/p&gt;

&lt;p&gt;That separability is the finding: warmth is nearly free. You don't pay for it with grade inflation; the rubric-and-booleans architecture keeps verdicts anchored while the prose register moves independently. Which means shipping a cold grader isn't rigor. It's just a product defect you haven't measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making tone a regression test
&lt;/h2&gt;

&lt;p&gt;The measurement is deliberately crude, because crude and automated beats sophisticated and manual. &lt;a href="https://github.com/michaeltuszynski/rubric-bench" rel="noopener noreferrer"&gt;rubric-bench&lt;/a&gt; golden cases now take a &lt;code&gt;feedbackForbidden&lt;/code&gt; list alongside &lt;code&gt;feedbackMustMention&lt;/code&gt; (terms that must never appear in feedback, like "unfortunately" or "you failed"), and every run keeps the full feedback text per case, so the &lt;a href="https://github.com/michaeltuszynski/rubric-bench/blob/main/examples/analyze-tone.ts" rel="noopener noreferrer"&gt;tone analysis script&lt;/a&gt; can report positive-acknowledgment rates and scold counts across whole runs. A prompt edit that keeps accuracy but drops the positive-opener rate from 90% to 60% now fails visibly, in CI, before a student sees it.&lt;/p&gt;

&lt;p&gt;Regex against feedback text is a blunt instrument and I'm comfortable with that. The alternative, an LLM judging the tone of an LLM's feedback, is a real technique, but it puts a second nondeterministic judge in your test suite, and you'd need a bench for the bench. Start with substring guards on the phrases you never want students to read; graduate to a tone judge only when the blunt version stops catching real regressions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The product frame
&lt;/h2&gt;

&lt;p&gt;Formative assessment lives or dies on retry behavior. The Dartmouth data's strongest feature was students returning to cumulative reviews a median of ~1.5 days apart — voluntary spaced retrieval, the thing &lt;a href="https://doi.org/10.1177/1529100612453266" rel="noopener noreferrer"&gt;decades of learning science&lt;/a&gt; says to maximize. Every piece of that loop runs on the student choosing to come back, and the feedback message is the last thing they read before choosing. This is why "the grader was accurate" and "the grading feature failed" can both be true: accuracy is a property of verdicts, retention is a property of the loop, and tone is the hinge between them.&lt;/p&gt;

&lt;p&gt;There's a business asymmetry here too. A too-lenient grader fails quietly and gets caught by the adversarial suite from &lt;a href="https://www.mpt.solutions/students-are-adversaries-red-teaming-an-llm-grader/" rel="noopener noreferrer"&gt;post 4&lt;/a&gt;. A too-cold grader fails loudly: screenshots, complaints, an instructor pulling the feature mid-term. The Dartmouth team's response (rip out constructed response, discover the replacement taught nothing) is what unmeasured tone costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks
&lt;/h2&gt;

&lt;p&gt;The metrics are proxies, and proxies saturate: a grader could open every message with a hollow "Good effort!" and score perfectly on positive-acknowledgment while being useless. The forbidden-terms list is English-specific and enumerable, and a genuinely different feedback register (a new model's house style) could pass every guard while feeling off in ways only students will tell you. Verdict-tone separability held on this model pair and this rubric architecture; a judge that assigns scores directly (no boolean criteria) would likely see verdicts drift with persona. And 60 cases of feedback is a tone sample, not a study — the real instrument is a mid-term student survey sitting next to the bench numbers.&lt;/p&gt;

&lt;p&gt;Next post: the platform itself. Lessons, quizzes, and the boring cumulative-review feature that carried the biggest effect size in the study — assembled into a runnable web app.&lt;/p&gt;

</description>
      <category>aiassessment</category>
      <category>learningscience</category>
      <category>edtech</category>
      <category>llmapplications</category>
    </item>
  </channel>
</rss>
