<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ted</title>
    <description>The latest articles on DEV Community by Ted (@henry_dan_81513dd35a2f540).</description>
    <link>https://dev.to/henry_dan_81513dd35a2f540</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2256153%2F4cd4eaa5-dfdf-46a4-9487-af0cbe57b775.jpeg</url>
      <title>DEV Community: Ted</title>
      <link>https://dev.to/henry_dan_81513dd35a2f540</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/henry_dan_81513dd35a2f540"/>
    <language>en</language>
    <item>
      <title>I Audited My AI's Briefing File. It Had Turned Into a Diary.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:13:25 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/i-audited-my-ais-briefing-file-it-had-turned-into-a-diary-1hhc</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/i-audited-my-ais-briefing-file-it-had-turned-into-a-diary-1hhc</guid>
      <description>&lt;p&gt;Coding agents like Claude Code read a plain-text briefing file before they do anything else. Mine is called &lt;code&gt;CLAUDE.md&lt;/code&gt;, and it sits in the home directory of the small server I run my automation on. It tells the agent who I am, what's installed, which services run where, what the cron jobs do, and the handful of rules I don't want broken. It loads into every session, whatever I'm working on.&lt;/p&gt;

&lt;p&gt;I'd been adding to it for six months. After two model releases in ten days, I ran Claude Code's new prompt audit on it (&lt;code&gt;/checkup prompt-audit&lt;/code&gt;). The audit reads your instruction files and flags text written for older models: ALL-CAPS warnings, "think step by step", rigid scripts. It writes a report and a patch and changes nothing until you say so.&lt;/p&gt;

&lt;p&gt;I expected it to find shouting. It found almost none. What it found instead was more useful, and it changed how I think about the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I expected
&lt;/h2&gt;

&lt;p&gt;The audit's main checklist is about prompting habits that used to help and now hurt. Older models needed forceful instructions to follow anything reliably, so people wrote &lt;code&gt;IMPORTANT:&lt;/code&gt; and &lt;code&gt;NEVER&lt;/code&gt; everywhere. Current models follow instructions closely and literally, so the same shouting makes them over-apply a rule and behave rigidly in situations it was never meant for.&lt;/p&gt;

&lt;p&gt;My file had exactly one of those: a rule in capitals about using a site-specific command-line tool only for the site it was built for. The fix was to say it at normal volume and keep the reason: "the other sites have no command center." Two other spots used capitals for emphasis, but they stated facts, not rules, so the audit left them alone.&lt;/p&gt;

&lt;p&gt;That was the whole outdated-prompting section of the report. One finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually found
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A quarter of the file was history.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The cron section was meant to list my scheduled jobs. Over the months its header had grown into an 880-word paragraph about what I'd &lt;em&gt;removed&lt;/em&gt;: which jobs were cut on which date, why, where the backups went, the traffic numbers that made me drop three sites, which repository was left untouched. Every entry had been accurate when I wrote it. None of it was something the agent needed to do anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Several facts had quietly gone wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The file listed the Gemini CLI at a specific path. It wasn't installed there, or anywhere.&lt;/li&gt;
&lt;li&gt;It gave a version for one of my agent gateways that was two months out of date.&lt;/li&gt;
&lt;li&gt;It listed four API keys in my environment file. There were more than a dozen.&lt;/li&gt;
&lt;li&gt;It listed the tools wired into my voice assistant and was missing one.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these would cause a crash. They'd cause a confident wrong answer: the agent telling me "use the Gemini CLI at this path" and spending a turn finding out it isn't there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One rule was a security habit dressed as documentation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Under my blog's repository it said: "Git push requires the token in the remote URL", followed by the exact command to paste a GitHub token into the repo's config. That was true once, because nothing else was set up to handle the login. But it's an instruction, so every time the agent pushed, it would write a live credential into a plain-text file. I'd already cleaned that exact pattern out of two other repos a few weeks earlier, and the instruction file was still teaching it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The flip
&lt;/h2&gt;

&lt;p&gt;I'd been treating &lt;code&gt;CLAUDE.md&lt;/code&gt; as documentation, a notebook about my server that the agent happens to read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It isn't documentation. It's a prompt that runs every session, and every sentence in it is a sentence the model acts on.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once you see it that way, the findings stop looking like housekeeping:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;History in a prompt isn't a record, it's noise the model has to work around.&lt;/strong&gt; "Removed on 2026-08-25, backup at this path, 82 days of rank history pruned" reads as context to me. To a model that treats everything in its instructions as relevant, it's a pile of half-relevant facts competing with the ones that matter. The one useful rule hidden in that paragraph ("these jobs were removed on purpose, so check before adding them back") was the part hardest to find.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A stale fact in a prompt isn't out-of-date documentation, it's a wrong instruction.&lt;/strong&gt; A wiki page with an old version number just sits there. The same line in a prompt gets acted on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A how-to in a prompt isn't a note, it's a standing order.&lt;/strong&gt; "Put the token in the URL" stopped being a description of how I once got a push to work. It became the default.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The audit's own guide puts it in one line: the job is to find "specific instructions that no longer fit", not to make prompts shorter. My file didn't have an old-model problem. It had the problem every long-lived config file gets: it kept accumulating and nothing ever removed anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;p&gt;The patch had nine hunks, each tied to one finding, so I could take or skip them one at a time. I took all of them.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Moved the history out.&lt;/strong&gt; The 880-word paragraph went verbatim into a separate &lt;code&gt;cron-history.md&lt;/code&gt; that isn't loaded automatically. The cron header is now one line with the job count and a pointer: "read it before re-adding a removed job." The schedule itself stayed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixed the facts.&lt;/strong&gt; Dropped the Gemini line, corrected the version, listed the real key names, and added the missing voice-agent tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kept rules, dropped stories.&lt;/strong&gt; Several sections had a current rule with an incident wrapped around it, like "Seen on 2026-09-22: a session started at 14:48 picked up the new package but still ran the old code…" The rule (restart the process after a config change, because clearing the session isn't enough) stayed, with its reason. The timestamped story went.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixed the token rule properly, not just the wording.&lt;/strong&gt; I set git to use the GitHub CLI's existing login for pushes (&lt;code&gt;gh auth setup-git&lt;/code&gt;), confirmed a dry-run push worked, and only then replaced the line with "plain &lt;code&gt;git push&lt;/code&gt;, never put a token in the remote URL." Then I checked every local repo for a token in its remote URL. There were none.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two smaller things went too. I removed account-balance figures from a list of fallback models, since balances drift and nothing on the machine could confirm them. I also moved four skills for a framework I'd stopped using out of the global skills folder: their descriptions were loading into every project.&lt;/p&gt;

&lt;p&gt;The file went from about 3,700 words to 2,760. That wasn't the goal, just a side effect of removing what didn't belong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the audit left alone, and why that matters
&lt;/h2&gt;

&lt;p&gt;The audit has an explicit list of things it must not delete, and it followed it. It kept:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the rules that carry a reason, like "don't re-optimize this page's title, the low click rate comes from the search results layout, not the title";&lt;/li&gt;
&lt;li&gt;two identical copies of a scaffold file in two project folders, because they agree;&lt;/li&gt;
&lt;li&gt;my one-line note that I prefer short, direct answers, because that's context about me, not an outdated trick.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That restraint is the part I'd have got wrong doing this by hand. My instinct with a bloated file is to cut it hard, and a hard cut removes exactly the lines that only I could have written: the reasons. The model can work out how to be thorough by itself. It can't work out why I don't trust a particular page's metrics.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you keep one of these files
&lt;/h2&gt;

&lt;p&gt;Some things I'm doing differently now:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Write the rule, not the story.&lt;/strong&gt; If a line describes something that happened, ask what the agent should &lt;em&gt;do&lt;/em&gt; differently because of it, and write that sentence instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give history its own file.&lt;/strong&gt; A changelog is valuable. It just shouldn't load into every session. Point to it from the briefing file, and the agent can read it when it's relevant.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Check facts against the machine, not your memory.&lt;/strong&gt; Paths, versions and key names drift silently. The audit caught four because it checked them, not because they looked wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read every how-to as a standing order.&lt;/strong&gt; If you wouldn't want the agent to do it every time without asking, it doesn't belong in the file as an instruction.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The file is shorter now, but that isn't the improvement. The improvement is that everything left in it is something I actually want the agent to act on.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
    <item>
      <title>What Is Jev? I Tested TypeSafe's Decision Model on My Comments and Email</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Sun, 04 Oct 2026 03:13:48 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/what-is-jev-i-tested-typesafes-decision-model-on-my-comments-and-email-4l1h</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/what-is-jev-i-tested-typesafes-decision-model-on-my-comments-and-email-4l1h</guid>
      <description>&lt;p&gt;In September 2026 a company called TypeSafe released Jev, and it isn't a chatbot. It never writes a sentence. You give it some text and a question with fixed answer options, and it returns one of your options plus how sure it is: "bot, 97%." Nothing to parse, nothing to read.&lt;/p&gt;

&lt;p&gt;I spent a day testing it on two real jobs: sorting the comments people leave on my blog posts, and sorting my email. This post covers what Jev is, how it works, what it's good and bad at, exactly how far I tested it, and the one thing that decided whether it worked, which turned out not to be the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Jev is
&lt;/h2&gt;

&lt;p&gt;TypeSafe calls Jev a "System One" model, borrowing the psychology term for fast, intuitive judgment as opposed to slow, deliberate reasoning. Chat models like Claude or GPT reason and write. Jev does one job: decide.&lt;/p&gt;

&lt;p&gt;That makes it a component, not an assistant. You don't talk to it. You put it inside a script at the point where plain code can't make a call. Is this comment spam? Does this email need me? Which team should this support ticket go to? Code can't answer those reliably with rules, and sending each one to a full chat model is slow and expensive for a decision you need thousands of times.&lt;/p&gt;

&lt;p&gt;What it can't do matters just as much:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It doesn't write.&lt;/strong&gt; No summaries, no replies, no explanations of its answer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It doesn't know facts.&lt;/strong&gt; It has no web access and isn't a knowledge base. It can't tell you whether a law changed or a price is right. It only judges the text you hand it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It only picks from your options.&lt;/strong&gt; If none of your options fits, it still picks one. It just tells you it isn't sure.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;One HTTP request. You send three things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;state&lt;/strong&gt;: the text to judge. It can be a plain string or named fields.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;questions&lt;/strong&gt;: one or more typed questions, all answered in the same call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;model&lt;/strong&gt;: &lt;code&gt;jev-latest&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are three question types:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;What you get back&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;choice&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;One option from your list, a probability for every option, and a confidence score&lt;/td&gt;
&lt;td&gt;"Is this comment genuine, promo, or bot?"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;noul&lt;/strong&gt; (yes/no)&lt;/td&gt;
&lt;td&gt;The probability of "yes"&lt;/td&gt;
&lt;td&gt;"Does this comment ask the author a question?"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;score&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A position on a scale you define (2–10 levels)&lt;/td&gt;
&lt;td&gt;"How frustrated is this customer: calm, frustrated, very angry?"&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here's a real request from my comment filter, shortened:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jev-latest"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"post_title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"…"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"comment"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"We need to write a short casual YouTube comment as a regular developer…"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"questions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Classify `comment`, left on the developer blog post titled `post_title`."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"criteria"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"genuine"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"A reader discussing the post's ideas, with no link to or plug for their own project"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"promo"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Mentions, links to, or plugs the commenter's own project, tool, repo, product or site"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"bot"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Reads like instructions written to an AI or leaked prompt text, or generic filler unrelated to the post"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"asks_question"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Does `comment` ask the author a direct question or explicitly invite a reply?"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the answer comes back as data:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"jev-1.13.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"answers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"choice"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bot"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"probabilities"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"…"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"one per option"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.97&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"asks_question"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"noul"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.79&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;criteria&lt;/code&gt; text is where all the work is. Hold that thought.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Price and speed.&lt;/strong&gt; You pay $0.042 per million tokens Jev reads; output is free because it writes nothing. In my tests a call took about 0.4 seconds (412 ms to 964 ms, almost all close to 430 ms).&lt;/p&gt;

&lt;h2&gt;
  
  
  How far I tested it
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Messages&lt;/th&gt;
&lt;th&gt;Questions per message&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Result after fixing the wording&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Blog comments&lt;/td&gt;
&lt;td&gt;15 real comments on my dev.to posts&lt;/td&gt;
&lt;td&gt;2 (choice + yes/no)&lt;/td&gt;
&lt;td&gt;$0.0004&lt;/td&gt;
&lt;td&gt;15 of 15 sensible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Email&lt;/td&gt;
&lt;td&gt;My last 50 Gmail Primary emails&lt;/td&gt;
&lt;td&gt;2 (choice + yes/no)&lt;/td&gt;
&lt;td&gt;$0.002&lt;/td&gt;
&lt;td&gt;All payout, security and action emails flagged; codes, receipts and newsletters sorted to routine&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Security question&lt;/td&gt;
&lt;td&gt;12 account-related emails&lt;/td&gt;
&lt;td&gt;1 (yes/no)&lt;/td&gt;
&lt;td&gt;under $0.001&lt;/td&gt;
&lt;td&gt;Real alerts 75–96%, everything else 47% or lower&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Everything, including the failed first attempts, cost under one cent. I didn't test the &lt;strong&gt;score&lt;/strong&gt; type, very long documents, or high volume. Treat this as one person's real-data trial, not a benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it failed first
&lt;/h2&gt;

&lt;p&gt;On the first run, Jev called a bot &lt;strong&gt;genuine&lt;/strong&gt;. The bot had pasted its own instructions instead of a comment: &lt;em&gt;"We need to write a short casual YouTube comment as a regular developer… Must start with specific reaction… Use lowercase start."&lt;/em&gt; Any human spots it instantly. Jev also called two self-promotional comments genuine: one plugged the writer's GitHub tool, the other said "I work on [product]."&lt;/p&gt;

&lt;p&gt;The confidence numbers told the story. On the twelve clear-cut reader comments Jev was 89–100% sure. On the bot it was 43% sure, and on the GitHub plug 34%. It wasn't confidently wrong. It was saying my options didn't fit.&lt;/p&gt;

&lt;p&gt;My first descriptions were judgments: genuine meant "engaging with the post's ideas in their own words", promo meant "mainly promotes the commenter's own product". Deciding whether something "mainly" promotes is exactly the call I wanted made for me. So I rewrote each option as something you can see in the text, the wording in the request above.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Comment&lt;/th&gt;
&lt;th&gt;First wording&lt;/th&gt;
&lt;th&gt;Concrete wording&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Bot that leaked its prompt&lt;/td&gt;
&lt;td&gt;genuine, 43%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;bot, 97%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reader plugging a GitHub tool&lt;/td&gt;
&lt;td&gt;genuine, 34%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;promo, 99%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"I work on [product]…"&lt;/td&gt;
&lt;td&gt;genuine, 81%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;promo, 100%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Two good comments ending in a link to the author's own site&lt;/td&gt;
&lt;td&gt;genuine&lt;/td&gt;
&lt;td&gt;genuine, 50–60%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The other ten&lt;/td&gt;
&lt;td&gt;genuine&lt;/td&gt;
&lt;td&gt;genuine, 93–100%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The unsure row is my favourite. Those two comments make real points and then link the writer's own site. Is that promotion? I'm not sure either, and Jev dropped to a coin flip on exactly those two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model wasn't judging the comments. It was matching them against my descriptions, and my descriptions were the thing that was wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Email repeated the lesson. Sign-in codes I'd requested myself came back as "needs me" (71–98% sure), because my description mentioned "a verification he must complete". Webmaster "improve your site" tips came back the same way. Explicit exclusions fixed both: "NOT one-time sign-in or verification codes, and NOT tips or suggestions from tools."&lt;/p&gt;

&lt;p&gt;The security question was the scary one. Google's "A new sign-in on Linux" alert scored 9%:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question wording&lt;/th&gt;
&lt;th&gt;Google sign-in alert&lt;/th&gt;
&lt;th&gt;Google app-password alert&lt;/th&gt;
&lt;th&gt;Sign-in codes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"Is this a security alert about Ted's own account (a new sign-in, password or recovery change…)?"&lt;/td&gt;
&lt;td&gt;9%&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;td&gt;9–22%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same question plus "true" and "false" descriptions&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;td&gt;3–30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Does &lt;code&gt;body&lt;/code&gt; tell Ted that something changed in his account's access: a new sign-in, a new device, a new password or app password…?"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;75%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3–20%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two ask whether the email belongs to a category, "security alert". The third asks whether the text reports a specific event you can point at.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building something safe with it
&lt;/h2&gt;

&lt;p&gt;Even well-worded, Jev will sometimes be wrong. How you use its answer matters more than its accuracy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tag, don't filter.&lt;/strong&gt; My comment watcher still sends me every comment, now tagged 🔗 promo or ❓ asks-you-a-question. The only thing it drops is a bot verdict at 85%+ confidence, and those go to a log I can check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Low confidence means "show me".&lt;/strong&gt; Anything under 70% gets an "unsure" tag instead of being acted on.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure means the old behaviour.&lt;/strong&gt; If the API errors, times out or my credit runs out, messages go through untagged. Jev can improve an alert. It can't make one disappear.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Back up the critical cases with plain rules.&lt;/strong&gt; Beside Jev's security score, a simple check flags any subject containing "security alert", "new sign-in", "password" or "passkey".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the questions and thresholds in one block.&lt;/strong&gt; They decide behaviour, so they're what to review. TypeSafe's own docs give the same advice.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both scripts are live now. My email watcher pings me for security, money and anything that needs me, and puts codes, receipts and newsletters into one evening summary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you try it?
&lt;/h2&gt;

&lt;p&gt;If you have a script that has to make many small yes/no or which-bucket decisions about text, and you're currently either writing fragile keyword rules or paying a chat model to answer one word at a time, Jev is worth an afternoon. Testing costs fractions of a cent.&lt;/p&gt;

&lt;p&gt;Just budget the afternoon for the questions, not the integration. Describe what's visible in the text, not your conclusion about it. Test on your own messages. And read the confidence column as carefully as the answers: a low number usually means your options don't fit the case.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>automation</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>My Site Became the Source, So My Mistake Became the Answer</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Thu, 01 Oct 2026 18:08:28 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/my-site-became-the-source-so-my-mistake-became-the-answer-3g22</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/my-site-became-the-source-so-my-mistake-became-the-answer-3g22</guid>
      <description>&lt;p&gt;I run a small travel directory. One of its listings is a boutique hotel whose owners hold a licence for an on-site cannabis lounge. For months the listing said the hotel &lt;em&gt;has&lt;/em&gt; a licensed lounge, and recommended it to anyone visiting in winter who wanted somewhere warm to consume.&lt;/p&gt;

&lt;p&gt;During a routine review of the pages that list it, that claim started to look shaky, so I began checking it. Partway through, I asked an AI search engine the question a traveller would ask: is this hotel cannabis-friendly?&lt;/p&gt;

&lt;h2&gt;
  
  
  The answer was mine
&lt;/h2&gt;

&lt;p&gt;Perplexity said yes, with a tidy breakdown. Consumption only in the licensed lounge. Bring your own. A capacity of about 25 people, described as upscale and private. It closed with practical advice: if you plan to consume, do it in the lounge and bring your own.&lt;/p&gt;

&lt;p&gt;Nearly every line was cited to my site. The 25-person capacity was cited to my site alone, because no other page on the internet says it.&lt;/p&gt;

&lt;p&gt;That sent me back to the listing to see where the number came from. It was in the listing's description, along with "a curated social hour" and "knowledgeable staff". None of it had a source. It reads like the lounge's plans, written up at some point as if they were the lounge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Licensed is not open
&lt;/h2&gt;

&lt;p&gt;The checking I'd already started was the other half of the answer. The only way to get the lounge's current state was to line sources up by date:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it says&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2022&lt;/td&gt;
&lt;td&gt;City approval, widely reported&lt;/td&gt;
&lt;td&gt;Preliminary approval for the licence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Early 2023&lt;/td&gt;
&lt;td&gt;Local news&lt;/td&gt;
&lt;td&gt;Not open yet; the ventilation design still being redrafted&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spring 2025&lt;/td&gt;
&lt;td&gt;Same outlet&lt;/td&gt;
&lt;td&gt;The owner about four years into trying to open it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Early 2026&lt;/td&gt;
&lt;td&gt;A local lounge directory&lt;/td&gt;
&lt;td&gt;Still being built&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Today&lt;/td&gt;
&lt;td&gt;The owners' own website&lt;/td&gt;
&lt;td&gt;Describes the lounge as pending&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every source agreed the licence is real. None of them said the lounge was open. The owners' own FAQ, written in the future tense, adds the part that matters most to a guest: cannabis is only allowed in that lounge, nowhere else on the property.&lt;/p&gt;

&lt;p&gt;My listing had collapsed two different facts into one. &lt;em&gt;Licensed&lt;/em&gt; is a state that was true in 2022 and is still true. &lt;em&gt;Open&lt;/em&gt; is an event that hasn't happened. The owners' marketing blurs them too, which is how the listing got written that way in the first place, and how an AI reading my page repeated it without hesitation. To a hurried researcher and to a language model, "holds a licence for a lounge" and "has a lounge" are the same sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Being cited is not being checked
&lt;/h2&gt;

&lt;p&gt;This is the part that changed how I think about the site.&lt;/p&gt;

&lt;p&gt;When a page is one source among many, its mistakes get diluted. When it becomes the source, the one an answer engine leans on for a niche question, its mistakes get passed on whole, with a citation attached that makes them look checked. The answer engine didn't verify the lounge. It verified that my page said so. The citation proved where the claim came from, not that it was true.&lt;/p&gt;

&lt;p&gt;And the people asking aren't careless. Most travellers will call a hotel before booking around a feature like this. But the one who doesn't arrives expecting a lounge that isn't there, and the page they trusted, and the answer engine that quoted it, are both wrong in the same way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first fix was too honest
&lt;/h2&gt;

&lt;p&gt;My first correction, drafted with the AI agent I use for site work, swung hard the other way. It said the lounge hadn't opened, and quoted the owners' rule that cannabis is allowed nowhere else on the property, on every page that mentioned the hotel.&lt;/p&gt;

&lt;p&gt;Every word of it was accurate, and it read like a warning. When I read it as a guest would, the problem was obvious: we'd taken a listing people were booking and turned it into a reason not to. The hotel is real, well reviewed, run by owners building around exactly this kind of guest, and holding a licence very few hotels have. None of that had stopped being true.&lt;/p&gt;

&lt;p&gt;So the second version keeps everything that's true and attractive, and changes only the one claim that wasn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A well-reviewed historic hotel whose owners hold a licence for an on-site cannabis lounge, opening soon. Until it opens, ask the hotel where cannabis is allowed before you book.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The verification label changed too, from "verified" to a new status that says exactly that: licensed, opening soon. The old system only had verified, host-declared and unconfirmed, and none of them could describe a property where we know exactly what the situation is and it simply hasn't finished happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I took from it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Write states with their dates.&lt;/strong&gt; "Holds a licence" and "opening soon" stay true until something changes. "Has a lounge" was false from the day it was written.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check a claim with a timeline, not a source.&lt;/strong&gt; Any single source here would have told me the lounge was real. Only the sequence told me it had been "about to open" for four years.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Know what you're the only source for.&lt;/strong&gt; The detail that only my page contained, the 25-person capacity, was the one with no source at all. If you're the only page saying something, nobody downstream will catch it for you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fix the claim, not the listing.&lt;/strong&gt; The correction that tells people what's true and still gives them a reason to go is better for the reader and the business than the one that just warns them off.&lt;/p&gt;

&lt;p&gt;The answer engine will catch up when it next reads the page. Until then, it's still quoting the old sentence, with my name on it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>seo</category>
      <category>webdev</category>
      <category>writing</category>
    </item>
    <item>
      <title>The Reviewer Checked the Wrong Pages</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Mon, 28 Sep 2026 11:46:09 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/the-reviewer-checked-the-wrong-pages-49gm</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/the-reviewer-checked-the-wrong-pages-49gm</guid>
      <description>&lt;p&gt;I run a small team of AI agents that answers research questions. One agent searches the web and takes notes, one writes the answer, and a third, the reviewer, checks the write-up before I see it. The reviewer is the expensive part, so I tested three versions of it on the same draft. That draft had a real mistake in it: it credited a hardware claim to a source that didn't say it.&lt;/p&gt;

&lt;p&gt;The expensive reviewer, Claude Sonnet at high effort, caught the mistake for about three cents. The cheap one, Claude Haiku, missed it for about one cent. I kept Haiku and gave it a new ability: before deciding, open two or three of the pages the draft cites and check they really say what the draft claims. Then I wrote that this closed the gap.&lt;/p&gt;

&lt;p&gt;I never re-ran the test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The comment
&lt;/h2&gt;

&lt;p&gt;A reader summed up the idea better than I had: the cheap reviewer didn't need to be smarter, it needed something to be wrong against. A reviewer that only reads the draft can judge whether it sounds right. One that holds the source can actually fail it.&lt;/p&gt;

&lt;p&gt;That's true, and it made me notice I had described a mechanism, not a result. So I pulled the original flawed draft out of the run log and gave it to the new reviewer. Same draft, same notes, same mistake. The only change was the ability to open pages.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;I ran it twice, plus once with the page-opening tool switched off to recreate the original setup.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Reviewer&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;th&gt;Caught the mistake?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Haiku, no tools (the original setup)&lt;/td&gt;
&lt;td&gt;revise, for other reasons&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku that can open pages, run 1&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;approve&lt;/strong&gt;, no notes&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku that can open pages, run 2&lt;/td&gt;
&lt;td&gt;revise, for other reasons&lt;/td&gt;
&lt;td&gt;no&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The version with the new ability did worse. On one run it approved the flawed draft outright.&lt;/p&gt;

&lt;p&gt;The reason was in the draft itself. It cited sources as "[5]", with no link. The citations only started carrying URLs the next day, so the reviewer had nothing it could actually open. Instead of saying so, it went looking. It opened an Apple specification page, a hardware review, a code repository: six pages that seemed relevant to the topic and weren't the source of anything in the draft. It read them, found nothing that contradicted the draft, and signed off.&lt;/p&gt;

&lt;p&gt;That is worse than the reviewer with no tools. The no-tools reviewer never claimed to have checked a source. This one did the motions of a source check against pages that couldn't fail the claim, and came back more confident, not less.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Two rules, and the second one had to live in code rather than in the instructions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Only the cited sources.&lt;/strong&gt; Before reviewing, the system collects every URL that actually appears in the draft and the research notes, and the reviewer is told to open only those. A request to open any other page is refused before anything is fetched.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Say "not checkable" when you can't check.&lt;/strong&gt; If the draft cites nothing with a link, the reviewer is told not to search or guess. Its notes always begin with "Citations not checkable: no source URLs", and that line is added by the code if the model leaves it out. A sign-off can no longer look like a source check that didn't happen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On the same link-less draft, the fixed reviewer opened nothing, both times, and both reviews started with that line.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the source check actually adds
&lt;/h2&gt;

&lt;p&gt;Fixing the failure still left the original question: does opening the sources help the cheap reviewer catch mistakes? So I took a recent draft that does cite real links and planted a false claim in it: "Node 22 also dropped support for Windows 10 [1]." The cited source is the official release announcement, which says nothing of the kind, and the research notes don't mention Windows at all.&lt;/p&gt;

&lt;p&gt;Every version caught it, including the one with no tools. The no-tools reviewer noticed that the research notes never mention Windows. The version that opened the sources went further: it named the official announcement and said the claim wasn't in it, which is stronger evidence and cost about two and a half times as much.&lt;/p&gt;

&lt;p&gt;So in my tests, opening the sources didn't catch anything that comparing the draft against the research notes missed. It produced better proof, not more catches. The mistake that started all this, a hardware claim sitting right next to figures the notes really did contain, was missed by the cheap reviewer in every setup I tried. Only the expensive one has caught it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell myself three days ago
&lt;/h2&gt;

&lt;p&gt;A tool is not a check. A reviewer that can open pages will open some, and if the right ones aren't available it will find pages that look close enough, and treat not finding a problem as finding none. The ability to look only helps when there is something specific to look at, and when "I couldn't look" is an answer the system is allowed to give.&lt;/p&gt;

&lt;p&gt;The smaller lesson is the one the reader's comment forced on me: the only way to know what a fix does is to re-run the case that made you want it. I had the flawed draft saved the whole time. Writing "this closed the gap" took less effort than checking whether it had.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>testing</category>
    </item>
    <item>
      <title>It Looked Finished on Day One</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Sun, 27 Sep 2026 02:36:30 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/it-looked-finished-on-day-one-4k0j</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/it-looked-finished-on-day-one-4k0j</guid>
      <description>&lt;p&gt;I wanted a team of AI agents that could answer a research question the way a small team of people would. One plans the work, one searches the web and reads the sources, one writes it up, and one checks the write-up before I see it. I also wanted to watch them do it: which agent is working, what it's reading, what it handed to whom, and what it cost.&lt;/p&gt;

&lt;p&gt;The app is called FORGE, and it runs on a small server in my house. This post covers how it got from an idea to something that works end to end, which took three days. Most of that time wasn't building features. It was finding the places where the app looked like it worked and didn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Day one: the part that looks finished
&lt;/h2&gt;

&lt;p&gt;The first version was a dashboard for a system that didn't exist yet.&lt;/p&gt;

&lt;p&gt;It had a workflow canvas where each agent is a card and each arrow is a handoff. It had a live event stream, a timeline, a replay scrubber, per-agent cost and token counters, an agent builder and a workflow designer. Behind it sat a simulator: fake agents on a fake clock, producing realistic-looking events so every screen had something to show.&lt;/p&gt;

&lt;p&gt;One design decision from that first day held up for everything that came after. &lt;strong&gt;A run is a log of events, not a record that gets updated.&lt;/strong&gt; "Researcher started", "tool call finished", "handed the notes to the writer": every screen is computed from that log. Live view and replay are the same code; replay just stops reading the log partway through. When real agents replaced the simulator later, the screens didn't change at all, because real agents emit the same events.&lt;/p&gt;

&lt;p&gt;It looked gorgeous, and none of it was real.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making the agents real
&lt;/h2&gt;

&lt;p&gt;Real agents meant real model calls, and I wanted them cheap. I started with three agents: a researcher, a writer and a reviewer.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Models:&lt;/strong&gt; most agents run GLM-5.3 Flash through OpenRouter, at about $0.15 per million input tokens. The reviewer runs on Claude.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Search:&lt;/strong&gt; DuckDuckGo through the &lt;code&gt;ddgs&lt;/code&gt; Python package. It's free and needs no key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reading pages:&lt;/strong&gt; a small script fetches a URL and pulls out the main text (trafilatura for web pages, pypdf for PDFs).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A server:&lt;/strong&gt; a small Node.js server runs the agents, stores every event, and streams them to the browser as they happen.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first live runs failed in ways the simulator never could.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every call timed out, instantly.&lt;/strong&gt; The server has no IPv6. When a hostname has several addresses, Node.js tries them one after another and gives each attempt only 250 milliseconds by default, which isn't enough on this network. The fix was two lines: prefer IPv4, and give each attempt three seconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Some replies came back empty.&lt;/strong&gt; The model had spent its whole output budget reasoning and had nothing left for the answer. It now gets one retry with double the room.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The researcher never stopped researching.&lt;/strong&gt; It kept searching and reading until it hit the step limit, then had no notes. Now every tool result tells it how many rounds it has left, and the last round forces it to write up what it has.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Some websites never answered.&lt;/strong&gt; One slow host could eat a minute. Page fetches now give up after 20 seconds, and a host that times out once is skipped for the rest of the run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honesty pass
&lt;/h2&gt;

&lt;p&gt;With real agents working, I went through the app page by page and asked one question: is this real or simulated?&lt;/p&gt;

&lt;p&gt;A lot of it was simulated and looked real. The model-provider page showed made-up usage. The tools page listed tools with success rates and response times for tools that had never run. The run history was full of sample runs.&lt;/p&gt;

&lt;p&gt;Then I asked the team a real question and got a confident, well-formatted answer to a completely different question. The run had used the simulator, which produces plausible text and ignores the question entirely. It was the default for most workflows, and nothing on screen said so.&lt;/p&gt;

&lt;p&gt;The rule I took from it: &lt;strong&gt;never ship a placeholder that looks real.&lt;/strong&gt; Label it or remove it. Simulated data is now clearly marked everywhere. The simulator survives only as a "Dry run" button in the workflow designer, for checking the wiring for free. You can't start a simulated run from the New Run screen at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  One store, not one per browser
&lt;/h2&gt;

&lt;p&gt;I use FORGE from a desktop and a laptop, and they showed different data. Everything lived in each browser's local storage, so each device had its own private copy.&lt;/p&gt;

&lt;p&gt;The server now owns a single SQLite database. Every change goes to the server, and every open tab gets pushed updates over a live connection. Each browser uploaded its old local data once, and the server merged it.&lt;/p&gt;

&lt;p&gt;This surfaced a subtler bug. Run IDs had been picked by each browser, and the laptop picked an ID the desktop had already used, overwriting a run. Now only the server issues run IDs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Spreading the cost
&lt;/h2&gt;

&lt;p&gt;I added a planner in front of the researcher. It writes a short plan with sub-questions and search queries, costs about a hundredth of a cent, and makes the research noticeably more focused.&lt;/p&gt;

&lt;p&gt;Then I checked the bill. One OpenRouter key was paying for everything, and it had $1.74 left. So Claude models now go straight to Anthropic's API, and everything else stays on OpenRouter. If Anthropic's key fails on billing, the call falls back to OpenRouter instead of failing the run.&lt;/p&gt;

&lt;p&gt;The reviewer was the expensive part. I tested three reviewers on the same flawed draft, one with a real error: it credited a performance claim to the wrong source.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Sonnet, high effort:&lt;/strong&gt; caught the error. About $0.03 per review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sonnet, lower effort:&lt;/strong&gt; approved the flawed draft. About $0.014.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Haiku:&lt;/strong&gt; sent the draft back for other reasons but missed that error. $0.009.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I chose Haiku. The next section is what closed that gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  Making it trustworthy
&lt;/h2&gt;

&lt;p&gt;One run hit seven failed web searches in a row. An outside review of the logs blamed DuckDuckGo throttling and suggested switching to a paid search API.&lt;/p&gt;

&lt;p&gt;The logs said otherwise. The search library rotates between several engines, and all four timed out in the same three minutes. The same queries worked in two to four seconds when I re-ran them. It was a short network outage on my side, and a paid API would have timed out too. So instead of switching providers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Search fails fast and retries:&lt;/strong&gt; a 12-second limit, one retry, and after three failures in a row a 30-second wait for the network to come back.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not enough sources is flagged in code, not just the prompt:&lt;/strong&gt; if the researchers read fewer than two pages, the answer carries a warning banner saying it's unverified. The writer is told too, but the flag doesn't depend on it listening.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The reviewer checks the sources itself:&lt;/strong&gt; it opens two or three of the pages the draft cites and checks they say what the draft claims. That's what the cheap reviewer was missing: a source it could check against. Pages the researcher already fetched are reused, so the check costs no extra time.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Reaching it from my phone
&lt;/h2&gt;

&lt;p&gt;I already run a personal assistant agent called Hermes that talks to me on Telegram. FORGE and Hermes now connect both ways.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FORGE → Telegram:&lt;/strong&gt; every finished run sends the answer, cost, time and a link to the run page. Failed runs send an ❌ with the error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Telegram → FORGE:&lt;/strong&gt; a &lt;code&gt;/forge&lt;/code&gt; command in Hermes hands a question to the team. Hermes can only wait three minutes for a command, and a run takes longer. So the command starts the run and returns straight away, and the answer arrives on Telegram when the team finishes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are now three levels: &lt;strong&gt;quick&lt;/strong&gt; (no reviewer, about half a cent, answers marked "not fact-checked"), &lt;strong&gt;verified&lt;/strong&gt; (the default, two to four cents) and a larger six-agent team for deep questions. I pick the level myself. I considered a classifier that decides whether a question needs checking. It would have been cheap, but it would sometimes skip checking on exactly the question that mattered.&lt;/p&gt;

&lt;h2&gt;
  
  
  Letting the agents talk
&lt;/h2&gt;

&lt;p&gt;Handoffs only go one way: the writer gets the notes, and that's it. If the notes leave a gap, the writer can only guess or leave it out.&lt;/p&gt;

&lt;p&gt;So agents can now ask a teammate a question mid-task. The writer can ask the researcher, and the researcher answers from its notes or does one quick lookup. In one run the writer spotted two researchers giving different numbers for the same fact and asked one to settle it. The researcher went back to the source and answered with the bill that set the figure.&lt;/p&gt;

&lt;p&gt;I also added a parallel team: a lead splits the question into three angles, three researchers work at once and post findings to a shared board, then the writer and reviewer finish.&lt;/p&gt;

&lt;p&gt;Its first run failed at the 20-minute limit while every agent was busy. One researcher's final write-up had taken five minutes on its own, because the model spent it reasoning. My stall detector watched for a stream that goes quiet, and a reasoning model streaming its thoughts never goes quiet.&lt;/p&gt;

&lt;p&gt;I measured the same three-sentence prompt. With no reasoning effort set, the model took 93 seconds and 3,200 reasoning tokens. With &lt;code&gt;effort: low&lt;/code&gt;, it took 9 seconds. Leaving the setting out didn't mean a sensible default; it meant no limit. After setting it for every call, the same parallel question finished in 5 minutes 10 seconds for four cents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching it from the outside
&lt;/h2&gt;

&lt;p&gt;The last piece was monitoring. I have a dashboard, Operator Pulse, that shows every scheduled job on the server and alerts me when one fails. FORGE now appears on it as a service: whether the server is up, recent runs, success rate, and how much OpenRouter credit is left.&lt;/p&gt;

&lt;p&gt;The first time the new FORGE node rendered, it was amber: &lt;strong&gt;$0.47 of credit left.&lt;/strong&gt; That key also pays for Hermes, so both would have stopped working without warning.&lt;/p&gt;

&lt;p&gt;FORGE can also run scheduled questions now. Each one is a small script that the dashboard tracks like any other job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it ended up
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0fm97msah0ajcoqze4q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe0fm97msah0ajcoqze4q.png" alt="A finished Quick Research run in FORGE: the planner, web researcher and report writer cards on the canvas, a progress bar reading 3 of 3, a total cost of $0.0012, and the live event stream on the right" width="800" height="390"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;A finished quick run: three agents, 1 minute 40 seconds, about a tenth of a cent. Every event on the right can be replayed.&lt;/em&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What runs&lt;/th&gt;
&lt;th&gt;Typical cost&lt;/th&gt;
&lt;th&gt;Typical time&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Quick&lt;/td&gt;
&lt;td&gt;planner → researcher → writer&lt;/td&gt;
&lt;td&gt;$0.005&lt;/td&gt;
&lt;td&gt;1–2 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Verified&lt;/td&gt;
&lt;td&gt;+ a reviewer that checks cited pages&lt;/td&gt;
&lt;td&gt;$0.02–0.04&lt;/td&gt;
&lt;td&gt;1–4 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Parallel&lt;/td&gt;
&lt;td&gt;lead → 3 researchers at once → writer → reviewer&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;td&gt;~5 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkhehrmd48ovh9y3809vz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkhehrmd48ovh9y3809vz.png" alt="The Runs list in FORGE: four live runs with their workflow, duration, agent count, tokens and cost, from $0.0008 for a quick run to $0.468 for the six-agent team" width="800" height="385"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every run is fully replayable. Every answer lists the pages it came from, and says so plainly when it couldn't verify enough of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What building it taught me
&lt;/h2&gt;

&lt;p&gt;The dashboard was done on day one. It had every screen, every counter and every animation, and it was all simulated. Everything after that was the same job repeated: find a place where the app &lt;em&gt;looks&lt;/em&gt; like it works and make it actually work.&lt;/p&gt;

&lt;p&gt;Some of those places were loud: timeouts, empty replies, a run that died at twenty minutes. The ones that mattered most were quiet: a confident answer to the wrong question, two devices each sure their data was the real one, a timeout that never fired because the model never stopped talking, a credit balance nobody was watching.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A demo shows that something can work. Making it trustworthy meant finding every place it only looked like it worked.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The Parameter Was Never Read</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Mon, 21 Sep 2026 09:22:32 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/the-parameter-was-never-read-1pca</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/the-parameter-was-never-read-1pca</guid>
      <description>&lt;p&gt;Every morning at 8:25 a script reads this blog's numbers out of Google Search Console — the tool Google gives site owners to see which searches showed their pages — and sends me a short digest on Telegram. Clicks, impressions, average position. Then two lists: top queries, top pages.&lt;/p&gt;

&lt;p&gt;I built it on day one, before the site had a single impression to report. It has run every morning since.&lt;/p&gt;

&lt;p&gt;Yesterday I was looking at something else entirely when the page list caught my eye.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;📄 Top pages
  / — 2 impr | 0 clicks | pos 8.0
  /contact — 2 impr | 0 clicks | pos 2.0
  /operations — 1 impr | 0 clicks | pos 7.0
  /posts/a-second-check-is-a-second-source-of-truth — 1 impr | 0 clicks | pos 9.0
  /posts/automating-seo-monitoring-with-ai-agents — 2 impr | 0 clicks | pos 11.5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The site had 290 impressions that week. The top page, according to the section labelled top pages, had two of them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sort I had been asking for
&lt;/h2&gt;

&lt;p&gt;The script asks Search Console for pages the obvious way. Give me pages, ten of them, ordered by impressions, biggest first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;sc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;startDate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;endDate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;dimensions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;page&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;rowLimit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;orderBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;fieldName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;impressions&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;sortOrder&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;DESCENDING&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That request has always returned 200. It has always returned rows. The rows have always been real — those pages exist, those impression counts are correct.&lt;/p&gt;

&lt;p&gt;I went to check what &lt;code&gt;orderBy&lt;/code&gt; actually does, and found that it does not exist.&lt;/p&gt;

&lt;p&gt;The request body the Search Analytics API accepts has ten fields: &lt;code&gt;aggregationType&lt;/code&gt;, &lt;code&gt;dataState&lt;/code&gt;, &lt;code&gt;dimensionFilterGroups&lt;/code&gt;, &lt;code&gt;dimensions&lt;/code&gt;, &lt;code&gt;endDate&lt;/code&gt;, &lt;code&gt;rowLimit&lt;/code&gt;, &lt;code&gt;searchType&lt;/code&gt;, &lt;code&gt;startDate&lt;/code&gt;, &lt;code&gt;startRow&lt;/code&gt;, &lt;code&gt;type&lt;/code&gt;. That is the whole list. Searching the client library's type definitions for the string &lt;code&gt;orderBy&lt;/code&gt; returns zero matches — not in the request, not anywhere in the Search Console surface.&lt;/p&gt;

&lt;p&gt;So for as long as this script has run, it has been attaching a field the API has no concept of. The API did not reject it. It did not warn. It read the fields it knew about, ignored the one it didn't, and answered.&lt;/p&gt;

&lt;h2&gt;
  
  
  What order the rows were in
&lt;/h2&gt;

&lt;p&gt;Here is the same window, asked for queries, printed exactly as the API returned them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; 1   1 astro-site-mu-eight.vercel.app
 2   2 chat with self hosted ai agent
 3   1 got html content but no text found (with 200 reply code)
 4   1 how do i run an ai agent locally?
 5   9 http://localhost:8766
 6   8 http://localhost:8766.
 7   6 http://localhost:8766/
 8   6 localhost:8766
 9   7 self hosted ai agent setup
10   2 this deployment is temporarily paused
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Alphabetical.&lt;/p&gt;

&lt;p&gt;The documented default is to sort by clicks, descending. Every row in that window has zero clicks. With nothing to rank by, what came back was ordered by key — and my report, which took the first five of whatever arrived, printed them under the heading &lt;strong&gt;Top queries&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Nothing about that output looks wrong. Five plausible search terms with plausible numbers beside them. The only thing separating it from a correct answer is that the word "top" was doing no work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that actually cost something
&lt;/h2&gt;

&lt;p&gt;Mis-ordering ten rows is cosmetic. I had ten queries that week and all ten came back, just shuffled.&lt;/p&gt;

&lt;p&gt;Pages were different. There were twenty-six pages with impressions that week, and &lt;code&gt;rowLimit: 10&lt;/code&gt; asked for ten. When the order is alphabetical, "the first ten" means the ten whose URLs sort earliest — &lt;code&gt;/&lt;/code&gt;, &lt;code&gt;/contact&lt;/code&gt;, &lt;code&gt;/operations&lt;/code&gt;, and whatever posts happen to start with &lt;code&gt;a&lt;/code&gt; and &lt;code&gt;b&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The busiest page on the site is &lt;code&gt;/posts/headless-oauth-loopback-callback&lt;/code&gt;. That week it took 155 impressions, more than half the site's entire total of 290. Its URL begins with &lt;code&gt;h&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It was not ranked low in the report. &lt;strong&gt;It was never in the response.&lt;/strong&gt; The truncation happened at the API, ten rows in, long before my code got to choose what to display. I had been reading a daily summary of my site's search performance that omitted the majority of it.&lt;/p&gt;

&lt;p&gt;Sorting the rows myself, after asking for all of them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;📄 Top pages
  /posts/headless-oauth-loopback-callback — 155 impr | pos 7.1
  /posts/the-audit-became-a-build-step — 51 impr | pos 5.9
  /posts/how-to-set-up-local-ai-agent — 17 impr | pos 37.4
  /posts/vercel-silent-build-failure — 10 impr | pos 12.9
  /posts/claude-built-my-astro-blog — 9 impr | pos 13.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same API, same week, same window. A different site.&lt;/p&gt;

&lt;h2&gt;
  
  
  It failed because the numbers were small
&lt;/h2&gt;

&lt;p&gt;This is the part I keep turning over.&lt;/p&gt;

&lt;p&gt;The fallback to alphabetical only happened because every row had zero clicks. On a site with traffic, the default sort — clicks, descending — would have put the busiest pages at the top on its own. The report would have been right. Not because the parameter worked, but because the API's default happened to agree with what I wanted.&lt;/p&gt;

&lt;p&gt;I would have gone on passing a field that does nothing, reading correct output, for as long as the site had clicks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The report was wrong in exactly the conditions it exists to report on.&lt;/strong&gt; It is a monitor for a site that is not yet getting traffic, and the absence of traffic is what broke it. Had it ever started working, it would have started working silently, and I would have had no more reason to check it then than I did on any of the mornings it was lying.&lt;/p&gt;

&lt;h2&gt;
  
  
  An ignored input is a default you didn't choose
&lt;/h2&gt;

&lt;p&gt;I have a rule on this site that &lt;a href="https://tedagentic.com/rules" rel="noopener noreferrer"&gt;defaults are policy&lt;/a&gt;: an unset value is not an absence, because something downstream always picks one. This is the same rule arriving from a direction I hadn't considered.&lt;/p&gt;

&lt;p&gt;I did set the value. I set it explicitly, in the request, with the right intent. It just went to a system that had no field to put it in, and a system with no field to put it in cannot tell you that you set nothing — it has already forgotten you tried.&lt;/p&gt;

&lt;p&gt;A strict API rejects unknown fields. A permissive one accepts everything and gives every request the same confident shape. That permissiveness is usually described as robustness. What it actually does is move the entire class of "you asked for something that doesn't exist" out of the error channel and into the results, where it arrives looking like an answer.&lt;/p&gt;

&lt;p&gt;I've written before about &lt;a href="https://tedagentic.com/posts/the-field-was-computed" rel="noopener noreferrer"&gt;an API that reported 44 updates and performed 24&lt;/a&gt;. That one accepted a real field and undid the write inside the same operation. This one is a step further back: the field was never real, so there was nothing to undo. Both return 200. Both hand you well-formed output. The difference only exists in documentation you have to go and read.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things I do now
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Check the parameter against the schema, not against the response.&lt;/strong&gt; A 200 tells you the request was accepted. It cannot tell you the request was understood. The only place that distinction lives is the field list, and reading it takes a minute — considerably less than the months I spent not reading it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sort where you control the sorting.&lt;/strong&gt; If ranking matters, pull a generous row limit and order the rows in your own code. Ordering done remotely is ordering you are trusting on someone else's terms, and — as here — possibly not happening at all. Ordering done locally is something you can look at.&lt;/p&gt;

&lt;p&gt;The report was never broken. Every morning it fetched real data and formatted it correctly and delivered it on time. It just answered a question I hadn't asked, under a heading that said I had.&lt;/p&gt;

</description>
      <category>api</category>
      <category>seo</category>
      <category>webdev</category>
      <category>debugging</category>
    </item>
    <item>
      <title>The Top Row Had 150 Clicks. It Wasn't a Listing.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Fri, 18 Sep 2026 18:24:51 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/the-top-row-had-150-clicks-it-wasnt-a-listing-21k9</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/the-top-row-had-150-clicks-it-wasnt-a-listing-21k9</guid>
      <description>&lt;p&gt;I track every click that leaves one of my sites. When someone taps a booking button, a small function records which listing it was for, which page it happened on and a few other details, and a dashboard adds them up. Its most useful panel is a table called Top Performing Entities: the listings people click most.&lt;/p&gt;

&lt;p&gt;I opened it to see which listings were earning their place. The top row had 150 clicks in 30 days, well ahead of second place. It carried the name of a real listing.&lt;/p&gt;

&lt;p&gt;It wasn't one. It was 29 different listings, stacked on top of each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five kinds of value in one column
&lt;/h2&gt;

&lt;p&gt;Each click row stores an &lt;code&gt;entity_id&lt;/code&gt; — the ID of the listing that was clicked. The dashboard groups by it. That's the whole design: count the rows per ID, sort, show the top ten.&lt;/p&gt;

&lt;p&gt;So I counted what the column actually held. Of the 806 clicks in the last 30 days:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;database UUID        441
"unknown"            157
slug or label        108
list position        100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Five different ideas of what an ID is — the slugs and the hand-written labels share a line — in one column. Each had a reasonable origin.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;unknown&lt;/code&gt; came from a button.&lt;/strong&gt; On a phone, every detail page shows a sticky bar at the bottom of the screen with the main call to action. The component that draws it takes the listing's ID as an optional prop, and falls back like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight tsx"&gt;&lt;code&gt;&lt;span class="nx"&gt;entityId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;entityId&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unknown&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of the four page types that used the bar passed an ID. Most of the site's visitors are on phones, so this was one of the most-clicked buttons on the site, and every click on it was logged as &lt;code&gt;unknown&lt;/code&gt;. All 157 that month came from it; 463 since June.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;List positions came from the main listings page.&lt;/strong&gt; The database gives every listing a UUID, but a shared TypeScript type declared &lt;code&gt;id: number&lt;/code&gt;. To satisfy it, the code that turned database rows into cards made one up:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;9000&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;index&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's a position in a list, not a property of a listing. Sort the list differently, filter it, add a listing, and the same number points somewhere else. It did: &lt;code&gt;9037&lt;/code&gt; was logged for two different listings, and so was &lt;code&gt;9010&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Slugs and labels came from blog posts&lt;/strong&gt;, whose links were written by hand — sometimes with the database ID, sometimes the URL slug, and sometimes a short label that matched nothing in the database at all.&lt;/p&gt;

&lt;p&gt;None of it was an error. Every insert succeeded, because the column was plain text and accepted all of it.&lt;/p&gt;

&lt;p&gt;It hadn't always been. It started as a UUID column, and months ago I loosened it to text so an events page could log IDs that weren't UUIDs. That was the one thing that could have refused &lt;code&gt;unknown&lt;/code&gt; and &lt;code&gt;9037&lt;/code&gt; — though it would have refused them by dropping the clicks, which is its own kind of wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the dashboard did with it
&lt;/h2&gt;

&lt;p&gt;Grouping by that column did two opposite things at once.&lt;/p&gt;

&lt;p&gt;It &lt;strong&gt;split&lt;/strong&gt; listings that should have been one row. The most-clicked real listing was spread across five different IDs: its UUID from its own page, a list position, a slug from one blog post, a label from another, and &lt;code&gt;unknown&lt;/code&gt;. Another was spread across seven. In the same 30 days, 76 listing names came back under 102 distinct IDs.&lt;/p&gt;

&lt;p&gt;And it &lt;strong&gt;merged&lt;/strong&gt; listings that should have been separate. Every &lt;code&gt;unknown&lt;/code&gt; click from every listing went into one bucket. The dashboard labelled each row with the name on the first click it happened to read, so 150 clicks from 29 listings wore one listing's name — and sat at the top of the table, because the sum of everyone's missing IDs will outscore any single listing.&lt;/p&gt;

&lt;p&gt;The panel next to it was worse, because it drew a conclusion. A "missed revenue" report flagged listings that got clicks but no affiliate clicks, as candidates for a booking link. It flagged six. Two of them had affiliate clicks all along, logged under a different ID for the same listing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An ID column isn't an identity. It's whatever each piece of code decided to send.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The column promises that one value means one thing, and nothing enforces that promise at the moment of writing. Each caller — a button, a list, a blog post someone wrote by hand — picks what to send, and the table stores it. When they disagree, the database has no way to know. The disagreement only surfaces where someone finally groups by the column and trusts the result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resolve, don't trust
&lt;/h2&gt;

&lt;p&gt;There were two fixes, and the order mattered.&lt;/p&gt;

&lt;p&gt;The first was to stop believing the label. The dashboard now loads the ID, slug and name of every listing and resolves each click to a real row — by ID first, then by slug, then by name, with names normalised so "The Example Inn" and "Example Inn", or "&amp;amp; Spa" and "and Spa", land on the same listing. Only then does it group: by the resolved row, not the logged value, and under the listing's name from the database rather than whatever the click carried.&lt;/p&gt;

&lt;p&gt;Before shipping it, I ran the same rule in SQL over every listing click ever recorded, to see what it would leave behind:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;matched by ID      1,219
matched by slug      156
matched by name      811
unmatched            140   (of 2,326)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The 140 are almost all listings deleted since — there's no row to resolve them to, but their logged names are consistent, so they still group correctly. The top row went from 150 anonymous clicks to the real most-clicked listing, at 110. The missed-revenue report went from six flags to four.&lt;/p&gt;

&lt;p&gt;The second fix was to stop sending bad IDs. The sticky bar now receives the listing's ID from every page that uses it. The listings page keeps each row's own ID instead of a list position, and the shared type now says &lt;code&gt;id: string&lt;/code&gt;, which is what it was all along. The hand-written blog list tracks the database ID instead of its label. Old rows keep their old values; the resolver handles those.&lt;/p&gt;

&lt;p&gt;Doing it in that order meant the dashboard was right about history the moment it shipped, not just about clicks from then on.&lt;/p&gt;

&lt;p&gt;A few days later the same lesson turned up one layer over. A small script checks every morning that each listing still carries a paid booking link. The site counted links from three booking domains as paid; the checker, written earlier, knew only two. The first listing to use the third domain would have been counted as unpaid on the day it started paying. Two pieces of code, one idea, two definitions — and each was right about its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  The build passed
&lt;/h2&gt;

&lt;p&gt;One detail from the fix, because it nearly shipped a quieter version of the same bug.&lt;/p&gt;

&lt;p&gt;The resolver builds its lookup with &lt;code&gt;new Map()&lt;/code&gt;. The build succeeded. The typechecker didn't:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="nx"&gt;TS2350&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Only&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="k"&gt;void&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;can&lt;/span&gt; &lt;span class="nx"&gt;be&lt;/span&gt; &lt;span class="nx"&gt;called&lt;/span&gt; &lt;span class="kd"&gt;with&lt;/span&gt; &lt;span class="nx"&gt;the&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;new&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="nx"&gt;keyword&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
&lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="nx"&gt;TS2558&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Expected&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;arguments&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;but&lt;/span&gt; &lt;span class="nx"&gt;got&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The dashboard file imports an icon library, and one of its icons is called &lt;code&gt;Map&lt;/code&gt;. In that file, &lt;code&gt;Map&lt;/code&gt; wasn't JavaScript's built-in — it was a React component, and calling &lt;code&gt;new&lt;/code&gt; on it throws. The bundler strips types without checking them, so it compiled without complaint.&lt;/p&gt;

&lt;p&gt;It wouldn't even have crashed. The lookup runs inside a data-fetching hook that catches errors, and I'd written the resolver to fall back to the logged names whenever the lookup wasn't available. So the dashboard would have loaded and looked normal, while quietly grouping by the very column this post is about — the fix switched off by its own safety net. &lt;code&gt;globalThis.Map&lt;/code&gt; put it right.&lt;/p&gt;

&lt;p&gt;I've written before that a &lt;a href="https://tedagentic.com/posts/fallback-chain-error-suppression" rel="noopener noreferrer"&gt;fallback chain is an error-suppression system&lt;/a&gt;, and that &lt;a href="https://tedagentic.com/posts/it-passed-because-it-never-looked" rel="noopener noreferrer"&gt;green is not verification&lt;/a&gt;. This was both, in four lines.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd check first
&lt;/h2&gt;

&lt;p&gt;If a dashboard ranks things by an ID that several parts of your code write, count what's actually in the column before trusting the ranking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt;
  &lt;span class="k"&gt;CASE&lt;/span&gt;
    &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;entity_id&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt; &lt;span class="s1"&gt;'^[0-9a-f]{8}-'&lt;/span&gt;          &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="s1"&gt;'uuid'&lt;/span&gt;
    &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;entity_id&lt;/span&gt; &lt;span class="o"&gt;~&lt;/span&gt; &lt;span class="s1"&gt;'^[0-9]+$'&lt;/span&gt;               &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="s1"&gt;'numeric'&lt;/span&gt;
    &lt;span class="k"&gt;WHEN&lt;/span&gt; &lt;span class="n"&gt;entity_id&lt;/span&gt; &lt;span class="k"&gt;IN&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'unknown'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;'null'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;''&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;THEN&lt;/span&gt; &lt;span class="s1"&gt;'placeholder'&lt;/span&gt;
    &lt;span class="k"&gt;ELSE&lt;/span&gt; &lt;span class="s1"&gt;'other'&lt;/span&gt;
  &lt;span class="k"&gt;END&lt;/span&gt; &lt;span class="k"&gt;AS&lt;/span&gt; &lt;span class="n"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;clicks&lt;/span&gt;
&lt;span class="k"&gt;GROUP&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;ORDER&lt;/span&gt; &lt;span class="k"&gt;BY&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="k"&gt;DESC&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If more than one kind comes back, the top of that table is a claim, not a count.&lt;/p&gt;

&lt;p&gt;And look hard at the fallbacks. &lt;code&gt;|| "unknown"&lt;/code&gt; looks like defensive code. It was a &lt;a href="https://tedagentic.com/posts/legal-data-audit-four-copies" rel="noopener noreferrer"&gt;default that set policy&lt;/a&gt; for 463 clicks, and the policy was: forget which listing this was.&lt;/p&gt;

&lt;p&gt;The top row wasn't my best listing. It was the one place every missing ID agreed to meet.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>database</category>
      <category>typescript</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Three AI Agents Corrected Each Other. None of Them Were Talking.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Tue, 15 Sep 2026 04:43:01 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/three-ai-agents-corrected-each-other-none-of-them-were-talking-5d2h</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/three-ai-agents-corrected-each-other-none-of-them-were-talking-5d2h</guid>
      <description>&lt;p&gt;CrewAI is an open-source Python framework for running several AI agents as a team. You give each agent a role and a goal, give the team a list of tasks, and the framework passes the work between them. Before building anything real with it, I wanted to answer one question: when people say the agents "talk to each other", what is actually happening?&lt;/p&gt;

&lt;p&gt;So I set up a small crew on the server in my house — a ThinkCentre that already runs my monitoring and automation — and gave it a practice topic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The crew
&lt;/h2&gt;

&lt;p&gt;Three agents, three tasks, run in order:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Research Analyst&lt;/strong&gt; — gathers the facts on a topic and marks anything it isn't sure of.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content Writer&lt;/strong&gt; — turns the research into a 300-word brief. Allowed to ask the Analyst questions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Editor&lt;/strong&gt; — checks the brief against the research, verifies the biggest claim with the Analyst, and sends weak sections back to the Writer. It finishes with a "Team notes" section listing every exchange.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Analyst and the Writer run on Qwen 3 235B through OpenRouter, a service that puts dozens of model hosts behind one API. The Editor runs on Claude Opus 4.8, direct from Anthropic. The cheap model does the volume; the expensive one makes the judgement calls.&lt;/p&gt;

&lt;p&gt;Setup is four commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv tool &lt;span class="nb"&gt;install &lt;/span&gt;crewai
&lt;span class="nv"&gt;CREWAI_DMN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true &lt;/span&gt;crewai create crew starter_crew &lt;span class="nt"&gt;--provider&lt;/span&gt; anthropic/claude-haiku-4-5
&lt;span class="nb"&gt;cd &lt;/span&gt;starter_crew &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; crewai &lt;span class="nb"&gt;install
&lt;/span&gt;crewai run
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That installed CrewAI 1.15.21, the version everything below was tested on. &lt;code&gt;CREWAI_DMN&lt;/code&gt; isn't in &lt;code&gt;crewai --help&lt;/code&gt;. I found it in the CLI's source: it switches off every interactive prompt, which is what you want when a script rather than a person is driving. It also means the scaffold never writes your &lt;code&gt;.env&lt;/code&gt; — the API keys are yours to add. The &lt;code&gt;--provider&lt;/code&gt; flag only chooses the starting model that gets written into each agent's file; I swapped those afterwards, to Qwen for the Analyst and the Writer and Claude Opus 4.8 for the Editor. The project itself is plain JSON: one file per agent, one &lt;code&gt;crew.jsonc&lt;/code&gt; for the tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The run
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F537ahjn09jh39r8tnjap.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F537ahjn09jh39r8tnjap.png" alt="The finished run: three tasks complete, the Editor's team notes, and three coworker tool calls in the activity log" width="800" height="387"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It took 118 seconds, 14,609 tokens in and 3,928 out.&lt;/p&gt;

&lt;p&gt;The Writer asked the Analyst a question before drafting. Then the Editor read the draft and stopped on a price: a capable home AI server for about $1,600. It asked the Analyst whether that held up. The answer came back that $1,600 is the graphics card alone — a complete machine runs $2,500–3,500 — and that a single 24 GB card can't run a 70-billion-parameter model at a usable speed anyway. The Editor sent the paragraph back to the Writer with the correction, took the revised version, and recorded both exchanges in its Team notes.&lt;/p&gt;

&lt;p&gt;Three agents, one caught error, one fix. It looked exactly like a team conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  It wasn't a conversation
&lt;/h2&gt;

&lt;p&gt;The activity log at the bottom of the run screen says what actually happened. Three entries:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;✓ ask_question_to_coworker     15.3s
✓ ask_question_to_coworker     20.4s
✓ delegate_work_to_coworker     8.9s
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those are tools. When an agent is allowed to delegate, CrewAI hands it two extra functions, the same way you'd hand an agent a web search or a file reader. Calling &lt;code&gt;ask_question_to_coworker&lt;/code&gt; with a coworker's name, a question and some context starts a fresh model call as that coworker, and whatever it returns comes back to the caller as the tool's result. The Editor never spoke to the Analyst. It called a function whose implementation happens to be another agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A conversation between agents isn't a conversation. It's a function call with a job title.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That changes how you build with it. The quality of the "discussion" is the quality of the arguments one model passes into a tool. The Editor only knew about the $1,600 because it was in the draft it was handed; the Analyst could only answer what the question contained. Nothing is shared between agents except what one call gives the next — and the handoffs between tasks work the same way, each task's output pasted into the next task's prompt.&lt;/p&gt;

&lt;p&gt;It also exposes the limit of what I watched. Nobody in that exchange looked anything up. These agents had no web access, so the Analyst "confirmed" the correction from the same kind of training data that produced the original number. The new figures — the $2,500–3,500 and the claim about the 24 GB card — are plausible. Neither is verified. The Analyst had no way to say "I don't know", and one of the operating rules I keep for this blog covers exactly that: &lt;a href="https://tedagentic.com/rules" rel="noopener noreferrer"&gt;a system with no way to say "I don't know" will answer anyway&lt;/a&gt;. A second model agreeing is a second opinion, not a second source. Giving agents tools can change that. Giving them coworkers doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five things that broke first
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The key was a placeholder.&lt;/strong&gt; My &lt;code&gt;.env&lt;/code&gt; held template text, not a key, and Anthropic answered &lt;code&gt;401 invalid x-api-key&lt;/code&gt;. One direct request to the API would have told me that before any framework was involved.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The provider wasn't installed.&lt;/strong&gt; The scaffold depends only on &lt;code&gt;crewai[tools]&lt;/code&gt;. Anthropic support is an optional extra, so the first run died with &lt;code&gt;Anthropic native provider not available&lt;/code&gt;. The fix is &lt;code&gt;uv add "crewai[anthropic]"&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory wanted a key I never gave it.&lt;/strong&gt; The scaffold switches crew memory on, and memory's default embedder is OpenAI's. With no OpenAI key, every memory lookup errored. &lt;code&gt;"memory": false&lt;/code&gt; until you configure a different embedder.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The newest Claude doesn't fit yet.&lt;/strong&gt; I wanted Opus 5 as the Editor, so I read CrewAI's Anthropic code first. Its thinking setting accepts only "enabled" or "disabled"; Opus 5 thinks by default and rejects the old "enabled with a budget" form; and CrewAI drops the thinking blocks when it rebuilds a tool-call turn. Agents calling each other through tools is exactly the path that would hit it. I didn't spend money finding out — Opus 4.8 costs the same per token and only thinks when asked.&lt;/p&gt;

&lt;p&gt;The fifth one is the one worth a section.&lt;/p&gt;

&lt;h2&gt;
  
  
  The limit I never set
&lt;/h2&gt;

&lt;p&gt;The first full three-agent run failed halfway. The Analyst finished its task. The Writer failed on every attempt — ten in a row — with this, passed back from the model host:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Requested token count exceeds the model's maximum context length of 131072 tokens.
You requested a total of 132086 tokens: 1014 tokens from the input messages
and 131072 tokens for the completion.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Writer's prompt was about a thousand tokens. The other 131,072 were the reply — under a reply limit I had never set.&lt;/p&gt;

&lt;p&gt;That was the whole problem. CrewAI only sends a reply limit (&lt;code&gt;max_tokens&lt;/code&gt;) when you configure one; I checked the code. With none sent, the reply was treated as allowed to fill the model's entire context window, my thousand tokens of input went on top, and the total was rejected for being over the window. The request didn't fail because it was big. It failed because it didn't say how big.&lt;/p&gt;

&lt;p&gt;I reproduced it without CrewAI: one request to the same host, a 12-token prompt, no &lt;code&gt;max_tokens&lt;/code&gt;. Same error — 12 plus 131,072 is 131,084, over by twelve. Why the Analyst survived, I can only guess: OpenRouter spreads requests across hosts, and my best explanation is that its requests landed on one that handled a missing limit differently. I didn't capture which host served them.&lt;/p&gt;

&lt;p&gt;Here is what I can show, and what I can't. The limit being enforced is 131,072 — the error comes back from that host's backend — while OpenRouter's listing for the same host advertises 262,144 tokens of context and 235,929 of output. What I can't show from outside is which layer wrote 131,072 into the reply field, OpenRouter or the host. That it matches the host's real window, not the advertised one, points at the host. Treat that as a hypothesis.&lt;/p&gt;

&lt;p&gt;The fix is one field. Every agent's model now carries &lt;code&gt;"max_tokens": 4096&lt;/code&gt; — 16,000 for the Claude editor. It's another case of a rule from the same list, &lt;a href="https://tedagentic.com/rules" rel="noopener noreferrer"&gt;defaults are policy&lt;/a&gt;: an unset value isn't an absence. Something downstream always chooses one, and here it chose the largest number it had.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An unset limit isn't no limit. It's the maximum.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs
&lt;/h2&gt;

&lt;p&gt;With only the Qwen models, a full three-agent run cost under half a cent. With Claude as Editor, the OpenRouter share stayed around $0.003. Anthropic doesn't report per-run cost to an API key, but even if every one of that run's 18,537 tokens had gone through Claude at Opus 4.8's list price — $5 per million input tokens, $25 per million output — it would come to about 17 cents.&lt;/p&gt;

&lt;h2&gt;
  
  
  I had a crew review this post
&lt;/h2&gt;

&lt;p&gt;Before publishing, I built a second crew and gave it this draft. A Cold Reader plays a stranger arriving from Google. A Reframe Critic looks for the one sentence where the idea flips. A Claude Editor checks every claim against the text and gives a verdict. Before any of them ran, plain code handled the mechanical checks — no private site names, a valid category, the image present, the links resolving — because no model should be trusted with those.&lt;/p&gt;

&lt;p&gt;The verdict was "fix first", and the two most useful catches were mine to own. I had stated two guesses as facts: why the Analyst's requests survived, and which layer filled in the 131,072. Both are labelled as guesses now. It also noticed that the setup command names one Claude model while the crew runs another, with nothing to reconcile them.&lt;/p&gt;

&lt;p&gt;One reviewer was simply wrong. The Cold Reader said a stranger "cannot access" the rules page these posts link to. It's a public page. It said so with exactly the same confidence as everything else it said.&lt;/p&gt;

&lt;p&gt;That review had the same shape as the run above. The reviewers didn't discuss the post. Each one was a function call handed the draft, and the Editor was handed their outputs. The good catches came from the Editor checking claims against the text in front of it. The false one came from an agent with no way to check anything, answering anyway.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you set one up
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Set &lt;code&gt;max_tokens&lt;/code&gt; on every agent's model. Don't let something downstream pick it.&lt;/li&gt;
&lt;li&gt;Turn memory off, or configure an embedder, before the first run.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;CREWAI_DMN=true crewai run&lt;/code&gt; in scripts — plain output that exits when it's done. Use plain &lt;code&gt;crewai run&lt;/code&gt; when you want to watch.&lt;/li&gt;
&lt;li&gt;Read the activity log, not the final report. The log is the actual conversation — and it's a list of function calls.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>crewai</category>
      <category>selfhosted</category>
    </item>
    <item>
      <title>The Error Was Real. The Page Was Fine.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Sat, 12 Sep 2026 03:10:00 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/the-error-was-real-the-page-was-fine-4g9o</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/the-error-was-real-the-page-was-fine-4g9o</guid>
      <description>&lt;p&gt;The report arrived from the hosting platform's AI monitoring, and it was admirably specific.&lt;/p&gt;

&lt;p&gt;A listings page on one of my sites, it said, &lt;em&gt;always&lt;/em&gt; falls back to a degraded list. The main database query fails every time, because it asks for a field that doesn't exist. The page silently drops to a simpler query, so the listings load without their city, state and country — which can break the grouping, the filters and the labels visitors see. It had errored 28 times in 24 hours.&lt;/p&gt;

&lt;p&gt;That is a good report. It names the symptom, the mechanism, the consequence and a count. I nearly went straight to the consequence and started looking for broken labels.&lt;/p&gt;

&lt;p&gt;The first half turned out to be exactly right. The second half never happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  The query had never worked
&lt;/h2&gt;

&lt;p&gt;The page asks the database for its listings along with a nested relationship: each listing's city, that city's state, and that state's country. In Supabase that is one request, written as a nested select:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;supabase&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;hotels&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`
    *,
    cities:city_id (
      name, slug,
      states:state_id (
        name, code,
        countries:country_id ( name, code )
      )
    )
  `&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I ran that exact request against the live database and got this back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;HTTP 400
{"code":"42703","message":"column countries_3.code does not exist"}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither the &lt;code&gt;states&lt;/code&gt; table nor the &lt;code&gt;countries&lt;/code&gt; table has a &lt;code&gt;code&lt;/code&gt; column. The database reports only the first missing column it trips over, so the error names countries, but both references were wrong. The request was malformed as written, so it could not succeed on any load, for any visitor, ever.&lt;/p&gt;

&lt;p&gt;Wrapped around it was a pattern you have probably written yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;tryWithRelationships&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;relationship query failed, falling back to simple query:&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;supabase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;hotels&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;select&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;*&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Try the rich version first. If it fails, degrade gracefully. It reads like defensive engineering, and it is. It is also, in this case, a description of a code path that had executed on every page load since the day it shipped. The "rich version" was not the primary path with a safety net beneath it. It was a request that failed, followed by the actual implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I checked what visitors saw
&lt;/h2&gt;

&lt;p&gt;The report's consequence — listings without their location — is a reasonable inference from that error. The failing request is the one that fetches location. If it never returns, location should be missing.&lt;/p&gt;

&lt;p&gt;Except the listings table stores its own copy. Each row carries plain &lt;code&gt;city&lt;/code&gt;, &lt;code&gt;state&lt;/code&gt; and &lt;code&gt;country&lt;/code&gt; columns, and those columns were populated on all 58 rows. The fallback's &lt;code&gt;select('*')&lt;/code&gt; returns them. And the code that formats each listing already reached for them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;city&lt;/span&gt;  &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;city&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cities&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Unknown&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;cities&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;states&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Unknown&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It uses the relationship when it exists and the row's own column when it doesn't. That had been written deliberately, with a comment explaining why: only 8 of the 58 rows had a &lt;code&gt;city_id&lt;/code&gt; set at all. So for 50 listings the relationship would have come back empty even on a day the query worked. Whoever wrote the formatter had already stopped trusting the relationship and routed around it.&lt;/p&gt;

&lt;p&gt;So every visitor got the fallback, and the fallback carried everything the page needed. The grouping worked. The filters worked. The labels were right. The error fired on every single load, and nobody who loaded the page ever saw it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An error rate counts what your code tried and failed. It does not count what anyone saw. When a fallback is doing its job, those two numbers stop having anything to do with each other.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The mirror image of the usual problem
&lt;/h2&gt;

&lt;p&gt;I have written before about a fallback chain that did the opposite. Four layers of image lookup, where a dead link in layer three was simply covered by layer four, and forty-nine broken URLs sat behind a default photo without a single error being raised. That was silence hiding damage.&lt;/p&gt;

&lt;p&gt;This was noise hiding the absence of damage. Same architecture, the same defensible habit of degrading instead of crashing, and the signal failed in the other direction. A fallback can make a real failure invisible. It can just as easily make a harmless failure look catastrophic, because the thing that fails is loud, and the thing that catches it is quiet.&lt;/p&gt;

&lt;p&gt;That's what the report got wrong, and it was a good-faith mistake. It read the error, which is about a query, and inferred the impact, which is about a page. The error message was completely accurate about the query. It simply contained no information about the page, because the page's outcome was decided one step later, by code the error never passed through.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually cost
&lt;/h2&gt;

&lt;p&gt;Not nothing, which is worth saying plainly, since "the page was fine" is not the same as "there was no bug."&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A failed request on every load.&lt;/strong&gt; 28 in a day is 28 wasted round-trips and 28 console errors, each one burying whatever real error might arrive next to it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A code path that misrepresented itself.&lt;/strong&gt; Anyone reading the component would believe the page had a richer mode it used when it could. It never had one. Code that describes behaviour it doesn't have is a trap for the next person who reads it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One field genuinely degraded.&lt;/strong&gt; Country had no column fallback. It read &lt;code&gt;row.cities?.states?.countries?.name || 'USA'&lt;/code&gt;, so every listing was labelled with a hardcoded default.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fix for the error was deleting two field names from the select. The request now returns 200 and resolves the relationship for the 8 rows that have one.&lt;/p&gt;

&lt;p&gt;The country field was the part that didn't go the way I expected. With the query fixed, I checked what came back for those 8 rows, expecting real country names at last. Every one of them was null. The &lt;code&gt;states&lt;/code&gt; table has a &lt;code&gt;country_id&lt;/code&gt; column, and it is empty on every row. So the country relationship cannot resolve for anything, and the label still falls back to &lt;code&gt;'USA'&lt;/code&gt;. It happens to be correct today, because every listing is in the US. It will quietly be wrong for the first one that isn't.&lt;/p&gt;

&lt;p&gt;That is the richer mode, working exactly as written, with nothing underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The useful version of the report
&lt;/h2&gt;

&lt;p&gt;I don't think the lesson is "distrust automated monitoring." The monitor found a real defect that had been live for as long as the page existed, and I would never have looked otherwise.&lt;/p&gt;

&lt;p&gt;The lesson is that an error with a 100% failure rate and zero visible impact is not a contradiction to be explained away. It is a finding in its own right, and a sharper one than the report made. &lt;strong&gt;If something fails every time and nobody notices, then nobody depends on it.&lt;/strong&gt; That path is dead code wearing the costume of the main path.&lt;/p&gt;

&lt;p&gt;So when an alert tells you something fails on every request, ask the second question before reaching for the consequence: &lt;em&gt;what did the user get instead?&lt;/em&gt; If the answer is "exactly what they needed," then the error isn't telling you the page is broken. It's telling you the part you thought was the system has never actually run — and that the part you thought was the backup is the system.&lt;/p&gt;

&lt;p&gt;The number was true. The story attached to it was a guess.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>debugging</category>
      <category>database</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>The API Said 44 Updated. Twenty Didn't Change.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Sat, 12 Sep 2026 03:08:39 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/the-api-said-44-updated-twenty-didnt-change-2eom</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/the-api-said-44-updated-twenty-didnt-change-2eom</guid>
      <description>&lt;p&gt;I keep a copy of every post from this blog on dev.to, the developer blogging platform. Each copy points back here with a canonical link, so search engines treat this site as the original. Last month I rewrote the descriptions on all of them — the short summary that shows on share cards and in dev.to's own search.&lt;/p&gt;

&lt;p&gt;Forty-five articles. A job for a script, not a morning of clicking.&lt;/p&gt;

&lt;p&gt;I did it the careful way. I updated one article first, and before and after I fingerprinted everything the update could plausibly disturb: the length of the body, the tags, the canonical URL, the title, the published state. Only the description had changed. Everything else was byte-identical. Then I ran the other forty-four.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;updated 44, failed 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every request came back &lt;code&gt;200 OK&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Out of habit rather than suspicion, I ran the checker again — the one that compares what's live against what I intended. Twenty articles still had their old description.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not a cache
&lt;/h2&gt;

&lt;p&gt;My first thought was lag. The endpoint that lists all your articles at once is exactly the kind of thing that serves a slightly stale snapshot. So I skipped it and fetched each of the twenty individually, from the endpoint that returns a single article fresh.&lt;/p&gt;

&lt;p&gt;Old descriptions, all twenty. The writes hadn't been delayed. They hadn't happened.&lt;/p&gt;

&lt;p&gt;So forty-four requests had gone out, all forty-four had been answered with success, and twenty of them had changed nothing. Nothing in the responses separated the twenty from the twenty-four. Same status, same shape, same body echoing back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The twenty had something in common
&lt;/h2&gt;

&lt;p&gt;Laid side by side, the difference was in the article bodies. Every one of the twenty began like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Zero&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Measurement"&lt;/span&gt;
&lt;span class="na"&gt;published&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;old&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;summary,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;still&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;sitting&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;here."&lt;/span&gt;
&lt;span class="na"&gt;tags&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;debugging, monitoring&lt;/span&gt;
&lt;span class="na"&gt;canonical_url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://tedagentic.com/posts/zero-is-not-a-measurement&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

The article itself starts here...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That block between the two &lt;code&gt;---&lt;/code&gt; lines is front matter: metadata written into the top of a markdown file, the same convention static site generators use. dev.to supports it. You can publish an article by sending the whole thing, metadata block included, and dev.to reads its settings from there.&lt;/p&gt;

&lt;p&gt;The twenty-five articles that updated correctly had no such block. Their bodies were just the article.&lt;/p&gt;

&lt;p&gt;Same author, same months, interleaved through the calendar. I had two kinds of article and had never known it, because from the outside — on the page, in the list, in the API — they looked identical.&lt;/p&gt;

&lt;p&gt;And the one I had tested first, so carefully, with all those fingerprints? It was one of the twenty-five. My probe had been perfectly representative of the articles where the method works, and had told me nothing at all about the ones where it doesn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  I was writing to the output of a function
&lt;/h2&gt;

&lt;p&gt;Here is what happened on each of those twenty saves.&lt;/p&gt;

&lt;p&gt;The request arrived with a new &lt;code&gt;description&lt;/code&gt;. It was well-formed, the field was a legitimate part of the article, and dev.to accepted it. Then, as part of the same save, dev.to read the front matter out of the body, found a &lt;code&gt;description:&lt;/code&gt; line there, and set the article's description from that. My value lasted exactly as long as it took to be overwritten by the block it was supposed to replace.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For those twenty articles, the description wasn't a field I could set. It was a computed value — read out of the body every time the article saved. I had been writing to the output of a function.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is why the response said 200. The request was valid and it was processed. The write landed and was undone inside a single operation, and the status code described the first half.&lt;/p&gt;

&lt;p&gt;Four of the twenty were stranger still. Their front matter had no &lt;code&gt;description:&lt;/code&gt; line at all. For those, dev.to had been generating the description from the opening words of the article. They had never had a description as data in the first place — only a summary produced on demand, which no update to the field was ever going to reach.&lt;/p&gt;

&lt;h2&gt;
  
  
  A copy can be fixed. A derived value can't.
&lt;/h2&gt;

&lt;p&gt;I've written before about copies of data drifting apart: a fact corrected in the database that never reached the page, prose that went on contradicting a dataset after the dataset was fixed. Those are real problems, but they share a comforting property. A copy exists. It is a stored thing that can be found, compared and corrected.&lt;/p&gt;

&lt;p&gt;A derived value has no independent existence. There is nothing in the field to fix, because the field is not where the value lives. Overwrite it and you have written to a view, and a view is rebuilt from its source whenever the system feels like it. The only place to change it is upstream.&lt;/p&gt;

&lt;p&gt;That is a different and worse failure than drift, because every check at the level of the field passes. You wrote it. The system confirmed it. For a moment, it was even true.&lt;/p&gt;

&lt;h2&gt;
  
  
  Changing the source
&lt;/h2&gt;

&lt;p&gt;So the fix was to stop writing to the field and write to the source instead: edit the &lt;code&gt;description:&lt;/code&gt; line inside each article's front matter, then send the whole body back. For the four with no such line, insert one.&lt;/p&gt;

&lt;p&gt;That meant pushing the entire body of twenty live articles, not one field. A mistake there could damage far more than a summary, so every edit went through two gates. The new body had to differ from the old by exactly one line — one changed, or for the four, one added — or the edit was rejected before it was sent. And after each save I fetched the article fresh and confirmed two things: the description was now the new one, and the tags were exactly as they had been.&lt;/p&gt;

&lt;p&gt;All twenty passed. The checker now reads forty-five of forty-five.&lt;/p&gt;

&lt;p&gt;(One aside, for anyone scripting against the same API: dev.to answers Python's default &lt;code&gt;urllib&lt;/code&gt; user agent with &lt;code&gt;403 Forbidden Bots&lt;/code&gt;. Send an ordinary user agent string and it works.)&lt;/p&gt;

&lt;h2&gt;
  
  
  You can't tell from the schema
&lt;/h2&gt;

&lt;p&gt;The part that stays with me is that nothing warned me. The description was listed in the API as a property of the article. It accepted writes. It returned 200. There was no flag marking it as derived for some articles and stored for others, and I suspect there often isn't, anywhere. Whether a field is data or a projection of other data is usually an implementation detail, invisible until you write to it and watch it snap back.&lt;/p&gt;

&lt;p&gt;So there are only two defences I trust now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read back from the source, not from the response.&lt;/strong&gt; The response tells you the request was accepted. A fresh fetch of the thing, from the endpoint that doesn't cache, tells you whether your value survived. Those are different questions, and the success of the first says nothing about the second.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A probe only proves its own kind.&lt;/strong&gt; One careful test validated my method for exactly the population that test article came from. If there are two kinds of thing hiding behind one interface, a single sample will happily vouch for the easy one. You find out there were two kinds when the second one fails — which is a strong argument for checking every result afterwards, rather than trusting a representative one beforehand.&lt;/p&gt;

&lt;p&gt;The API never lied to me. It said it had accepted my request, and it had. I was the one who heard "it's done."&lt;/p&gt;

</description>
      <category>api</category>
      <category>webdev</category>
      <category>debugging</category>
      <category>programming</category>
    </item>
    <item>
      <title>Clicks Rose 73%. I Shut the Site Down.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Tue, 25 Aug 2026 13:21:04 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/clicks-rose-73-i-shut-the-site-down-5fnb</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/clicks-rose-73-i-shut-the-site-down-5fnb</guid>
      <description>&lt;p&gt;In June I fixed a real problem on a small editorial site I run.&lt;/p&gt;

&lt;p&gt;The site was a client-side app — the kind where the server sends a nearly empty HTML shell and the browser assembles the actual page afterward. That works fine for people. It works badly for search crawlers, because the description a search engine shows underneath your link in the results comes from the HTML the server sent. Every page on this site was sending the same generic one. Fifty-odd pages, one description, all identical.&lt;/p&gt;

&lt;p&gt;So I built a step into the deploy that renders each page's real title and description into the file the server hands out. Crawlers started seeing per-page text instead of one repeated boilerplate line.&lt;/p&gt;

&lt;p&gt;It worked. I can show you that it worked.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers said yes
&lt;/h2&gt;

&lt;p&gt;Two things get measured here. &lt;strong&gt;Impressions&lt;/strong&gt; are how many times one of your pages appeared in someone's search results. &lt;strong&gt;Clicks&lt;/strong&gt; are how many times someone actually clicked through. Click-through rate is the second divided by the first — of the people who saw you, what fraction picked you.&lt;/p&gt;

&lt;p&gt;The fix was aimed squarely at click-through rate. A better description should win more of the people already seeing you, without changing where you rank.&lt;/p&gt;

&lt;p&gt;Comparing the 28 days before the fix to the 28 days ending this week:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Clicks&lt;/th&gt;
&lt;th&gt;CTR&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Before the fix&lt;/td&gt;
&lt;td&gt;11&lt;/td&gt;
&lt;td&gt;0.47%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Now&lt;/td&gt;
&lt;td&gt;19&lt;/td&gt;
&lt;td&gt;0.56%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Click-through rate up about 19%. Clicks up 73%.&lt;/p&gt;

&lt;p&gt;That is the shape you want. The intervention targeted a specific mechanism, the mechanism moved, and it moved in the predicted direction. If I'd put that on a dashboard and walked away, I'd have called it a win and gone looking for the next thing to fix.&lt;/p&gt;

&lt;p&gt;Two months later I removed the site from all of my monitoring and retired it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The number I wasn't watching
&lt;/h2&gt;

&lt;p&gt;Here's the third column, the one that isn't on anybody's dashboard.&lt;/p&gt;

&lt;p&gt;Not &lt;em&gt;how many clicks did the site get&lt;/em&gt; — &lt;strong&gt;how many distinct pages earned anything at all.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Pages with ≥1 impression&lt;/th&gt;
&lt;th&gt;Pages with ≥1 click&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Before the fix&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A month later&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Now&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Ten. Out of about fifty.&lt;/p&gt;

&lt;p&gt;And the same five pages have earned every single click for three consecutive months. Not five pages of a rotating cast — the identical five. The top five pages account for 91% of all impressions the site receives.&lt;/p&gt;

&lt;p&gt;So both readings are true at once. The clicks went up 73%. The surface producing those clicks contracted by nearly a quarter. I got better at converting a shrinking audience.&lt;/p&gt;

&lt;h2&gt;
  
  
  How an aggregate hides a shrinking surface
&lt;/h2&gt;

&lt;p&gt;This isn't a paradox, it's arithmetic, and it's worth being precise about because the same trap sits under a lot of small-system metrics.&lt;/p&gt;

&lt;p&gt;Clicks is a &lt;strong&gt;sum over pages&lt;/strong&gt;. Any sum can rise while its terms are disappearing, as long as the survivors grow faster than the losses. On a big site this washes out — thousands of pages, and one page's fate is noise. On a site where five pages produce 91% of everything, the total &lt;em&gt;is&lt;/em&gt; those five pages. It stops being a measure of the site and becomes a measure of a handful of URLs.&lt;/p&gt;

&lt;p&gt;The distinction that matters: clicks measured &lt;strong&gt;how well my best pages were doing&lt;/strong&gt;. Coverage measured &lt;strong&gt;whether the site could still get new pages in front of anyone&lt;/strong&gt;. Those are different questions, and only the second one tells you whether there's a business here.&lt;/p&gt;

&lt;p&gt;My fix improved the first. It was never capable of improving the second — better descriptions help pages that already appear in results. It does nothing for pages that never appear at all. When I audited this site in June, nineteen of its pages had been crawled and passed over, and eighteen had never been crawled at all. Better descriptions were never going to reach either group. Two months on, the count of pages earning anything is lower than when I started, so whatever happened to those thirty-seven, it wasn't a breakthrough.&lt;/p&gt;

&lt;p&gt;I shipped a change that could only help the pages that needed help least. And the metric I'd chosen to judge it by was the one metric that couldn't tell me that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The detail that ended the argument
&lt;/h2&gt;

&lt;p&gt;One more number, because it's the one that made the decision obvious rather than merely defensible.&lt;/p&gt;

&lt;p&gt;The site's single best query is a person's name — a historical figure the site has a long article about. Over the last month that query put the page in front of people 454 times, at an average position of &lt;strong&gt;8.6&lt;/strong&gt;. That's the bottom of the first page of results. Respectable. Hard-won.&lt;/p&gt;

&lt;p&gt;It earned &lt;strong&gt;one&lt;/strong&gt; click. A click-through rate of 0.22%.&lt;/p&gt;

&lt;p&gt;At position eight or nine you'd expect somewhere between 1.5% and 3%. Getting a fifth of the floor isn't a snippet problem, and it isn't something a better description fixes. It's what happens when you rank underneath a Wikipedia entry and one of those boxed summaries search engines assemble on the right-hand side. The question gets answered on the results page. The user's need is met before your link is a candidate. You are technically visible and functionally absent.&lt;/p&gt;

&lt;p&gt;Which reframes the whole asset. The queries this site ranks for are queries that get answered without anyone leaving the search results. That is not a fixable snippet. That is the market telling you what your content is worth to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I got wrong
&lt;/h2&gt;

&lt;p&gt;I laid all of this out and then offered to set up a checkpoint: re-run the exact comparison in a month, deliver it automatically, decide then with fresh numbers.&lt;/p&gt;

&lt;p&gt;The answer I got back was that eight months of no movement was already the measurement.&lt;/p&gt;

&lt;p&gt;That's correct, and I should have said it myself instead of proposing more data collection. The trend had twelve months in it. Best month in the entire year: 16 clicks — this month, the one that supposedly shows the fix working. Second best: 15, back in March, three months before the fix existed. The whole twelve-month range is 4 to 16. Impressions had &lt;em&gt;peaked&lt;/em&gt; in February and were down 49% from that peak. Every one of those facts was already in front of me when I offered the checkpoint.&lt;/p&gt;

&lt;p&gt;There's a failure mode here that looks like rigor and isn't. When a trend is genuinely ambiguous, more measurement is the right call. When a trend has been flat for eight months across every framing you can apply to it, proposing another month of measurement isn't rigor — it's deferring a decision while appearing thorough. The tell is whether you can describe, in advance, a result that would change your mind. I couldn't. I'd already concluded the only remaining lever was a multi-month authority build that wasn't going to happen. Wanting one more datapoint was me not wanting to say so.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd watch instead
&lt;/h2&gt;

&lt;p&gt;For anything small enough that a handful of pages dominates the totals, the aggregate is close to useless as a health signal. It tells you how your winners are doing, and your winners are usually fine — that's what makes them winners.&lt;/p&gt;

&lt;p&gt;The two numbers that actually carried information:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distinct pages earning any impression, tracked over time.&lt;/strong&gt; Rising means the system still has reach. Falling means it's consolidating toward a few survivors, which is what dying looks like from the inside before it looks like anything on a revenue chart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Whether anything outside the incumbent set has broken through.&lt;/strong&gt; Not "did clicks grow" but "did a page that wasn't earning before start earning." On this site the answer was no, for three straight months, while clicks climbed 73%.&lt;/p&gt;

&lt;p&gt;If both of those are flat or falling, a rising total isn't recovery. It's a smaller thing being measured more favourably. I had that data the whole time, in the same account, one query away. I just wasn't asking for it, because the number I was asking for kept coming back green.&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>seo</category>
      <category>webdev</category>
      <category>devops</category>
    </item>
    <item>
      <title>Every Page Returned 200. There Was No Database.</title>
      <dc:creator>Ted</dc:creator>
      <pubDate>Mon, 24 Aug 2026 15:23:46 +0000</pubDate>
      <link>https://dev.to/henry_dan_81513dd35a2f540/every-page-returned-200-there-was-no-database-jb3</link>
      <guid>https://dev.to/henry_dan_81513dd35a2f540/every-page-returned-200-there-was-no-database-jb3</guid>
      <description>&lt;p&gt;My travel site was completely non-functional for about four hours today and never once went down.&lt;/p&gt;

&lt;p&gt;That is not wordplay. Every route returned HTTP 200 the entire time. Time to first byte was normal. The HTML came back full of content — property names, headings, structured data, the lot. An uptime monitor pointed at any page on the site would have recorded a flawless afternoon.&lt;/p&gt;

&lt;p&gt;Nobody could book anything, because there was nothing to book.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually broke
&lt;/h2&gt;

&lt;p&gt;The site's backend runs on a managed platform — a hosted wrapper around Postgres that bundles the database, authentication, file storage and serverless functions into one project. The workspace was on the free tier. Its allowance ran out and the project was paused.&lt;/p&gt;

&lt;p&gt;There was no option to settle a balance and carry on. The dashboard offered exactly one action: upgrade to a $25/month plan. That is worth separating from the usual story about an operator forgetting to pay a bill, because it is a different failure — the free tier had been running production for nine months, and the notice that it would stop was the stopping itself.&lt;/p&gt;

&lt;p&gt;The first thing I checked was DNS, and that told me how complete this was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;dig +short @1.1.1.1 &amp;lt;project&amp;gt;.supabase.co
&lt;span class="nv"&gt;$ &lt;/span&gt;dig +short @8.8.8.8 &amp;lt;project&amp;gt;.supabase.co
&lt;span class="nv"&gt;$ &lt;/span&gt;dig +short @1.1.1.1 supabase.co
76.76.21.21
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two independent resolvers, no answer. The parent domain resolved fine, so this was not a DNS outage — the subdomain had been withdrawn. NXDOMAIN, not a pause page.&lt;/p&gt;

&lt;p&gt;One exhausted allowance took four subsystems with it at once: the database, every login, every uploaded image, and the edge function that records outbound clicks. They felt like separate concerns right up until they shared a fate.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a visitor saw
&lt;/h2&gt;

&lt;p&gt;Here is the part that has stayed with me.&lt;/p&gt;

&lt;p&gt;The site is a React SPA with a prerender step: at build time, every route is rendered to static HTML with its data baked in. That HTML is what a crawler gets, and it is why &lt;code&gt;curl&lt;/code&gt; looked reassuring:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://&amp;lt;site&amp;gt;/hotels | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s2"&gt;"Patterson Inn"&lt;/span&gt;
Patterson Inn
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The content was right there. So for a few minutes I believed the damage was limited to logged-in features.&lt;/p&gt;

&lt;p&gt;Then I rendered the page in a real browser instead of reading its source. React hydrated, fired its query at a host that no longer existed, got nothing, and re-rendered the component with an empty array.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Found 0 verified stays.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every filter read zero. The listings area showed skeleton placeholders. The page was beautifully styled, fully responsive, instantly loaded, and contained nothing whatsoever.&lt;/p&gt;

&lt;p&gt;I checked the one thing that actually pays for the site:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://&amp;lt;site&amp;gt;/hotels | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"affiliate"&lt;/span&gt;
0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not a single outbound booking link exists in the static HTML. Those URLs come from the database at render time. So there was no accidental fallback, no degraded-but-working state. Just a confident, empty shop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why nothing alerted me
&lt;/h2&gt;

&lt;p&gt;I found out because I happened to open the platform dashboard for an unrelated reason.&lt;/p&gt;

&lt;p&gt;Think about what a conventional check would have needed to notice. HTTP status? 200. Response time? Fine. Does the HTML contain expected text? Yes — the prerendered content was intact. Is the page non-empty? Very. Does the site load? Beautifully.&lt;/p&gt;

&lt;p&gt;Every signal a normal monitor collects was healthy, because every one of them measures whether the server &lt;em&gt;answered&lt;/em&gt;, not whether the answer &lt;em&gt;meant&lt;/em&gt; anything. The failure lived entirely in the gap between those two questions.&lt;/p&gt;

&lt;p&gt;That gap is not exotic. Any site that renders static shells and fills them from an API has it. The more thorough your prerendering, the wider it gets — because the static layer keeps looking healthy long after the dynamic layer has died.&lt;/p&gt;

&lt;h2&gt;
  
  
  The crawler got a better experience than the customer
&lt;/h2&gt;

&lt;p&gt;There is a genuine consolation, and it is a strange one.&lt;/p&gt;

&lt;p&gt;Googlebot does not execute JavaScript on the schedule a visitor's browser does. It was served the prerendered HTML — 623 routes of intact content, correct structured data, working internal links. Throughout the outage, the search engine's view of the site was perfectly healthy.&lt;/p&gt;

&lt;p&gt;So the prerendering protected the index and abandoned the customer. For a nine-month-old domain still being assessed, that is the more expensive of the two to lose, and I would not have chosen differently. But it is worth being clear about what was protected and what was not, because "the prerendering saved us" is only half true and the other half was every booking that afternoon.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resolving is not serving
&lt;/h2&gt;

&lt;p&gt;One more thing, because it cost me the backup.&lt;/p&gt;

&lt;p&gt;Before paying, I set a script running that polled DNS and would export everything the moment the host came back — the idea being that if the alert came at 3am, the data would already be on my disk.&lt;/p&gt;

&lt;p&gt;It fired at 15:42. It captured nothing.&lt;/p&gt;

&lt;p&gt;DNS came back roughly forty minutes before the service did. My readiness check probed the API root, got a &lt;code&gt;401&lt;/code&gt;, and treated that as life — a 401 is a valid HTTP response from a live server, after all. It then requested every table and received &lt;code&gt;521&lt;/code&gt; for each one, which is the CDN saying it cannot reach the origin. The script dutifully reported sixteen tables as "not exported" and exited pleased with itself.&lt;/p&gt;

&lt;p&gt;The fix is to probe something that only works when the thing actually works:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# before: any response means alive&lt;/span&gt;
&lt;span class="nv"&gt;code&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code}'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOST&lt;/span&gt;&lt;span class="s2"&gt;/rest/v1/"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$code&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="s2"&gt;"000"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;break&lt;/span&gt;

&lt;span class="c"&gt;# after: a real table, and only 200 counts&lt;/span&gt;
&lt;span class="nv"&gt;code&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOST&lt;/span&gt;&lt;span class="s2"&gt;/rest/v1/hotels?select=name&amp;amp;limit=1"&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"apikey: &lt;/span&gt;&lt;span class="nv"&gt;$KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$code&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;"200"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same mistake as the monitoring, one layer down. I asked "did something answer?" when the question was "did the thing I need work?"&lt;/p&gt;

&lt;p&gt;The re-run pulled 91MB: 220 countries, 143 dispensaries, 58 properties, and 563 storage objects, none of which failed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;Uptime is a measurement of whether your server responds. For a static site those are the same thing, which is why the convention exists and why it goes unexamined. For a data-driven site they are different questions, and the difference is the entire business.&lt;/p&gt;

&lt;p&gt;A 500 would have been better. A 500 pages an on-call rota, trips a monitor, sends an email. What I had instead was a site that passed every automated check while converting nothing, and would have kept doing it for as long as I did not personally look.&lt;/p&gt;

&lt;p&gt;So the health check I actually needed was never about the server. It was about the data:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Fetch the listings endpoint. Assert the count is greater than zero.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two lines. It would have caught this within a minute, and it is the only check in this entire incident that would have.&lt;/p&gt;

&lt;p&gt;The corollary is a build I have not made yet: the prerender already pulls every listing at build time, so the client could fall back to that baked-in data when a live fetch fails, instead of rendering zeros. Stale listings still sell. Zero listings sell nothing, and look immaculate doing it.&lt;/p&gt;

</description>
      <category>monitoring</category>
      <category>devops</category>
      <category>webdev</category>
      <category>database</category>
    </item>
  </channel>
</rss>
