<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Lisandro Reinoso</title>
    <description>The latest articles on DEV Community by Lisandro Reinoso (@lisandro_reinoso_d12ac7b9).</description>
    <link>https://dev.to/lisandro_reinoso_d12ac7b9</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4121361%2F75a7fbf5-cc44-4d69-ae7a-be25e934dfe6.png</url>
      <title>DEV Community: Lisandro Reinoso</title>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lisandro_reinoso_d12ac7b9"/>
    <language>en</language>
    <item>
      <title>Questions for a chatbot</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Thu, 01 Oct 2026 15:25:26 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/questions-for-a-chatbot-2ic0</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/questions-for-a-chatbot-2ic0</guid>
      <description>&lt;p&gt;Today I have a file open with an empty table. At the top are the numbers from two weeks ago: out of 68 answers, one passed. At the bottom, the hypotheses I wrote so I wouldn't cheat myself when measuring again. The table in the middle, the one that would say whether the chatbot got better, has been empty for 16 days.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I ask it
&lt;/h2&gt;

&lt;p&gt;The chatbot answers questions about pasture growth rates with a chart and a text that explains it. Its user is a rural worker who doesn't code. To find out whether it answers well, I built a golden set: 68 questions, none of them from real users. Twenty came from a first round of examples and 48 from a second.&lt;/p&gt;

&lt;p&gt;A golden set is a fixed list of questions with what you expect from each answer, which you rerun after every change to see whether things got better. If whoever grades only says "pass" or "fail", you know how much fails. My auditor, a separate agent, returns something besides the verdict: &lt;code&gt;faltantes&lt;/code&gt; (what each answer was missing), with a type from a closed vocabulary (data, domain knowledge, tool, prompt) and a key so they can be summed across questions. That way you know what to fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  1 of 68
&lt;/h2&gt;

&lt;p&gt;In the September 14 baseline, 1 passed; 31 passed with reservations, 35 were rejected and 1 errored out.&lt;/p&gt;

&lt;p&gt;That number wasn't the useful part. The useful part was 205 gaps, which grouped into 102 keys and mapped to 22 tickets: 16 new and 6 that were already in the backlog. Today 11 are done, 5 cancelled and 6 open. The three most frequent gaps: not knowing the dataset's reference date (24), having no rainfall column (19) and not being able to filter by field and paddock at the same time (16).&lt;/p&gt;

&lt;p&gt;Rainfall didn't get fixed. There is no weather data, and I decided the chatbot should say so instead of inventing a proxy. A gap that gets resolved by saying "I don't know".&lt;/p&gt;

&lt;h2&gt;
  
  
  The pending measurement
&lt;/h2&gt;

&lt;p&gt;Before measuring again I wrote hypotheses with a threshold: rejected answers, from 35 down to 20 or fewer; the reference date, from 24 down to 5 or fewer. And a counter-hypothesis: the change that lets the model call no tool at all could make questions that were chartable worse.&lt;/p&gt;

&lt;p&gt;The second run generated the answers, but the audit stage, about 28 dollars, was cut off: the API credit ran out. It has been blocked since September 15.&lt;/p&gt;

&lt;p&gt;I know what I fixed. I don't know if it got better.&lt;/p&gt;

&lt;p&gt;When you evaluate something you built, does it hand you a score or a list of what's missing? And do you write down beforehand what you expect to change?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>testing</category>
      <category>chatbot</category>
    </item>
    <item>
      <title>A reader saw what my fix didn't do</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Wed, 30 Sep 2026 21:17:05 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/a-reader-saw-what-my-fix-didnt-do-53bp</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/a-reader-saw-what-my-fix-didnt-do-53bp</guid>
      <description>&lt;p&gt;The previous note ended with a question, and today someone answered. They didn't answer the question: they corrected the fix I had presented as the end of the story. Before replying I opened the scripts to see if they were right. They were.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the reader said
&lt;/h2&gt;

&lt;p&gt;A reminder: two agents wrote "I'm going to wait" and stopped, the harness cut them off after ten minutes, and the fix was a single line.&lt;/p&gt;

&lt;p&gt;For the reader, that sentence was the useful signal, not the bug: it describes an intention, not a durable wait. They proposed that every deferred task leave a record the machine owns, with an id, an explicit deadline, and the next action, and that someone reconcile it separately. And that "success" should mean that mechanism exists, not that the agent promised to use it.&lt;/p&gt;

&lt;p&gt;A handle, without the jargon: when one process hands work to another and moves on, the only thing that makes that wait durable is a record the machine keeps, with an identifier for the task, an explicit deadline, and what to do when it expires. Someone can check against it later to see what happened. A sentence from the agent is not that.&lt;/p&gt;

&lt;h2&gt;
  
  
  The weak version
&lt;/h2&gt;

&lt;p&gt;My fix exports &lt;code&gt;CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS&lt;/code&gt; as &lt;code&gt;0&lt;/code&gt;: no ceiling, wait forever. It removes the deadline instead of making it explicit. It makes the promise true by making it infinite, leaving nothing to reconcile.&lt;/p&gt;

&lt;p&gt;The only reconciliation I have lives at the ticket level and arrives after the fact: &lt;code&gt;run-ticket.sh&lt;/code&gt; reads the final JSON; if &lt;code&gt;is_error&lt;/code&gt; is true, the ticket ends up &lt;code&gt;blocked&lt;/code&gt;; if there's no verdict, it leaves the record untouched. The two turns from that early morning, which together cost USD 2.79, closed with &lt;code&gt;success&lt;/code&gt;. Nothing marked them as failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed
&lt;/h2&gt;

&lt;p&gt;Three days earlier a first comment had arrived, on a different note, and I decided not to capture it. This one I did: the same day it became a ticket in the template every new project is born from, with what the reader named (an explicit deadline per run, a block with the reason "deadline", closing with the handle and not with a sentence) and a pending decision they also saw: what to do when the deadline expires with the work half done.&lt;/p&gt;

&lt;p&gt;That same early morning the ceiling had cut off a third project. Today the line is in one of thirteen scripts, and in this project the ticket that adds it has been pending for seven days. The weak fix didn't even propagate; the strong one is still a ticket.&lt;/p&gt;

&lt;p&gt;I published that the problem was a line that doesn't travel, and a reader showed me the line wasn't the solution either. When an agent of yours delegates and says "I'll wait," what does the machine own of that wait, and who reconciles it if it expires?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>"The agent said it would wait, and didn't"</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Thu, 17 Sep 2026 11:35:33 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/the-agent-said-it-would-wait-and-didnt-1keb</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/the-agent-said-it-would-wait-and-didnt-1keb</guid>
      <description>&lt;p&gt;This morning, before launching the forks that refined this note's tickets, I exported an environment variable by hand from outside the script that starts the loop. I set it that way every time I run this project. Five nights ago that variable didn't exist anywhere, and not knowing that cost USD 2.79 and zero finished tickets.&lt;/p&gt;

&lt;h2&gt;
  
  
  I'm going to wait
&lt;/h2&gt;

&lt;p&gt;In the early hours of September 12th, in the chatbot project, the loop ran in parallel for the first time: two tickets at once, each with its own coordinator. Each coordinator delegated the work to a coder, which started a background task while continuing its own turn. Both wrote almost the same thing before closing: "I'm going to wait for the completion notification," with &lt;code&gt;success&lt;/code&gt; status. No error in sight; the process reported that it was waiting.&lt;/p&gt;

&lt;p&gt;It wasn't waiting: it had finished, and it wasn't coming back. Ten minutes later — six hundred seconds — the process that actually was waiting cut off both coders with the same message: &lt;code&gt;Background tasks still running after 600s; terminating. Set CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0 to wait indefinitely.&lt;/code&gt; The record went untouched. Cost: USD 2.79. Two coordinators, zero tickets done.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ten minutes
&lt;/h2&gt;

&lt;p&gt;Running something in headless mode — no interactive terminal, &lt;code&gt;claude -p&lt;/code&gt; answers once — lets you delegate to an agent "in the background": it starts another one and keeps going with its own work while the second one works apart, and finds out when the second one signals it's done. It's fine for a short question. But the harness ships with a default wait ceiling — ten minutes —, meant for a one-off query, not for a loop that writes, reviews, and fixes code. If it takes longer, the harness doesn't warn that it got tired: it kills the process underneath, and the one above is no longer around to find out.&lt;/p&gt;

&lt;p&gt;Eighteen minutes after that run started, the fix arrived: one line of code, plus four lines of comments, in the loop script. The next run finished ten tickets that same morning.&lt;/p&gt;

&lt;h2&gt;
  
  
  One line, one repo
&lt;/h2&gt;

&lt;p&gt;What's notable isn't the fix itself — the error message suggests it. The same ceiling had already cut off a run in this project one day earlier, with USD 0.81 lost, and it got logged as pending instead of fixed on the spot. It's still pending six days later. Of eleven loop scripts across the ecosystem of projects I work with — nine active ones plus the two templates every new project is born from — only one has the line: the one that suffered it that morning. A project that starts today is born without the fix.&lt;/p&gt;

&lt;p&gt;A one-line fix has the same propagation problem as a ten-page convention: if nothing pushes it between repos, someone has to remember to push it. How many one-line fixes do you have sitting in a single place, with the same bug waiting in the rest?&lt;/p&gt;

</description>
      <category>claude</category>
      <category>agents</category>
      <category>devops</category>
      <category>ai</category>
    </item>
    <item>
      <title>"Nobody designed the frontmatter"</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Wed, 16 Sep 2026 15:44:06 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/nobody-designed-the-frontmatter-fa0</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/nobody-designed-the-frontmatter-fa0</guid>
      <description>&lt;p&gt;I open any file in these repos in VS Code, and before the first heading the Outline panel shows a block of names: &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;numero&lt;/code&gt;, &lt;code&gt;date&lt;/code&gt;, &lt;code&gt;sesion_fuente&lt;/code&gt;, &lt;code&gt;fecha_creacion&lt;/code&gt;, &lt;code&gt;description&lt;/code&gt;, &lt;code&gt;tags&lt;/code&gt;, &lt;code&gt;destacado&lt;/code&gt;, &lt;code&gt;slug&lt;/code&gt;. This post carries those nine at the top, and will gain a tenth once it's published. Yesterday I asked myself who decided each one. Short answer: nobody, all at once.&lt;/p&gt;

&lt;p&gt;That block is the &lt;em&gt;frontmatter&lt;/em&gt;: a few &lt;code&gt;key: value&lt;/code&gt; lines between two &lt;code&gt;---&lt;/code&gt; at the top of a Markdown file. It isn't part of the text; it's what the file says about itself — title, date, tags, where it came from — in a format a script can read without understanding the prose. Nobody writes it from scratch: every file type has a template with the block already built and the values empty, and whoever creates the file copies it and fills it in; some fields get filled in later by a script — the hook that closes a session, the step that publishes a note. It's there so the site can order the notes, so a hook can fill in dates, so a radar can cross sessions against notes already published. And it's there, above all, for as long as someone reads it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four cradles
&lt;/h2&gt;

&lt;p&gt;Tracing the origin took an afternoon: seven families of fields, born in four places. The sessions one — &lt;code&gt;alias&lt;/code&gt;, &lt;code&gt;fecha_inicio&lt;/code&gt;, &lt;code&gt;estado&lt;/code&gt; — was born on August 16 in another project, when two sessions in a row settled how to name and pick a work session. Six days later, in the template I use to spin up projects, a two-field classification was born, &lt;code&gt;domain&lt;/code&gt; and &lt;code&gt;type&lt;/code&gt;, with a closed vocabulary, to tag rules, agents, tickets, and decisions alike. The third cradle is this project: between September 3 and 15, across five rounds, the fields that identify a note. The fourth, on the 11th: the English version, in a separate file with its own block.&lt;/p&gt;

&lt;h2&gt;
  
  
  What traveled and what didn't
&lt;/h2&gt;

&lt;p&gt;Each family travels by a different mechanism, and none of them is all-or-nothing: the initial scaffold copies everything; the sync only updates agents, rules, and scripts; a hook fills in a handful of fields when each session closes. The result: of the tickets in a project prior to the classification, 1 out of 84 has it; in a later one, 178 out of 178; here, where it was born, 48 out of 215. And no script reads those two fields at all: the rule that defines them calls them, literally, "metadata".&lt;/p&gt;

&lt;p&gt;The second gap is stranger. The same hook, when registering a new session, fills in its identifier with a regular expression; since that field starts out empty, it also swallows the next line — the start-date one. 254 out of 260 sessions in one project were born without it, and nobody noticed until another script had to infer it from the file name.&lt;/p&gt;

&lt;h2&gt;
  
  
  What holds up
&lt;/h2&gt;

&lt;p&gt;The four fields that identify a note — &lt;code&gt;slug&lt;/code&gt;, &lt;code&gt;numero&lt;/code&gt;, &lt;code&gt;fecha_creacion&lt;/code&gt;, &lt;code&gt;tags&lt;/code&gt; — have what the classification never did: a reader that fails loudly. The build breaks if one is missing. And a piece's publication status isn't stored: it's recalculated every time it's queried. Less convenient, impossible to drift out of sync. A field doesn't survive because of the intent it was created with, but because someone, or something, still reads it.&lt;/p&gt;

&lt;p&gt;Nobody designed the block above end to end: it has the shape of its own history. Of the ones heading your files, how many do you know who wrote, and how many do you know someone still reads?&lt;/p&gt;

</description>
      <category>metadata</category>
      <category>documentation</category>
      <category>claude</category>
      <category>ai</category>
    </item>
    <item>
      <title>"Two readers, two documents"</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Tue, 15 Sep 2026 14:04:05 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/two-readers-two-documents-8hp</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/two-readers-two-documents-8hp</guid>
      <description>&lt;p&gt;Today, in the Matching repo, there are two files called &lt;code&gt;README.md&lt;/code&gt; with the same modification time: &lt;code&gt;2026-08-18 04:43:43&lt;/code&gt; for the one at the root, &lt;code&gt;04:43:24&lt;/code&gt; for the one inside the code folder. Almost a month has gone by without anyone touching them. Next to them, in the same repository, &lt;code&gt;CLAUDE.md&lt;/code&gt; has 886 lines and its last edit is from September 10 — twenty-three days after those two READMEs. Three documents in the same repo: two frozen, one alive. The difference isn't chance. It's who reads them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why two and not one
&lt;/h2&gt;

&lt;p&gt;The only documented session from that August 18th closed at 04:48 without recording what time it had started. The goal, a single line: "Create the project's README.txt." The original request was actually double —a README "for the general-purpose assistant" and another inside the code folder— though I didn't even get that folder's name right, and along the way it became clear both should be &lt;code&gt;.md&lt;/code&gt;, not &lt;code&gt;.txt&lt;/code&gt;. The second one replaced Vite's default boilerplate outright, without anything else living alongside it.&lt;/p&gt;

&lt;p&gt;What ended up in each one isn't symmetric, it's complementary. The one at the root is a guide for whoever picks up work in the repo, person or assistant: folder structure, agents, skills, backlog rules, the session system, the ticket system. The other is the guide to the product itself: stack (React 19, Vite 8, Supabase), setup, scripts, environment variables, deploy. That same exploration left, along the way, two findings logged to the backlog as out of scope —the outdated stack reference and a duplicated costs folder—, both resolved today.&lt;/p&gt;

&lt;h2&gt;
  
  
  What aged and what didn't
&lt;/h2&gt;

&lt;p&gt;The code README still says the production deploy "is in &lt;code&gt;blocked&lt;/code&gt; status" and that the app shouldn't be assumed to be deployed yet. The real deploy happened on August 22 —it's in the changelog, with a Vercel URL and all— so that file has spent 24 days claiming something that stopped being true the same month it was written. The agent table in the root README didn't hold up either: it lists six, the agents folder has twelve files today, and one of those original six doesn't even live there anymore — the costs one moved up to the user level, outside the repo.&lt;/p&gt;

&lt;p&gt;What didn't age is as telling as what did. The root README's two pointers into &lt;code&gt;CLAUDE.md&lt;/code&gt; sections still resolve today: the agent routing section and the session-naming convention section are exactly where the README says they are. The folder structure it describes still exists as described. And the claim that there's no &lt;code&gt;.claude/skills/&lt;/code&gt; or &lt;code&gt;.claude/commands/&lt;/code&gt; in the repo —skills are global to the user— is still true. What rotted were the states: a ticket, a file count. What held up were the pointers and the conventions.&lt;/p&gt;

&lt;p&gt;That same August 18th, the project's changelog records six product changes: the public landing page, the guided matching assistant, the candidates view, the intro description, the topic search, and the rename of "Materia" to &lt;strong&gt;Área&lt;/strong&gt; across the whole stack. No session documents any of them. The day I wrote how the repo gets documented is, also, the day the least got documented.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stayed
&lt;/h2&gt;

&lt;p&gt;The convention survived even though the files didn't. The template I use for every new project still ships both as boilerplate: 391 lines for the &lt;code&gt;CLAUDE.md&lt;/code&gt; one, 137 for the &lt;code&gt;README.md&lt;/code&gt; one. Of the twelve projects I have today, all twelve have a &lt;code&gt;README.md&lt;/code&gt; and ten have a &lt;code&gt;CLAUDE.md&lt;/code&gt;. The split by audience won; keeping it up to date in every project didn't. Which of your two documents is up to date today, and what does that tell you about who's actually reading it?&lt;/p&gt;

</description>
      <category>documentation</category>
      <category>claude</category>
      <category>ai</category>
      <category>programming</category>
    </item>
    <item>
      <title>"Day three: the loop leaves a trace"</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Mon, 14 Sep 2026 17:38:06 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/day-three-the-loop-leaves-a-trace-2p53</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/day-three-the-loop-leaves-a-trace-2p53</guid>
      <description>&lt;p&gt;Day one I built the agent loop. Day two, the session record. Day three I ran &lt;code&gt;./run-ticket.sh US-04&lt;/code&gt; and nothing happened: no error, no output, no ticket executed. The loop had been running for two days, and that early morning, for the first time, there was no way to see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A script that said nothing
&lt;/h2&gt;

&lt;p&gt;The cause was small and a little embarrassing. &lt;code&gt;run-ticket.sh&lt;/code&gt; looks up the ticket by file name, but the sprint's twenty user stories lived in &lt;code&gt;01-Stories/20260815Sprint/&lt;/code&gt; as &lt;code&gt;NN_slug.md&lt;/code&gt;, without the &lt;code&gt;US-&lt;/code&gt; prefix. &lt;code&gt;grep -i "US-04"&lt;/code&gt; found nothing, and with &lt;code&gt;set -euo pipefail&lt;/code&gt; active the &lt;code&gt;find | grep | head&lt;/code&gt; pipeline failed and killed the script before it could print the error message I had written myself. The bug wasn't just US-04's: it affected all twenty stories in the sprint equally.&lt;/p&gt;

&lt;p&gt;The fix added a second content-matching attempt — look for the &lt;code&gt;# US-04 — ...&lt;/code&gt; heading inside the file when the name isn't enough — and neutralized &lt;code&gt;pipefail&lt;/code&gt; with &lt;code&gt;|| true&lt;/code&gt; on both pipelines, so the existing error block would actually run when it should. Verification turned up a second bug inside the first one: without anchoring the search to the heading, &lt;code&gt;US-01&lt;/code&gt; matched &lt;code&gt;TECH-011_smtp_resend.md&lt;/code&gt; because that ticket mentioned "Discovered during US-01" in its body. It was fixed by anchoring the &lt;code&gt;grep&lt;/code&gt; to &lt;code&gt;^# US-01&lt;/code&gt;. I closed that session at 03:06 without having run &lt;code&gt;US-04&lt;/code&gt; end to end — fixing the matching wasn't the same as invoking the whole loop, and that was left for later.&lt;/p&gt;

&lt;h2&gt;
  
  
  One file per run
&lt;/h2&gt;

&lt;p&gt;The next session, that same early morning, tackled something different: the loop ran, but its output only lived in the terminal. If it got cut off, if I wanted to compare two runs of the same ticket, or simply know how much a run had cost, there was nowhere to look. The decision was to save each run in its own timestamped file — &lt;code&gt;history/&amp;lt;TICKET_ID&amp;gt;_&amp;lt;timestamp&amp;gt;.json&lt;/code&gt; — instead of a fixed file per ticket that each new run would overwrite. That night I was already generating some of those files by hand, without a timestamp, running real tickets in parallel; I was asked whether I'd rather align the convention to that, and I chose to keep the timestamp anyway, so as not to lose the history of repeated runs of the same ticket.&lt;/p&gt;

&lt;p&gt;The implementation was a &lt;code&gt;tee&lt;/code&gt; at the end of the pipeline: &lt;code&gt;claude ... | tee "$HISTORY_FILE"&lt;/code&gt;, verified with &lt;code&gt;bash -n&lt;/code&gt; and an isolated simulation of the pipe, without invoking &lt;code&gt;claude&lt;/code&gt; for real so as not to spend money on a test. I closed that session at 03:46. The first run under the new convention, &lt;code&gt;US-06_20260817-034659.json&lt;/code&gt;, started that very minute.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the record saw that same night
&lt;/h2&gt;

&lt;p&gt;From then on, every run left a trace, and that same night the new record captured something that used to get lost. Between that early morning and the following night, 29 valid runs were left in &lt;code&gt;history/&lt;/code&gt; — close to 500 turns in total, around USD 55 — 20 finished &lt;code&gt;DONE&lt;/code&gt; and 5 &lt;code&gt;BLOCKED&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;US-11 was the one that showed the most. It ran four times: two early &lt;code&gt;BLOCKED&lt;/code&gt;s without even reaching the coder, a third one cut off at 49 turns by "Credit balance is too low" — there, the coder had touched &lt;code&gt;database.types.ts&lt;/code&gt;, a file generated by Supabase, and along the way had broken a flow from another already-finished user story; the coordinator caught it and returned the ticket before the credit ran out — and a fourth, with a narrow fix this time, that finished &lt;code&gt;DONE&lt;/code&gt;. The ticket was still marked &lt;code&gt;blocked&lt;/code&gt; in the record anyway: a false positive from &lt;code&gt;run-ticket.sh&lt;/code&gt; itself, which looked for the word "BLOCKED" anywhere in the text instead of in the status line, and which was fixed by hand after documenting it. The system that finally let me see the loop was born with its own reading bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stayed
&lt;/h2&gt;

&lt;p&gt;Today, a month later, the three decisions from that early morning are still intact in the template I copy into every new project: &lt;code&gt;pipefail&lt;/code&gt; active, the heading fallback, &lt;code&gt;tee&lt;/code&gt; writing to &lt;code&gt;history/&lt;/code&gt;. That's more than 500 files spread across eight projects. Every run of the loop — including the one that wrote this post — leaves a record in some &lt;code&gt;history/&lt;/code&gt;, and that's where, among other things, the cost of each one comes from.&lt;/p&gt;

&lt;p&gt;Day three didn't add a new feature to the product. It added the ability to look back and know what had happened. What do you keep from every agent run, and who reads it afterward?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>"Keeping a record: working across many sessions"</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Sat, 12 Sep 2026 21:29:04 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/keeping-a-record-working-across-many-sessions-1g2k</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/keeping-a-record-working-across-many-sessions-1g2k</guid>
      <description>&lt;p&gt;Day one closed in the small hours and day two started half an hour later, with a short session that didn't touch the product: it defined how sessions would be named, where they would be recorded, and which script would close them on its own. The problem it solved wasn't one of order but of continuity. Every conversation with Claude Code ends, and I wanted to be able to work across many of them —open one, close it, come back the next day or from another project— without losing what had been decided in the previous one. With some distance, it is one of the sessions that weigh the most of all the ones I have: without it this post wouldn't exist, because the material every post in this log comes from is, precisely, the file of each session.&lt;/p&gt;

&lt;h2&gt;
  
  
  A short session and a convention
&lt;/h2&gt;

&lt;p&gt;What got written that night fits in four lines. A naming format, &lt;code&gt;YYYYMMDD-VNN_nombre&lt;/code&gt;: the date, a counter that restarts at &lt;code&gt;V01&lt;/code&gt; every day, and a goal in two or three words. An index, &lt;code&gt;docs/sessions.md&lt;/code&gt;, a Markdown table. A closing &lt;strong&gt;hook&lt;/strong&gt; —a script Claude Code runs on its own every time the model finishes responding— that updates the end time and the status. And a decision on how to find out the session's real identifier, the UUID you need to resume it: take the most recent file in Claude Code's internal sessions directory.&lt;/p&gt;

&lt;p&gt;And one more line, almost at the bottom: &lt;em&gt;"session V01 of 08/15 registered retroactively"&lt;/em&gt;. Day one, which I covered in the previous post, exists as a session because day two wrote it down. The record started by recording backwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the same day already outgrew
&lt;/h2&gt;

&lt;p&gt;That same afternoon, the third session of the day redid half of the design (the second one was opened and closed empty, with no goal: the first record of a session that never was). A single table wasn't enough. Each session got its own file, with a metadata header and fixed sections —goal, decisions, work done, next steps— and the index was reduced to a lightweight list. &lt;code&gt;list_sessions.py&lt;/code&gt; appeared, which reads those headers and shows the sessions ordered by activity, along with the instruction that every conversation should start by showing that list and asking: do we resume one or open another? That is where the real goal got stated: that the work could be spread across many sessions without each one starting from scratch.&lt;/p&gt;

&lt;p&gt;The detail I like most is in the metadata. The fourth session of the day has the same identifier as the third: the "most recent file" heuristic failed the very day it was written, with two sessions opened almost at once. And the times: the first session of the day has a start time; the next three, only the date; the ones from the day after literally say &lt;em&gt;"(time not available)"&lt;/em&gt;. The system was born without knowing how to read a clock.&lt;/p&gt;

&lt;h2&gt;
  
  
  What stayed and what moved
&lt;/h2&gt;

&lt;p&gt;The naming format is today identical, letter by letter, in the template I replicate in every project. There are 426 session files spread across ten projects; 165 of them were opened by the system on its own, using the text of the first message as the name, one for every ticket an agent executed from the terminal without me in the conversation —day two didn't foresee that. The fixed sections of the third session are the same ones in the file of the session I'm writing this in.&lt;/p&gt;

&lt;p&gt;What moved is the responsibility. The closing hook stopped living in my user folder and now travels with each project, and it responds to two events: one refreshes the end time on every turn, the other marks the session as closed exactly once, when the conversation ends. The start —list and ask— stopped being a paragraph in the instructions file: a start hook injects real data (tickets, sessions open for more than 48 hours, alerts), a rule propagated from the template defines the protocol, and the global instructions guarantee the minimum if the other two layers are missing. Three layers for what day two solved with a paragraph, because the paragraph diverged in three out of eight projects as soon as it was copied.&lt;/p&gt;

&lt;p&gt;And there is something day two couldn't know: that the date in the name would end up being the date of this post. Posts in this log carry the date of their source session, not the day they are written. This one carries August 16. That night's convention decided, without meaning to, the order in which everything else is read.&lt;/p&gt;

&lt;p&gt;What it didn't decide is whether this is the right way. Four hundred files, three layers of hooks and rules, a counter that restarts every day: it works, and it is what let me write this. But I built it in one night and kept patching it ever since, without ever comparing it with anything else. If you work with agents across many sessions —how do you keep the record? Is this the best way to manage sessions, or just the one I happened to land on?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Anatomy of a skill</title>
      <dc:creator>Lisandro Reinoso</dc:creator>
      <pubDate>Fri, 11 Sep 2026 22:50:34 +0000</pubDate>
      <link>https://dev.to/lisandro_reinoso_d12ac7b9/anatomy-of-a-skill-295g</link>
      <guid>https://dev.to/lisandro_reinoso_d12ac7b9/anatomy-of-a-skill-295g</guid>
      <description>&lt;p&gt;This started with a surprise, not with a plan. I needed a survey of the state of the art on a topic I was considering working on, so I ran &lt;code&gt;deep-research&lt;/code&gt;, a skill that ships with Claude Code, and it came back with a report I didn't expect from a single run —27 sources, 123 claims extracted, 25 verified, and of those 18 confirmed and 7 refuted— with an executive summary, caveats, and open questions. What struck me wasn't the volume (I already covered the 109 agents in the previous post) but that all of it came out of &lt;strong&gt;a single skill&lt;/strong&gt;. I wanted to see how it was built.&lt;/p&gt;

&lt;p&gt;It wasn't easy to find: it doesn't live in &lt;code&gt;.claude/skills/&lt;/code&gt; or in &lt;code&gt;~/.claude&lt;/code&gt;. It's compiled into the Claude Code binary as a &lt;em&gt;bundled workflow&lt;/em&gt;, and a comment in the code tells its origin: &lt;em&gt;"Ported from bughunter architecture"&lt;/em&gt;. It's 349 lines of JavaScript. Reading them was the research.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a prompt is written
&lt;/h2&gt;

&lt;p&gt;Inside there are three prompts, and all three follow the same shape: a &lt;strong&gt;role&lt;/strong&gt; in the title, the &lt;strong&gt;context&lt;/strong&gt; (the original question plus the specific input), a &lt;strong&gt;task&lt;/strong&gt; as a numbered checklist, an explicit &lt;strong&gt;decision criterion&lt;/strong&gt;, and the &lt;strong&gt;output format&lt;/strong&gt;. The verifier is the clearest one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;VERIFY_PROMPT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;claim&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;v&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;## Adversarial Claim Verifier (voter &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;v&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/3)&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Be SKEPTICAL. Try to REFUTE this claim. ≥2/3 refutations kill it.&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;## Claim under review&lt;/span&gt;&lt;span class="se"&gt;\n\"&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;claim&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;claim&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\"\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;**Supporting quote:** &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;claim&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;quote&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\"\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;## Checklist&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;1. Is the claim actually supported by the quote, or is it an overreach?&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;2. WebSearch for contradicting evidence.&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;3. Is the source quality sufficient for the claim's strength?&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;4. Is the claim outdated?&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;5. Is this a marketing claim / cherry-picked benchmark / forum speculation?&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;**refuted=false** ONLY if: well-supported, current, and source quality matches claim strength.&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Default to refuted=true if uncertain.&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s2"&gt;Structured output only. Evidence MUST be specific.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's exactly the framework I use when I review one of my own prompts —role, context, task, format, constraints— but with two details I don't usually include: the decision criterion is written as a rule (&lt;code&gt;refuted=false&lt;/code&gt; &lt;strong&gt;only&lt;/strong&gt; if…) and ties resolve by default toward the conservative side (&lt;em&gt;"Default to refuted=true if uncertain"&lt;/em&gt;). A prompt that doesn't say what to do when in doubt leaves that decision to the model, and that's where the tidy-but-unfounded answers show up.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a skill is built
&lt;/h2&gt;

&lt;p&gt;A skill is more than a long prompt. What I saw in the file falls into six pieces:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Trigger metadata.&lt;/strong&gt; The &lt;code&gt;meta&lt;/code&gt; declares the name, the description, and the five phases, but the key is &lt;code&gt;whenToUse&lt;/code&gt;: it's what the model reads to decide whether to invoke the skill, and it includes a prior instruction —if the question is underspecified, ask two or three clarifying questions before starting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Tuning constants at the top.&lt;/strong&gt; No magic numbers buried in the code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;VOTES_PER_CLAIM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;REFUTATIONS_REQUIRED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAX_FETCH&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAX_VERIFY_CLAIMS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;25&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. One schema per agent.&lt;/strong&gt; Each of the five agent types returns JSON validated against a schema. That's what makes the pipeline composable: one agent's output is the next one's typed input.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Prompts as functions.&lt;/strong&gt; &lt;code&gt;SEARCH_PROMPT(angle)&lt;/code&gt;, &lt;code&gt;FETCH_PROMPT(source, angle)&lt;/code&gt;, &lt;code&gt;VERIFY_PROMPT(claim, v)&lt;/code&gt;: they take the input and return the text. The prompt isn't copied, it's instantiated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Explicit orchestration.&lt;/strong&gt; Search and fetch go through &lt;code&gt;pipeline()&lt;/code&gt;: each angle moves on to fetching its sources as soon as it's done, without waiting for the others. Before verification there's a barrier —and the comment says so: &lt;em&gt;"Barrier here is intentional"&lt;/em&gt;— because the pool of claims has to be complete before it can be ranked. Then, nested &lt;code&gt;parallel()&lt;/code&gt;: 25 claims × 3 votes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;6. Defensive design.&lt;/strong&gt; Every early exit (no claims, everything refuted, failed synthesis) returns a useful result with &lt;code&gt;stats&lt;/code&gt; instead of throwing. A &lt;code&gt;null&lt;/code&gt; vote counts as an abstention, not as a free pass:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;survives&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;valid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="nx"&gt;REFUTATIONS_REQUIRED&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;refuted&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;REFUTATIONS_REQUIRED&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And the final result carries &lt;code&gt;agentCalls&lt;/code&gt;: the skill computes its own cost (&lt;code&gt;1 + angles + sources + claims × 3 + 1&lt;/code&gt;). That's where the 109 came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  Trimming the pattern
&lt;/h2&gt;

&lt;p&gt;The proof that I understood the pattern was reusing it. A couple of later runs on that same topic got cut off by the session limit before verification, and instead of repeating them in full I wrote &lt;code&gt;reverify-linea-c&lt;/code&gt;: same &lt;code&gt;meta&lt;/code&gt;, same constants, same verdict schema, and the same three-vote verifier, but with no Scope, Search, or Fetch —those phases were replaced by an array of 22 already-extracted claims. The whole skeleton inherited, one array of my own. All 22 came back confirmed.&lt;/p&gt;

&lt;p&gt;What I take away: a good skill isn't a long prompt, it's a short, well-formed prompt, instantiated many times by an orchestration that knows where to wait and where not to, and that measures what it spends. And it reads in an afternoon.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
