<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Evgenii Arsentev</title>
    <description>The latest articles on DEV Community by Evgenii Arsentev (@arsentev).</description>
    <link>https://dev.to/arsentev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4106171%2Fe4a438a3-4658-4364-9c9a-6a8c5b396ad4.jpg</url>
      <title>DEV Community: Evgenii Arsentev</title>
      <link>https://dev.to/arsentev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/arsentev"/>
    <language>en</language>
    <item>
      <title>Nobody with a crawler token asked for my llms.txt - what I saw in 15 days of nginx logs</title>
      <dc:creator>Evgenii Arsentev</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:28:35 +0000</pubDate>
      <link>https://dev.to/arsentev/nobody-with-a-crawler-token-asked-for-my-llmstxt-what-i-saw-in-15-days-of-nginx-logs-o94</link>
      <guid>https://dev.to/arsentev/nobody-with-a-crawler-token-asked-for-my-llmstxt-what-i-saw-in-15-days-of-nginx-logs-o94</guid>
      <description>&lt;p&gt;I run a digital health startup as CEO (and a personal site, arsentev.ai, about practical AI). I put /llms.txt and /llms-full.txt on my site, and I wanted to see who actually reads them.&lt;/p&gt;

&lt;p&gt;The two files are a plain-text summary of the site for language models. They were first served on 9 September 2026. I kept the complete nginx access logs of one origin server for 28 August - 11 September 2026. It is 15 days and 345,808 parsed requests.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I counted
&lt;/h2&gt;

&lt;p&gt;A request was a crawler if the User-Agent contained one of 22 tokens published by search and AI operators (GPTBot, ClaudeBot, PerplexityBot and similar). 46,155 requests carried such tokens, and 20 distinct tokens appeared. None of them requested /llms.txt or /llms-full.txt. I also checked every path ending in llms.txt, including /.well-known/llms.txt. No crawler token there either.&lt;/p&gt;

&lt;p&gt;Then I looked only at 9-11 September, after the files were first served. In these days crawler tokens made 15,906 requests. Of them 698 were for /robots.txt (11 tokens) and 282 for /sitemap.xml. So the crawlers were on the site, they just didn't ask for my files.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who asked for them
&lt;/h2&gt;

&lt;p&gt;All 67 requests for the context files came from clients without a crawler token. 53 of them were command-line HTTP clients (curl and the like).&lt;/p&gt;

&lt;p&gt;About a third of crawler-token requests bypassed the CDN. None of those came from the operators' published IP ranges, so I think much of that traffic is imitation. The robots.txt and sitemap numbers don't rest on it.&lt;/p&gt;

&lt;p&gt;One more thing I noticed. Nothing in any standard tells a crawler that llms.txt exists. A sitemap is different, because robots.txt can point to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  My mistake
&lt;/h2&gt;

&lt;p&gt;First I published different numbers: 331,758 / 44,005 / 577. They came from a log that was cut at about 12:09 UTC on the last day. I ran everything again on the complete log and corrected the report. The numbers above are from the complete log.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;This is one origin. Three hostnames are pooled in one log (host wasn't logged). After deployment I have only about two days of exposure. User agents are self-asserted, and I did not record links to the files. So it is absence on one server, not proof for the web.&lt;/p&gt;

&lt;p&gt;After this I added $host to my nginx log format. The next study can split the hostnames.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I suggest
&lt;/h2&gt;

&lt;p&gt;Just check your own access log before you decide that anyone reads your llms.txt. Look for requests to it by user-agent. Grep is enough for this.&lt;/p&gt;

&lt;p&gt;The report, the scripts (Python standard library only) and the aggregate counts without IP addresses are here: &lt;a href="https://doi.org/10.5281/zenodo.23018429" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.23018429&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Research index: &lt;a href="https://arsentev.ai/research" rel="noopener noreferrer"&gt;https://arsentev.ai/research&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI assistance: drafted and edited with an AI tool; the data, measurements and conclusions are my own.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>llm</category>
      <category>devops</category>
    </item>
    <item>
      <title>My company has no full-time developers, DevOps or designers. Here's what three $200 Claude subscriptions do instead</title>
      <dc:creator>Evgenii Arsentev</dc:creator>
      <pubDate>Mon, 28 Sep 2026 20:30:15 +0000</pubDate>
      <link>https://dev.to/arsentev/my-company-has-no-full-time-developers-devops-or-designers-heres-what-three-200-claude-46a7</link>
      <guid>https://dev.to/arsentev/my-company-has-no-full-time-developers-devops-or-designers-heres-what-three-200-claude-46a7</guid>
      <description>&lt;p&gt;At night the agent shifts run by the clock. New builds of educational games get built, and a separate fresh agent checks each one by screenshots. The mail agent of one business line starts every hour, night too. In the morning the results are ready. In the businesses I run there are no full-time developers, no DevOps, no designers, no analysts and no product manager. At US median pay a developer, a product manager, an analyst and a designer come to about $38,500 a month (Bureau of Labor Statistics, see the note at the end). The agents doing that work run on three $200 Claude subscriptions, so $600 a month (about 64 times less).&lt;/p&gt;

&lt;p&gt;There is me, there is a product builder, and there are AI agents that run on a schedule on separate servers (not on my laptop). I'm a medical doctor by education and a CEO. This is not a forecast, it's what runs today, what I kept for people and what didn't work.&lt;/p&gt;

&lt;h2&gt;
  
  
  The order we did it in
&lt;/h2&gt;

&lt;p&gt;Designers went first, as soon as agents got good enough at design. Analysts went a long time ago too. The product manager role is closed too, and Claude does its tasks now (tasks that took days now take minutes). The developer was the last one.&lt;/p&gt;

&lt;p&gt;We kept our former CTO on a small retainer as a safety net, in case an agent breaks something serious. So far we haven't needed him once. Everything our DevOps used to do, the agents do.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is on the schedule
&lt;/h2&gt;

&lt;p&gt;Each business line has its own agent shift. It starts by the clock, does its job and stops. I only start things, the agents do the rest.&lt;/p&gt;

&lt;p&gt;Deploys. Releases to all our projects and zones are done by agents.&lt;/p&gt;

&lt;p&gt;Our own CRM. We didn't buy one. The agents built a CRM from scratch, and the agents also work inside it. So why buy a CRM if we built our own?&lt;/p&gt;

&lt;p&gt;Email. Every business line has its own mail shift, and the agent sorts the mailbox of that line only.&lt;/p&gt;

&lt;p&gt;Meetings. Scheduling meetings is on the agents too.&lt;/p&gt;

&lt;p&gt;Technical support. Technical support is done by agents.&lt;/p&gt;

&lt;p&gt;Content. On a news and learning site I founded, agents publish news four times a day and scan new open-source AI releases three times a day. The games project I mentioned works overnight, and the builder's opinion about its own work doesn't count - only the fresh agent's check does.&lt;/p&gt;

&lt;p&gt;And not everything is an agent. Backups for example are plain scripts on a timer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The supervisor of the supervisor
&lt;/h2&gt;

&lt;p&gt;This is the part I like most. In one of our US business lines we have a call-center team. The operators and their supervisor are real people. But the supervisor gets tasks from an agent.&lt;/p&gt;

&lt;p&gt;Every day the agent goes through the operators' work - the calls, their transcripts, the results in the CRM. Then it writes to the human supervisor what to do today and tomorrow (with a number and a deadline), who needs help, who should be replaced and what the plan is for each person. Each operator gets a separate message with 3-4 small skills to fix, taken from the transcripts of their last calls. Every evening each skill gets a ✅ or ❌, and after three ✅ in a row the skill is closed.&lt;/p&gt;

&lt;p&gt;The agent asks three things in every shift. Is each operator better or worse than last time? Did the supervisor improve the work of their people (by the people's numbers, not their own)? What should the supervisor do today and tomorrow? If the team is stuck, that is written as the supervisor's failure. The agent only suggests. Decisions about people are made by people and are not posted in the team channel, and the operators were told openly that calls are checked every day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Subscription, not tokens
&lt;/h2&gt;

&lt;p&gt;The agent work runs on subscription, not on paying per token through the API. On a flat price the students on my course experiment more boldly. The limits get eaten by re-reading context - in my own measurement more than 85% of the modeled spend was context work (DOI &lt;a href="https://doi.org/10.5281/zenodo.22759216" rel="noopener noreferrer"&gt;10.5281/zenodo.22759216&lt;/a&gt;). And every shift is short, it does one thing and exits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where people stayed
&lt;/h2&gt;

&lt;p&gt;An agent can do the work, but it can't be accountable for it. Every agent workflow has a named person who owns the result and answers for it as if a colleague had done the work.&lt;/p&gt;

&lt;p&gt;Checks stay mandatory. Agents don't show doubt (a wrong result comes with the same confidence as a correct one). Nothing ships just because an agent finished it. It ships because it passed a check. And an agent gets access for a task, not a standing role.&lt;/p&gt;

&lt;h2&gt;
  
  
  What didn't work for us
&lt;/h2&gt;

&lt;p&gt;First, vague tasks. A vague task given to a person gets clarified over coffee. Given to an agent, it gets done confidently and wrong. So the main management skill for me now is writing the task - the goal, the limits, what "done" means and how we check it.&lt;/p&gt;

&lt;p&gt;Second, measuring the wrong thing. The first AI judge of calls in our CRM scored them against a script checklist. It measured if the operator followed the script. I think we should judge an operator by whether the partner gets to a purchase, not by the checklist.&lt;/p&gt;

&lt;p&gt;Third, screens nobody opens. Our CRM has supervisor pages, and in 30 days they were opened 9 times. What works is the plain daily message to each person.&lt;/p&gt;

&lt;p&gt;Fourth, AI summaries garbled a supervisor's name. So for names we trust only the spelling a person confirmed.&lt;/p&gt;

&lt;h2&gt;
  
  
  A case I documented
&lt;/h2&gt;

&lt;p&gt;A company whose case I documented replaced a video studio with Claude agents - 1,497 videos across 67 channels in 101 days, at $2.88 per video. With people it would take 42 to 78 staff and cost 138 to 935 times the AI subscriptions (DOI &lt;a href="https://doi.org/10.5281/zenodo.22802229" rel="noopener noreferrer"&gt;10.5281/zenodo.22802229&lt;/a&gt;). It's not our company, and I only wrote the analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would tell another CEO
&lt;/h2&gt;

&lt;p&gt;Just start with one function where the result is easy to check. Name a person who owns the result and track the cost per result from the first week. I look at the cost of a task that was done and accepted (with reruns and review time), not the price of the tool.&lt;/p&gt;

&lt;p&gt;Keep the shifts short. But don't clear the context after every single task either. In my own small experiment that was about a third more expensive than clearing it every three tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the limit moved
&lt;/h2&gt;

&lt;p&gt;Before, we were limited by people and the job market. Now the limit is Claude tokens and computing power - memory and GPUs. That's why I rent separate servers, including GPU machines, so the agents don't run on a laptop that chokes.&lt;/p&gt;

&lt;p&gt;I think memory is becoming the new gold. And I think for everyday agent work we'll need something like what ASICs did for bitcoin, a class of chips made for this kind of load.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Note on the numbers: US median annual pay from the U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (bls.gov/oes) - software developers (15-1252) $135,980, project management specialists (13-1082, the closest BLS category to a product manager) $102,320, data scientists (15-2051) $120,230, web and digital interface designers (15-1255) $104,000. Monthly = annual / 12. These are market numbers, not our salaries.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Evgenii Arsentev, MD, PhD — CEO, digital health; founder, arsentev.ai&lt;/strong&gt;&lt;br&gt;
Evgenii Arsentev is a medical doctor and a digital health CEO whose businesses run almost entirely on AI agents. He is the founder of arsentev.ai.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://arsentev.ai/press/no-developers-three-claude-subscriptions" rel="noopener noreferrer"&gt;arsentev.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI assistance: drafted and edited with an AI tool; the data, measurements and conclusions are my own.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>startup</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I Ran a Radar on 490 Open-Source AI Projects. Then I Checked Them Again</title>
      <dc:creator>Evgenii Arsentev</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:51:11 +0000</pubDate>
      <link>https://dev.to/arsentev/i-ran-a-radar-on-490-open-source-ai-projects-then-i-checked-them-again-cba</link>
      <guid>https://dev.to/arsentev/i-ran-a-radar-on-490-open-source-ai-projects-then-i-checked-them-again-cba</guid>
      <description>&lt;h2&gt;
  
  
  The selection
&lt;/h2&gt;

&lt;p&gt;A radar I run at arsentev.ai/github selects and reviews open-source AI projects from GitHub, Hugging Face and Hacker News, one page per project. Between 11 June and 12 September 2026 it selected 490 projects: 422 from GitHub, 63 from Hugging Face, 5 from Hacker News. By kind that's 427 repositories, 48 Hugging Face models, 15 Hugging Face Spaces.&lt;/p&gt;

&lt;p&gt;By month of selection: June 70, July 134, August 204, September (up to the 12th) 82.&lt;/p&gt;

&lt;p&gt;Median GitHub stars at the moment of selection was 493. Top languages across the repositories: Python 135, TypeScript 104, JavaScript 51, Rust 19, Swift 16. On Hugging Face the most common tasks were text-generation 16, image-text-to-text 16, image-to-video 7, text-to-speech 6.&lt;/p&gt;

&lt;h2&gt;
  
  
  One snapshot so far
&lt;/h2&gt;

&lt;p&gt;There is one snapshot of this data, taken on 12 September 2026, the last day of selection. Depending on the project, the gap between when it was selected and the snapshot ranges from 0 days to about three months.&lt;/p&gt;

&lt;p&gt;On that snapshot, of the 427 GitHub repositories: 416 were still available, 10 deleted (the API returns 404), 1 archived.&lt;/p&gt;

&lt;p&gt;Of the available repositories, 19% have no licence or a licence GitHub doesn't recognise. The breakdown: MIT 218, Apache-2.0 93, NOASSERTION (GitHub can't tell) 43, no licence file 36, AGPL-3.0 16, GPL-3.0 8.&lt;/p&gt;

&lt;p&gt;71.6% had a push in the last 30 days.&lt;/p&gt;

&lt;p&gt;The sample is curated, not random. It's what the radar picked, so it says nothing about GitHub as a whole.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tracking going forward
&lt;/h2&gt;

&lt;p&gt;From 12 September a script takes a daily snapshot of every selected project. For GitHub repositories that's stars, forks, open issues, licence, last push, archived status. For Hugging Face models and Spaces it's likes, downloads, last modified. Removed repositories stay in the data with their status attached, so the set doesn't drift toward only the ones that are still around.&lt;/p&gt;

&lt;p&gt;The dataset is open, under CC BY 4.0, on Zenodo, with three files: items.csv, snapshots.csv and stats.json. Every number above comes from the build script running on those CSV files. A new version comes out on the 1st of each month, and the series DOI always points to the latest one: &lt;a href="https://doi.org/10.5281/zenodo.22730450" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.22730450&lt;/a&gt;. The numbers in this piece are from the 2026-09-13 edition: &lt;a href="https://doi.org/10.5281/zenodo.22730451" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.22730451&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three checks before you build on one
&lt;/h2&gt;

&lt;p&gt;Before picking up a repository from a list like this, or from any trending list, these are the three things I'd check myself:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;does the repository still open&lt;/li&gt;
&lt;li&gt;is the licence field on GitHub set to a known licence, not blank and not NOASSERTION&lt;/li&gt;
&lt;li&gt;was there a recent push&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these take more than a minute, and all three are visible right on the repository page, before you clone anything or read a line of code.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Evgenii Arsentev, PhD. The dataset and the radar pages are at &lt;a href="https://arsentev.ai/github" rel="noopener noreferrer"&gt;arsentev.ai/github&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI assistance: drafted and edited with an AI tool; the data, measurements and conclusions are my own.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>github</category>
      <category>datascience</category>
    </item>
    <item>
      <title>All my agent's tests were green, and they told me nothing</title>
      <dc:creator>Evgenii Arsentev</dc:creator>
      <pubDate>Sat, 26 Sep 2026 08:11:32 +0000</pubDate>
      <link>https://dev.to/arsentev/all-my-agents-tests-were-green-and-they-told-me-nothing-3n9n</link>
      <guid>https://dev.to/arsentev/all-my-agents-tests-were-green-and-they-told-me-nothing-3n9n</guid>
      <description>&lt;p&gt;I'm CEO of AskDocDoc (telehealth) and in our company the software development is done by AI agents. This is about a fully green test run that told me nothing (and it was my own experiment).&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I ran a coding agent on twelve fixed programming tasks under six context-clearing policies (a fresh session every 1, 2, 3, 4, 6 and 12 tasks). Six replicates each, same tasks, same order - 36 runs. Every run ended with the full test suite as a gate. My reasoning was simple - a cost comparison is worthless if the cheaper policy also did less work.&lt;/p&gt;

&lt;p&gt;The gate reported 4,086 tests executed and 4,086 passed. Zero failures (not in one run, not in one condition). And in the same files the modelled cost per run went from $2.165 under the most aggressive clearing to $1.622 at the observed optimum, and back to $1.780 when the context was never cleared. That's a 33.5% spread between the most and least expensive of six conditions (it's the gap between the extremes of six means, not a property of one policy).&lt;/p&gt;

&lt;p&gt;So if you read only the 36 green reports you conclude the six setups are equivalent. The cost data in the same deposit says they are not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found when I looked closer
&lt;/h2&gt;

&lt;p&gt;The 4,086 is a grand total. Per condition it was 701, 697, 687, 699, 661 and 641 tests. Per run it went from 97 to 127 - that's 30% between the least and most checked run. In my own report I wrote that the design held "the verification suite constant". It wasn't, and I corrected it in the paper.&lt;/p&gt;

&lt;p&gt;The reason is boring. The task prompts told the agent to create the test file itself and run &lt;code&gt;node --test&lt;/code&gt;. Nine of twelve prompts set a minimum number of cases (between five and ten) and none set a maximum. So the agent wrote the criterion and then satisfied it. A run that did less would not go red - it would just write fewer tests.&lt;/p&gt;

&lt;p&gt;I also checked the tempting story - longer sessions did less work and the gate didn't notice. It's not supported by the data (it doesn't survive removal of one endpoint condition, and it doesn't survive Bonferroni correction over the twelve columns I looked at). So I report it as a negative result.&lt;/p&gt;

&lt;p&gt;And zero failures constrains less than it feels. With the run as the unit, 0 in 36 is consistent with a per-run failure probability up to 8.3% at 95% confidence. With the test as the unit, 0 in 4,086 gives 0.073%. Both are correct, and they differ by two orders of magnitude. So a pass count means nothing until the record says what the unit was.&lt;/p&gt;

&lt;p&gt;This isn't only my corpus. Nicolás Rocchia found a checker where a missing companion file leaves no record at all, so the same artifact reports 40/40 with the file and 38/38 without it - passing both times (&lt;a href="https://lists.w3.org/Archives/Public/public-agent-conformance/2026Sep/0011.html" rel="noopener noreferrer"&gt;his message&lt;/a&gt; to the W3C list). His reading - the report gets smaller, not worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the record should have said
&lt;/h2&gt;

&lt;p&gt;Every green mark in my runs was correctly reported (as far as the record shows). The problem is a property of the set, and no per-check field can carry that. So the proposal in my paper is one field and one list at the result-set level - did any check change state across the compared conditions, and which ones did.&lt;/p&gt;

&lt;p&gt;The field needs three values, not a boolean - &lt;code&gt;some&lt;/code&gt;, &lt;code&gt;none&lt;/code&gt; and &lt;code&gt;incomparable&lt;/code&gt;. In my runs the agent wrote a fresh suite every time, so no check exists across conditions. The honest value for my study is &lt;code&gt;incomparable&lt;/code&gt;. This is the record it should have emitted:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"compared"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"ucurve-p1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ucurve-p2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ucurve-p3"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
               &lt;/span&gt;&lt;span class="s2"&gt;"ucurve-p4"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ucurve-p6"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ucurve-p12"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"check_set_origin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"generated_by_run"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"discrimination"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"incomparable"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"shared_checks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;shared_checks&lt;/code&gt; says how many checks the conditions really have in common. &lt;code&gt;check_set_origin&lt;/code&gt; says if the suite was &lt;code&gt;fixed&lt;/code&gt;, &lt;code&gt;generated_by_run&lt;/code&gt; or &lt;code&gt;mixed&lt;/code&gt;. Under &lt;code&gt;incomparable&lt;/code&gt; the list of checks that moved must be absent, not empty (an empty list is exactly what &lt;code&gt;none&lt;/code&gt; emits, and the two must not be confused).&lt;/p&gt;

&lt;p&gt;If you compare agent runs, the practical part is short. Log the test count next to pass/fail for every run. If the agent writes its own tests, cross-run comparison of verdicts is not just noisy - it's ill-defined. And a record that can't say &lt;code&gt;incomparable&lt;/code&gt; will say &lt;code&gt;none&lt;/code&gt;, and a reader will hear "equivalent".&lt;/p&gt;

&lt;p&gt;One honest limit - my claim that the suite could have gone red comes from how the harness was built, not from a mutation experiment. The repair is to run one. The idea grew out of the W3C Agent Conformance discussion and nothing in it is adopted yet.&lt;/p&gt;

&lt;p&gt;The paper with all the numbers is open on Qeios - &lt;a href="https://doi.org/10.32388/0BV3Z8" rel="noopener noreferrer"&gt;A Suite That Does Not Discriminate&lt;/a&gt; (DOI 10.32388/0BV3Z8). My other measurements are at &lt;a href="https://arsentev.ai/research" rel="noopener noreferrer"&gt;arsentev.ai/research&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Evgenii Arsentev, PhD&lt;br&gt;
CEO, AskDocDoc&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI assistance: drafted and edited with an AI tool; the data, measurements and conclusions are my own.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>97% of what my coding agent billed for was re-reading its own context</title>
      <dc:creator>Evgenii Arsentev</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:11:32 +0000</pubDate>
      <link>https://dev.to/arsentev/97-of-what-my-coding-agent-billed-for-was-re-reading-its-own-context-4o2</link>
      <guid>https://dev.to/arsentev/97-of-what-my-coding-agent-billed-for-was-re-reading-its-own-context-4o2</guid>
      <description>&lt;p&gt;Token counters tell you how much you spent. I wanted a different number: &lt;strong&gt;how much of what I paid for was new text the model actually produced.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So I measured it against a complete local log corpus of my own agentic coding work: &lt;strong&gt;722 sessions, 150,902 model calls, 34.56 billion tokens.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;By tokens billed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Token class&lt;/th&gt;
&lt;th&gt;Share of all tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cache read (re-reading context)&lt;/td&gt;
&lt;td&gt;97.05%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache write (storing context)&lt;/td&gt;
&lt;td&gt;2.57%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output (generation)&lt;/td&gt;
&lt;td&gt;0.38%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fresh input&lt;/td&gt;
&lt;td&gt;0.01%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For every token the models produced, roughly &lt;strong&gt;256 tokens of cached context were read back in.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Cache pricing softens that ratio, but not the conclusion. Under published per-token list rates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Share of cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Re-reading context&lt;/td&gt;
&lt;td&gt;55.89%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Writing context to cache&lt;/td&gt;
&lt;td&gt;31.89%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Generating new text&lt;/td&gt;
&lt;td&gt;12.18%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fresh input&lt;/td&gt;
&lt;td&gt;0.04%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Context handling 87.8%, generation 12.2%.&lt;/strong&gt; Re-reading context alone costs &lt;strong&gt;4.59x&lt;/strong&gt; as much as everything the models wrote. And cache writing is not a rounding error: at 31.89% it is the second-largest line, larger than generation. Caching does not make context free, it moves part of the price to the moment state is stored.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it looks like this
&lt;/h2&gt;

&lt;p&gt;An agent doesn't send a fresh question on every step. It sends the whole conversation so far: instructions, every file it read, every command it ran. Then it adds one step and sends everything again. A token that enters the context on step 1 of a 60-step run is paid for 60 times.&lt;/p&gt;

&lt;p&gt;I put a number on that with a derived measure, &lt;strong&gt;input amplification&lt;/strong&gt;: total input tokens billed in a session, divided by that session's peak context size. Across the 590 sessions with at least three model calls, the median is &lt;strong&gt;23.7x&lt;/strong&gt;, the 90th percentile &lt;strong&gt;124.5x&lt;/strong&gt;, the maximum &lt;strong&gt;4,182x&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The median session pays for its peak working state about 24 times. One session in ten pays for it more than 125 times. A session's cost is not mainly set by how big its context is, nor by how much it writes, but by the product of context size and how many times that context gets traversed.&lt;/p&gt;

&lt;p&gt;Cost is also concentrated: an estimated 80% of the corpus total falls on 21 of 722 sessions (2.9%), and sessions longer than 200 model calls - 8% of sessions - account for an estimated 92.8% of spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part you can actually change
&lt;/h2&gt;

&lt;p&gt;None of the above is a knob. How often you clear the context is.&lt;/p&gt;

&lt;p&gt;I ran a controlled comparison: twelve fixed programming tasks under six context-clearing policies - a fresh session every 1, 2, 3, 4, 6 and 12 tasks - six replicates each, with the model, the tasks and their order held constant. 36 runs.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Clear after every&lt;/th&gt;
&lt;th&gt;Mean modeled cost per run&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1 task&lt;/td&gt;
&lt;td&gt;$2.165&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2 tasks&lt;/td&gt;
&lt;td&gt;$1.854&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;3 tasks&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.622&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4 tasks&lt;/td&gt;
&lt;td&gt;$1.704&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6 tasks&lt;/td&gt;
&lt;td&gt;$1.679&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12 tasks (never)&lt;/td&gt;
&lt;td&gt;$1.780&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cost is not monotone in session length. It falls, bottoms out at three tasks, then rises again.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Clearing after &lt;strong&gt;every&lt;/strong&gt; task costs &lt;strong&gt;33.5%&lt;/strong&gt; more than clearing every third (p = 0.0022, Holm-adjusted 0.011).&lt;/li&gt;
&lt;li&gt;Every second task: &lt;strong&gt;+14.3%&lt;/strong&gt; (p = 0.0022, Holm 0.011).&lt;/li&gt;
&lt;li&gt;Going from clearing every task to every third cuts cost by &lt;strong&gt;25.1%&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The reason the constant-clearing end is expensive: a cache write costs &lt;strong&gt;20x&lt;/strong&gt; a cache read. Clear after every task and you keep paying to rebuild state you just threw away. At the optimum the spend splits 39.0% cache read, 32.3% cache write, 28.7% output.&lt;/p&gt;

&lt;p&gt;Two things I want to state plainly rather than round off:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The middle is a plateau, not a point.&lt;/strong&gt; Every fourth task is +5.1% against every third (p = 0.17), every sixth +3.5% (p = 0.45). The data cannot separate 3 from 4 or 6. "Clear every few tasks" is the finding; "clear every third" is just where the observed minimum landed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never clearing is not significantly worse.&lt;/strong&gt; It comes out +9.7% against every third, p = 0.046, Holm-adjusted 0.14. An earlier version of this work claimed both extremes were significantly more expensive than the optimum. That claim is withdrawn for never clearing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measuring it yourself
&lt;/h2&gt;

&lt;p&gt;I packaged the measurement as a small open-source tool, &lt;a href="https://github.com/arsentev-ai/contextburn" rel="noopener noreferrer"&gt;contextburn&lt;/a&gt;. It reads the transcripts your agent already writes on your machine and makes no network calls.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;contextburn
contextburn detail 24
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also runs as an MCP server, so the agent can check its own efficiency mid-session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add contextburn &lt;span class="nt"&gt;--&lt;/span&gt; uvx contextburn mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Two traps, in case you write your own counter
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Streaming logs record the same model call more than once.&lt;/strong&gt; An early snapshot and a final record share one message id. Count both and you double the call; keep only the first and you halve the output. Take the element-wise maximum per message id.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Prices change the answer, and they will change yours.&lt;/strong&gt; Both datasets behind this post were re-published with corrected costs, and the errors were mine. Cache reads had been priced at a rate that applies only to a newer model, and every cache write had been priced at the 5-minute rate when the logs show mostly 1-hour writes, billed at 2x input rather than 1.25x. Token counts, call counts and the composition table did not move at all. Every dollar figure did: the corpus split went from 83.5% / 16.5% to the &lt;strong&gt;87.8% / 12.2%&lt;/strong&gt; above. If you report a cost-weighted share, it is only as current as your price table - and worth re-deriving before you quote it anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;This is a single-practitioner case study, not a sample of a population, and the dollar figures are a model applied to logs rather than an invoice.&lt;/p&gt;

&lt;p&gt;Cache-write duration cannot be verified for part of one machine's logs. If every write were instead a 5-minute write, the corpus split would be 86.2% / 13.8% - the direction holds under every assumption I can test, only the size moves.&lt;/p&gt;

&lt;p&gt;In the controlled runs, "no test failures" means tests the agent wrote itself, and the count differed by condition: 116.8 tests per run on average when clearing after every task, 106.8 when never clearing. The three-task reference condition was selected after the fact, as the observed minimum.&lt;/p&gt;

&lt;p&gt;I'd like to see the same measurement run against other people's logs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The data
&lt;/h2&gt;

&lt;p&gt;Report, datasets and recomputation scripts are public:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Corpus report and dataset: &lt;a href="https://doi.org/10.5281/zenodo.22759216" rel="noopener noreferrer"&gt;10.5281/zenodo.22759216&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Peer-reviewed report: &lt;a href="https://doi.org/10.32388/0BV3Z8" rel="noopener noreferrer"&gt;10.32388/0BV3Z8&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Context-clearing experiment, 36 runs: &lt;a href="https://doi.org/10.5281/zenodo.22759217" rel="noopener noreferrer"&gt;10.5281/zenodo.22759217&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Evgenii Arsentev, PhD - Chief Executive Officer. I measure how AI agents spend money, and publish the data and the scripts. More at &lt;a href="https://arsentev.ai" rel="noopener noreferrer"&gt;arsentev.ai&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI assistance: drafted and edited with an AI tool; the data, measurements and conclusions are my own.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>claudecode</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
