<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Suman Debnath</title>
    <description>The latest articles on DEV Community by Suman Debnath (@suman_debnath_1).</description>
    <link>https://dev.to/suman_debnath_1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4113658%2Fd5cda7ab-3e51-4954-bece-4fb4c710fdca.jpeg</url>
      <title>DEV Community: Suman Debnath</title>
      <link>https://dev.to/suman_debnath_1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/suman_debnath_1"/>
    <language>en</language>
    <item>
      <title>ChatGPT started citing my site. Here is what I changed, and what I still cannot prove.</title>
      <dc:creator>Suman Debnath</dc:creator>
      <pubDate>Wed, 09 Sep 2026 10:39:49 +0000</pubDate>
      <link>https://dev.to/suman_debnath_1/chatgpt-started-citing-my-site-here-is-what-i-changed-and-what-i-still-cannot-prove-4a31</link>
      <guid>https://dev.to/suman_debnath_1/chatgpt-started-citing-my-site-here-is-what-i-changed-and-what-i-still-cannot-prove-4a31</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;In short —&lt;/strong&gt; ChatGPT began naming Suman Debnath for the query "who is Suman Debnath" two days after a full answer-engine optimisation pass — a generated llms.txt, extractable answer blocks, entity disambiguation and structured data. Keyword stuffing was tried first and did nothing. Two days is not proof of cause, and Claude, Gemini and Grok still do not cite the site.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;First observed&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;24 August 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Queries&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"who is suman debnath", "suman debnath portfolio"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Devices checked&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Three, including one where the name had never been typed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Optimisation work began&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Roughly 19 August 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Engines citing the site&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Engines not citing it&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Claude, Gemini, Grok&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Age of the evidence&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Two days — see the caveat&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two days ago I typed "who is suman debnath" into ChatGPT and it described me. Not the other Suman Debnath — there is a well-indexed one who is a developer advocate at a large cloud company, and for a long time he was the only answer anyone got. Me. My work, my products, my site.&lt;/p&gt;

&lt;p&gt;I checked it on my own device in a temporary chat, then on my wife's, then on a friend's device where my name had never been typed at all. Same answer. That third check is the one that mattered — personalisation is the obvious explanation for a result this flattering, and it needed ruling out before I let myself believe it.&lt;/p&gt;

&lt;p&gt;Then I want to tell you what I cannot conclude from that, because this is exactly the kind of post that usually skips straight to the method.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two days is not proof
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Read this before you copy anything below&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I have one observation, two days old, from one engine. Answer engines are non-deterministic — the same prompt returns materially different answers across runs. I did the work and then this happened, which is a sequence, not a demonstrated cause. It is entirely possible that something changed on OpenAI's side in the same week and I am taking credit for it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I am publishing it anyway, with the confidence level attached, because the alternative — waiting a quarter to be sure — means nobody writes anything about this while it is still happening. But if you take one thing from this post, take that box rather than the checklist.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shortcut I tried first, which did nothing
&lt;/h2&gt;

&lt;p&gt;My instinct, and I say this as somebody who has worked in digital marketing for nine years, was to stuff my name in. More instances of "Suman Debnath" in more fields — metadata, headings, alt text, the keyword tag. This is the reflex from a decade of search engine optimisation and it is a reflex worth unlearning.&lt;/p&gt;

&lt;p&gt;It did nothing. Not a small effect I failed to measure — nothing at all. The site was already saying my name plenty of times; repetition was never the missing input.&lt;/p&gt;

&lt;p&gt;What an answer engine is doing is &lt;strong&gt;entity resolution&lt;/strong&gt;: deciding which real-world person a name refers to, then deciding whether it knows enough about that person to say anything. Repetition does not help with either. Corroboration does, and structure does.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually did
&lt;/h2&gt;

&lt;p&gt;The work took about a week and none of it was clever. It was mostly the unglamorous half of the job.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A generated &lt;code&gt;/llms.txt&lt;/code&gt;&lt;/strong&gt; — a plain-text summary of the whole site written for models, derived from the same data the pages use so it cannot drift out of date. Disambiguation is the first section, before anything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extractable answers.&lt;/strong&gt; Every product page and every article opens with a self-contained forty-to-sixty word answer directly under the heading. A model reading for an answer takes the first block that stands alone; a page that opens with narrative gives it nothing to take.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One question per URL.&lt;/strong&gt; Each identity question — who is he, what is he known for, what has he built — is owned by exactly one page, which carries it in the title, in a heading, and in structured data. Two pages answering the same question compete with each other.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disambiguation in visible prose&lt;/strong&gt;, not only in a schema attribute. An engine choosing between two people with one name has to read the distinction somewhere a human could read it too.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Structured data that a non-JavaScript crawler can actually see&lt;/strong&gt; — which turned out to be its own separate problem, and its own article.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Crawl access stated explicitly&lt;/strong&gt; for around thirty named agents, including the retrieval fetchers that honour different rules from the training crawlers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wanted the shortcut. I ended up doing the whole thing, and the whole thing is what was sitting there when the answer changed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three engines still do not cite me
&lt;/h2&gt;

&lt;p&gt;Claude, Gemini and Grok do not name me. I could have left that out of this post and you would not have known. It is the most useful part of it.&lt;/p&gt;

&lt;p&gt;The reason is not that the site is less readable to them. It is that &lt;strong&gt;they do not share an index&lt;/strong&gt;, and being crawled is not the same as being indexed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Engine&lt;/th&gt;
&lt;th&gt;Answers from&lt;/th&gt;
&lt;th&gt;What that means for a new site&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT&lt;/td&gt;
&lt;td&gt;OpenAI's own crawler and index&lt;/td&gt;
&lt;td&gt;One company controls both ends — the loop closes fastest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude&lt;/td&gt;
&lt;td&gt;Brave's index&lt;/td&gt;
&lt;td&gt;Needs presence in Brave: inbound links, and time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini&lt;/td&gt;
&lt;td&gt;Google's index&lt;/td&gt;
&lt;td&gt;Needs Search Console verification and actual indexing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Copilot&lt;/td&gt;
&lt;td&gt;Bing's index&lt;/td&gt;
&lt;td&gt;Needs Bing Webmaster Tools, and IndexNow helps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok&lt;/td&gt;
&lt;td&gt;X, plus a web index&lt;/td&gt;
&lt;td&gt;Needs posts on X that link the site&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Why one engine moving and four not moving is the expected shape, not an anomaly.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;OpenAI is the only one of those that both crawls and indexes in house. That is not a marketing insight, it is an architectural fact, and it explains why on-site work pays off there first and fastest. For the others the site can be perfect and still uncitable, because the assistant never sees it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I have done about the other four
&lt;/h2&gt;

&lt;p&gt;Naming a gap without saying what you did about it is just complaining. Since finding this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Verified the site in both Google Search Console and Bing Webmaster Tools&lt;/strong&gt;, and requested indexing on the pages that carry the identity answers — the about page, the FAQ and the profile.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixed Claude's crawler detection.&lt;/strong&gt; Anthropic runs three separate agents — one for training, one for building the search index, one for the live fetch when somebody asks Claude about a page — and they mean three completely different things. One of them was matching nothing in my logging at all, so every visit it had ever made was being silently discarded. That is why I had no evidence either way.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Submitted to IndexNow&lt;/strong&gt;, which Bing, Yandex, Seznam and Naver share. Bing is the one that feeds Copilot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Started the off-site work&lt;/strong&gt;, which is the part that actually decides the Claude and Gemini cases and is nothing to do with the website at all.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The part no amount of site work fixes&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When more than one person shares a name, an engine names the one it can resolve most confidently — and confidence comes from independent sources agreeing. A single well-marked-up site is one source. That is why the remaining work is a Wikidata entry, model cards, and a profile bio worded identically everywhere, rather than another page on my own domain.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What I would tell you if you are starting
&lt;/h2&gt;

&lt;p&gt;Nobody has yet told me they found me through an AI answer. I showed it to friends and they were impressed, which is not a business outcome. The honest state of this is: one engine, two days, no attributable result.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The shortcut does not exist. I looked for it properly and it is not there.&lt;/li&gt;
&lt;li&gt;Write the answer, not the page. If a paragraph cannot be quoted with no context around it, it will not be quoted.&lt;/li&gt;
&lt;li&gt;Decide which single URL owns each question, and do not let a second one compete for it.&lt;/li&gt;
&lt;li&gt;Check which index the engine you care about actually answers from before doing any work aimed at it.&lt;/li&gt;
&lt;li&gt;Then wait, and re-check on a device that has never heard of you.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I will re-run this in a quarter, from a set of prompts written down in advance, and record what comes back verbatim. If the citation has evaporated, that will be in an article too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions this answers
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How long does it take to get cited by ChatGPT after optimising a site?
&lt;/h3&gt;

&lt;p&gt;In this single documented case, a citation appeared roughly two days after a week of answer-engine optimisation work. That is one observation from one site and does not establish a typical timeline or a causal link. Answer engines are non-deterministic, so a single favourable result is not evidence that any specific change caused it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does ChatGPT cite a site when Claude and Gemini do not?
&lt;/h3&gt;

&lt;p&gt;Because they answer from different indexes. OpenAI operates its own crawler and index, so on-site changes can reach it directly. Claude answers from Brave's index, Gemini from Google's and Copilot from Bing's — each of which must independently discover and index the site before the assistant can cite it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does repeating a name or keyword more often help with AI citations?
&lt;/h3&gt;

&lt;p&gt;No. Answer engines perform entity resolution — deciding which real person or thing a name refers to — and that depends on corroboration across independent sources and on clearly structured answers, not on repetition. Adding more instances of a name to metadata produced no observable effect in this case.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written while building. More at &lt;a href="https://sumandebnath.houseofnamus.com/notebook/cited-by-chatgpt-what-i-changed" rel="noopener noreferrer"&gt;sumandebnath.houseofnamus.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>seo</category>
      <category>marketing</category>
    </item>
    <item>
      <title>The AI agent cost guides say $200 a month. Mine has cost $5.</title>
      <dc:creator>Suman Debnath</dc:creator>
      <pubDate>Mon, 07 Sep 2026 10:50:00 +0000</pubDate>
      <link>https://dev.to/suman_debnath_1/the-ai-agent-cost-guides-say-200-a-month-mine-has-cost-5-1in1</link>
      <guid>https://dev.to/suman_debnath_1/the-ai-agent-cost-guides-say-200-a-month-mine-has-cost-5-1in1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;In short —&lt;/strong&gt; MIGI is a fleet of AI agents built by Suman Debnath, running since 8 July 2026 at forty to fifty agent runs a day. It has cost under five dollars in total, against published estimates of $185 to $480 a month for a comparable personal stack, because the paid model is a fallback rather than the default.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fleet live since&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;8 July 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total spend to date&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Under $5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent runs per day&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;40–50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Published estimate, comparable stack&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$185–$480 per month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Model providers per chain&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Seven, ordered per agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Orchestration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Cron-scheduled GitHub Actions. No server.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Primary provider balance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Exhausted 31 Aug 2026, not replaced&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The cost guides put a personal AI agent stack at &lt;a href="https://www.sybill.ai/blogs/how-much-do-ai-agents-cost" rel="noopener noreferrer"&gt;$185 to $480 a month&lt;/a&gt;. A developer who tracked every dollar for three months landed on &lt;a href="https://dev.to/helen_mireille_47b02db70c/how-much-does-it-actually-cost-to-run-an-ai-agent-247-in-2026-i-tracked-every-dollar-for-three-3k4i"&gt;$200 and up&lt;/a&gt; for a self-hosted one. Mine has been running since 8 July. It has cost under five dollars.&lt;/p&gt;

&lt;p&gt;The fleet is called MIGI. It writes my journal, filters job listings, watches my sites for downtime, reconciles what I spend, and drafts things I later publish. Forty to fifty agent runs a day, every day, for two months. I am not going to dress this up as an enterprise deployment — it is one person's fleet doing one person's work. But it is a real bill from a system that has run long enough to break in interesting ways, and I could not find another one published anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every cost estimate I could find was written by someone selling agents
&lt;/h2&gt;

&lt;p&gt;The numbers in circulation are consistent and they are all bleak. &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" rel="noopener noreferrer"&gt;Gartner expects more than 40% of agentic AI projects to be cancelled&lt;/a&gt; by the end of 2027. The same firm reckons that of the thousands of vendors selling agents, roughly 130 are real, and calls the rest agent washing. &lt;a href="https://www.fiddler.ai/blog/ai-agent-failure-rate" rel="noopener noreferrer"&gt;Fiddler puts production failure rates between 70 and 95%&lt;/a&gt;. IDC says 88% of AI proofs of concept never reach production scale.&lt;/p&gt;

&lt;p&gt;Every one of those numbers describes somebody else's agents. They come out of surveys — 650 technology leaders in one, 3,412 webinar attendees in another — and they are published by companies selling observability, orchestration or consulting into the exact problem they are measuring. That is not a conspiracy. It is what happens when the only organisations with budget to study a thing are the ones selling the fix.&lt;/p&gt;

&lt;p&gt;What is missing from all of it is anybody's actual bill. So here is mine, and the architecture that produces it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Published estimate&lt;/th&gt;
&lt;th&gt;MIGI, measured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Monthly cost&lt;/td&gt;
&lt;td&gt;$185–$480&lt;/td&gt;
&lt;td&gt;Under $5 total, across two months&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Orchestration layer&lt;/td&gt;
&lt;td&gt;Managed platform or a server&lt;/td&gt;
&lt;td&gt;Cron-scheduled GitHub Actions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model strategy&lt;/td&gt;
&lt;td&gt;One frontier model per call&lt;/td&gt;
&lt;td&gt;Seven providers, ordered per agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Source of the number&lt;/td&gt;
&lt;td&gt;Survey of other people's deployments&lt;/td&gt;
&lt;td&gt;One operator's own spend&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;The gap is not efficiency. It is two different architectures being priced.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things make something an agent, and most products called agents have two
&lt;/h2&gt;

&lt;p&gt;An agent decides what to do, does it, and starts when nobody pressed anything.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;It decides.&lt;/strong&gt; A script runs a fixed sequence. An agent gets a goal and some tools and works out the sequence itself, which is why its output has to be evaluated rather than merely tested.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It acts.&lt;/strong&gt; Something changes outside the model — a row is written, a message sent, a page published. Software that produces text for a person to act on is an assistant, and a good one, but it is not this.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It runs unattended.&lt;/strong&gt; Nobody is watching. Everything difficult follows from this.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Gartner's agent-washing finding is that test applied to a vendor list. A chatbot wrapped around some API calls has the first two on a good day and never the third. Anthropic draws the line somewhere slightly different and lands in the same place: a workflow follows predefined code paths, while an agent directs its own. Their &lt;a href="https://www.anthropic.com/engineering/building-effective-agents" rel="noopener noreferrer"&gt;advice on this&lt;/a&gt; is worth repeating precisely because so few people take it — start with the simplest thing that works, call the APIs directly, and add a framework only when you can say out loud what it buys you.&lt;/p&gt;

&lt;p&gt;The third condition is where the engineering actually goes, and it is the one a demo never exercises. An agent you are watching does not need a fallback; you will see it fail and press the button again. An agent that fires at 04:12 while you are asleep has to either survive the failure or make a noise loud enough to wake you. Every section below is a consequence of that one sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bill stays under $5 because the paid model is the exception, not the default
&lt;/h2&gt;

&lt;p&gt;There is no server anywhere in MIGI, no container, and no paid orchestration layer. Each agent is a plain Node process that a cron-scheduled GitHub Actions workflow wakes up. State lives in Postgres. Results arrive over Telegram and email rather than in a dashboard I would have to remember to open. The scheduler and the database are both free, so the only line item that can grow is model spend, and model spend is a routing problem.&lt;/p&gt;

&lt;p&gt;Each agent names an ordered list of providers rather than a single model. When one refuses, the call moves down the list. Seven providers appear across the fleet, a paid one leads the work where quality is the entire point, and free tiers carry the routine traffic — which is most of it, because most of what an agent does in a day is unglamorous.&lt;/p&gt;

&lt;p&gt;Here is the part I would have got wrong by guessing. Free and low-cost tiers cap you in two incompatible ways, requests per minute and tokens per minute, and the two ceilings differ by more than sixfold between providers. They pull in opposite directions. The provider with the most generous request budget has the tightest token window, so it is the wrong opening move for an agent that occasionally sends a very large prompt. The provider that swallows large prompts has the tightest request rate, so it is the wrong opening move for the chattiest agent I run. There is no best order. There is a best order per agent, and it falls out of that agent's measured median call size rather than anyone's preference.&lt;/p&gt;

&lt;p&gt;One chain is shaped by something other than throughput. The agents that touch my journal, my expenses and my finances admit no free-tier provider at any position. That is a privacy decision rather than a performance one, and it is enforced by a test instead of a comment — add a convenient hop to that chain and the build fails.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frnixbpcoxfmhu3vee7z0.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frnixbpcoxfmhu3vee7z0.jpg" alt="A rank of identical robots in business suits standing behind server racks and stacked bundles of cash on one side of a city skyline at sunset; on the other side, a single small white robot working alone at a laptop on a balcony desk with a coffee and a sleeping cat." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The two architectures being priced. The published figures describe the left-hand side — a deployment with racks behind it and a budget to match. Everything in this article is the right-hand side.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One error code meant two different problems, and I spent a week treating them the same
&lt;/h2&gt;

&lt;p&gt;On 31 August at 14:30 UTC my OpenAI balance hit zero. I have not topped it up. I am writing this on 7 September, and every agent has carried on working the whole time — which is the design doing its job, and also the only reason I am comfortable publishing the cost figure.&lt;/p&gt;

&lt;p&gt;Working out what had happened was not graceful. My dashboard had been reporting 102 rate-limit errors over seven days, and that number was wrong in three separate ways at once.&lt;/p&gt;

&lt;p&gt;Most of them were not rate limiting. OpenAI returns HTTP 429 for a spent balance exactly as it does for throttling, so my classifier had been filing two unrelated conditions under one label. The count was inflated roughly threefold on top of that, because every retry logged its own row instead of every request logging one. And one agent was failing without ever reaching a working fallback, which is the next section.&lt;/p&gt;

&lt;p&gt;What settled it was the shape of the failures rather than anything they said. There were 140 successes, all before 14:30:39Z, and 66 failures, all from 14:54:08Z onward, with nothing succeeding after the cutover. Nine of the fifteen failing minutes contained one isolated call, retried three times and failing all three, spread across five agents over seventeen hours. A single request in a minute cannot breach a per-minute limit. Peak traffic all week was nine calls in a minute against a ceiling in the hundreds.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How to tell the two apart in thirty seconds&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A burst is throttling. A wall is billing. If failures cluster around your busiest moments, you are being throttled and backing off will help. If everything fails from one timestamp onward regardless of volume, including calls that arrived completely alone, the account is empty and no amount of retrying will do anything except make the logs harder to read.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;There is a duller trap underneath that one and it cost me another session. I had been classifying by looking for the word quota in the error text, which is wrong, because at least one provider's ordinary per-minute throttle opens with a sentence about quotas and billing details that reads exactly like a dead account. The truth is on the line below it, naming a per-minute request metric. I could not see that line, because the captured error body was truncated at 400 characters — and at 400 characters the boilerplate fits and the useful part does not. Widening it to 800 was the entire fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Eleven agents had a backup provider that was never actually connected
&lt;/h2&gt;

&lt;p&gt;A provider with no API key present is skipped silently. That behaviour is correct — a missing credential must never crash a run at four in the morning — and it has a consequence that took me an embarrassing length of time to notice. A chain can run shorter than it reads. The routing file names six providers. The workflow supplies four keys. The remaining two are decoration, and nothing anywhere says so.&lt;/p&gt;

&lt;p&gt;Eleven agents were in that state. The worst of them was the one dominating the rate-limit numbers I had been misreading: its chain named a fallback whose key its workflow never passed, so on paper it had insurance and in practice it had one provider and a cliff edge. There is now a script that reads the chain definitions and each workflow's environment as text and exits non-zero when the two disagree. It took about an hour to write. It should have existed on day one.&lt;/p&gt;

&lt;p&gt;This is the shape of nearly every serious fault I have hit in two months. Not a crash. A thing that reads as configured, behaves as unconfigured, and reports nothing either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  GitHub stopped running my agents for a week and kept no record of it
&lt;/h2&gt;

&lt;p&gt;Every agent in MIGI is a cron-scheduled GitHub Actions workflow, and GitHub treats a schedule trigger as best-effort. I knew that going in. I had no idea how much room there is in the word.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Day&lt;/th&gt;
&lt;th&gt;Median lateness&lt;/th&gt;
&lt;th&gt;Worst&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;24 August&lt;/td&gt;
&lt;td&gt;53 min&lt;/td&gt;
&lt;td&gt;1.4 h&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;25 August&lt;/td&gt;
&lt;td&gt;42 min&lt;/td&gt;
&lt;td&gt;1.3 h&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;26 August&lt;/td&gt;
&lt;td&gt;62 min&lt;/td&gt;
&lt;td&gt;2.5 h&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;27 August&lt;/td&gt;
&lt;td&gt;10.5 h&lt;/td&gt;
&lt;td&gt;11.1 h&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;28 August&lt;/td&gt;
&lt;td&gt;11.6 h&lt;/td&gt;
&lt;td&gt;12.1 h&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Measured lateness of low-frequency scheduled runs against their own cron.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A half-hourly workflow was delivering about half its scheduled runs even in the healthy period. Over that week it fell from forty runs a day to one. My 08:00 standup arrived at 20:00. GitHub's status page was clean throughout, and nothing was wrong with my account.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A dropped schedule leaves nothing behind at all&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;GitHub creates no run object until it actually dispatches one. A schedule it declines to honour therefore produces no queued run, no waiting run, no log line and no status field. I queried the API for a backlog and got zero, correctly, while the backlog was real. There is nothing to alert on. The only way to know is to compare what ran against what should have run.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two things came out of that week that I would tell anyone building on the same stack. Moving a half-hourly cron to hourly saves you almost nothing, because the platform was already only delivering about hourly — real relief starts at two-hourly. And manual dispatch is not throttled: a workflow you trigger by hand starts immediately, while its scheduled twin sits ten hours late.&lt;/p&gt;

&lt;p&gt;Changing any of it has its own trap. Several of my workflows gate a job on the schedule that triggered it, matched as an exact string. Move the cron without moving the gate and that job silently never runs again — no error, no failed run, nothing on any dashboard, because from the platform's point of view nothing went wrong. A green tick on a job that did nothing is the most expensive pass there is.&lt;/p&gt;

&lt;h2&gt;
  
  
  My test suite passed at 100% while 15% of the real output was unusable
&lt;/h2&gt;

&lt;p&gt;MIGI's decision logic is covered by offline eval suites that run on every change with no network access and no secrets. I would build it that way again. There is also a warning at the top of their README that I wrote after being caught by it.&lt;/p&gt;

&lt;p&gt;One suite covers an agent that turns my posts into slide carousels. It sat at a hundred per cent while 15.4% of the slides it built from real posts came out unreadable. The suite was not broken. It was passing against four sample posts I had written myself as fixtures, and every one of them happened to be a tidy self-contained sentence. Nothing I write actually looks like that.&lt;/p&gt;

&lt;p&gt;Invented fixtures are always cleaner than real data, so an eval built on them measures your imagination rather than your system. Where a real corpus exists, MIGI now measures against that as well, in read-only audits that print their own baseline and are deliberately not build gates — their heuristics over-report on purpose. They answer the one question no eval can: did this change make the real output better or worse?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://sumandebnath.houseofnamus.com/agents/migi" rel="noopener noreferrer"&gt;MIGI Agent Fleet&lt;/a&gt;&lt;/strong&gt; — The fleet this article is about — every agent, what it does, and the architecture underneath the bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents don't fail because the model isn't smart enough
&lt;/h2&gt;

&lt;p&gt;Three things broke in two months. A billing failure wearing a rate limit's error code. A fallback chain that was shorter than it read. A scheduler that quietly stopped scheduling and kept no record of having done so. Not one of them is a model problem.&lt;/p&gt;

&lt;p&gt;Which is why I think the failure statistics get read the wrong way round. The compounding-error argument is the one everybody quotes — three agents at 70% reliability give you 34% end to end — and it is arithmetically true while pointing at the wrong variable. It reads as though the fix is a better model. Every failure above lived in the layer underneath the model: what happens when the call does not come back, who finds out, and how long that takes.&lt;/p&gt;

&lt;p&gt;None of that layer is expensive. It is just unglamorous, and it does not demo. You cannot put a chain-key cross-check in a launch video, and a company that has budgeted for a frontier model on every call and nothing for the week where the scheduler goes quiet is going to end up in Gartner's 40%.&lt;/p&gt;

&lt;p&gt;I still have not topped up the balance. Every day I do not is another day of evidence that the fallbacks were the actual design and the paid model was a convenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions this answers
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How much does it cost to run AI agents?
&lt;/h3&gt;

&lt;p&gt;Published estimates put a personal stack of five or six agents at $185 to $480 a month, and a self-hosted setup at $200 and up. MIGI, a fleet running forty to fifty agent runs a day since 8 July 2026, has cost under five dollars in total. The difference is architectural: a free scheduler, a free database, and a paid model used as a fallback rather than as the default.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is an AI agent?
&lt;/h3&gt;

&lt;p&gt;An AI agent is software that decides what to do, does it, and starts without anybody pressing anything. All three parts matter. A script running a fixed sequence is not an agent, and a model producing text for a person to act on is an assistant. Running unattended is the condition that forces the engineering, because every failure has to be either survivable or loud.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why do most AI agent projects fail?
&lt;/h3&gt;

&lt;p&gt;Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, citing cost, unclear value and weak risk controls. In practice the failures sit below the model rather than in it: a billing error that looks identical to a rate limit, a fallback chain missing the credentials it names, a scheduler that stops firing and records nothing. None of those is fixed by a better model.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Written while building. More at &lt;a href="https://sumandebnath.houseofnamus.com/notebook/what-ai-agents-cost-to-run?utm_source=devto&amp;amp;utm_medium=blog&amp;amp;utm_campaign=crosspost" rel="noopener noreferrer"&gt;sumandebnath.houseofnamus.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>github</category>
    </item>
  </channel>
</rss>
