<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Michael Hairetis</title>
    <description>The latest articles on DEV Community by Michael Hairetis (@michaelhairetis).</description>
    <link>https://dev.to/michaelhairetis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4117615%2Ffb885c49-b01e-4edc-81fd-c62c8ae87c11.jpeg</url>
      <title>DEV Community: Michael Hairetis</title>
      <link>https://dev.to/michaelhairetis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/michaelhairetis"/>
    <language>en</language>
    <item>
      <title>I spent a week building an optimization, A/B tested it, and deleted it</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Mon, 28 Sep 2026 11:57:00 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/i-spent-a-week-building-an-optimization-ab-tested-it-and-deleted-it-573</link>
      <guid>https://dev.to/michaelhairetis/i-spent-a-week-building-an-optimization-ab-tested-it-and-deleted-it-573</guid>
      <description>&lt;p&gt;I had a 25,000-token system-prompt overhead on every agent call. Obvious problem, obvious solution: build a session-reuse layer so the overhead is paid once and amortised across many calls.&lt;/p&gt;

&lt;p&gt;So I built it. Session pooling, lifecycle management, identity tracking, the works. It was the most intricate thing in that part of the codebase and I was quietly pleased with it.&lt;/p&gt;

&lt;p&gt;Then I A/B tested it against not having it.&lt;/p&gt;

&lt;p&gt;No difference.&lt;/p&gt;

&lt;p&gt;The provider's prompt cache hits on prefix content regardless of session identity. Two completely independent calls inside the cache window already got the benefit. The optimisation I had carefully engineered was buying something I already had for free, and had been getting for free the entire time I was building the thing to get it.&lt;/p&gt;

&lt;p&gt;I deleted the whole layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The lesson is not "measure before optimising."&lt;/strong&gt; Everyone says that and everyone nods and nobody does it, because the optimisation is obviously going to help.&lt;/p&gt;

&lt;p&gt;The sharper version is this: &lt;strong&gt;when your semantics are stateless, write stateless code.&lt;/strong&gt; I had reached for a stateful design to solve a problem the platform had already solved statelessly. The complexity was not buying performance, it was buying a mental model that did not match the system underneath.&lt;/p&gt;

&lt;p&gt;Three other things from the same project, all in the same shape:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Structure beats trust.&lt;/strong&gt; A single agent given a complex task fails unpredictably. It loses subgoals and declares victory early. Rather than trying to make one agent reliable, the pipeline does the managing and catches failures at phase boundaries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is a routing problem.&lt;/strong&gt; I had one heavyweight backend answering both "which of these three pipelines" and "go do two minutes of multi-step research." Making a heavyweight agent answer a classification question costs heavyweight money for a featherweight answer. Splitting them made both paths faster and cheaper.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Old heuristics still earn their keep.&lt;/strong&gt; Asking a model "is this task finished" costs about five cents a check. Pattern matching on the agent's own activity signals, with the model as fallback for genuinely ambiguous cases, cut average per-request cost roughly tenfold.&lt;/p&gt;

&lt;p&gt;None of these were model problems. None were solved by a better prompt.&lt;/p&gt;

&lt;p&gt;The full write-up is on my site, including why I think the infrastructure half is the half that actually decides whether the thing works in production.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/the-boring-parts-of-agentic-ai.html" rel="noopener noreferrer"&gt;https://openred.space/blog/the-boring-parts-of-agentic-ai.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>programming</category>
      <category>performance</category>
    </item>
    <item>
      <title>98.7% better on seasonal, 9.8% worse on a random walk: the most useful benchmark I ran</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Sat, 26 Sep 2026 12:29:59 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/987-better-on-seasonal-98-worse-on-a-random-walk-the-most-useful-benchmark-i-ran-nep</link>
      <guid>https://dev.to/michaelhairetis/987-better-on-seasonal-98-worse-on-a-random-walk-the-most-useful-benchmark-i-ran-nep</guid>
      <description>&lt;p&gt;Before pointing Google's TimesFM 3.0 at anything real, I ran a calibration baseline. Two synthetic series with known properties, same model, same settings, held out.&lt;/p&gt;

&lt;p&gt;Seasonal with drift: &lt;strong&gt;98.7% better&lt;/strong&gt; than the naive baseline, &lt;strong&gt;100%&lt;/strong&gt; direction accuracy.&lt;br&gt;
Random walk: &lt;strong&gt;9.8% worse&lt;/strong&gt; than naive, 47% direction.&lt;/p&gt;

&lt;p&gt;The naive baseline is the crudest forecast available: next value equals last value.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Row two is the important one, and it is not a defect.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A random walk is unpredictable by construction. Each step is independent of every step before it. A model that appeared to forecast one would be reporting structure that does not exist. In a backtest that looks like skill. In production it looks like losing money.&lt;/p&gt;

&lt;p&gt;The correct behaviour on an unpredictable series is to fail, ideally about as badly as naive. That is what happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read together the two rows are a map.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Where it works: demand forecasting, energy load, web traffic, call volume, inventory, capacity planning, sensor telemetry. Domains where next week resembles last week plus a trend.&lt;/p&gt;

&lt;p&gt;The commercial argument there is not really accuracy, it is that you do not fit a model per series. Ten thousand SKUs, ten thousand zero-shot forecasts, about 8 ms each.&lt;/p&gt;

&lt;p&gt;Where it does not: anything closer to a random walk than to a seasonal series.&lt;/p&gt;

&lt;p&gt;I then spent two days confirming that market tape sits firmly on the second side of that line.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.hashnode.dev/a-forecasting-model-that-correctly-refuses-to-forecast?utm_source=hashnode&amp;amp;utm_medium=feed" rel="noopener noreferrer"&gt;https://openred.hashnode.dev/a-forecasting-model-that-correctly-refuses-to-forecast?utm_source=hashnode&amp;amp;utm_medium=feed&lt;/a&gt;&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>datascience</category>
      <category>python</category>
      <category>ai</category>
    </item>
    <item>
      <title>Your menu needs an API, and the hard part is not the payment</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Fri, 25 Sep 2026 22:39:00 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/your-menu-needs-an-api-and-the-hard-part-is-not-the-payment-b11</link>
      <guid>https://dev.to/michaelhairetis/your-menu-needs-an-api-and-the-hard-part-is-not-the-payment-b11</guid>
      <description>&lt;p&gt;Direct ordering has existed for years and most independent restaurants still send most of their off-premise volume through marketplaces at 15% to 30% commission. There is a good reason, and agents change it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The objection that mattered:&lt;/strong&gt; direct ordering made you responsible for your own demand. The marketplace was a place customers already were, with a recommendation engine and a promotions budget. Going direct meant keeping 15 to 30 points and inheriting a marketing job you had no staff for. A high-margin channel nobody uses loses to a low-margin channel that is full.&lt;/p&gt;

&lt;p&gt;When a customer tells an assistant to order from a named restaurant, the demand routing is already done. The agent is looking for an endpoint that can take the order. If you have one it goes there. If not it falls back to whoever does, and you pay commission on demand you had already earned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direct ordering stops being a marketing problem and becomes an integration problem.&lt;/strong&gt; Integration problems are tractable.&lt;/p&gt;

&lt;p&gt;The build list, honestly:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A catalogue as structured data. Items, modifiers, prices, tax treatment, real availability, prep time. The single most common blocker, because most independents do not have an accurate machine-readable menu anywhere, including inside their own POS.&lt;/li&gt;
&lt;li&gt;A quote endpoint. Priced basket including tax and fees, with a short TTL. Not the same as publishing a menu, and where naive integrations break.&lt;/li&gt;
&lt;li&gt;Payment acceptance. Two live standards on HTTP 402: x402 (Coinbase, stablecoin-native, most volume) and MPP (Stripe and Tempo, also routes cards via Shared Payment Tokens). You do not get to pick which one the customer's agent speaks, so sit behind a processor that handles both.&lt;/li&gt;
&lt;li&gt;Order injection into the POS. Toast, Square and Clover all have APIs and all differ. This is where the hours go, and why "add an ordering page" is not the same project.&lt;/li&gt;
&lt;li&gt;A status channel back to the agent.&lt;/li&gt;
&lt;li&gt;Discoverability in whatever registry agents consult.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;The part nobody has built:&lt;/strong&gt; exceptions. Out of stock, substitution approval, partial refund on a payment that already settled. The payment protocols solve payment and say nothing about the commercial relationship around a failed order. Ask anyone selling you agent-readiness how they handle it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/your-menu-needs-an-api.html" rel="noopener noreferrer"&gt;https://openred.space/blog/your-menu-needs-an-api.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>fintech</category>
      <category>api</category>
      <category>architecture</category>
    </item>
    <item>
      <title>I built the classifier tier in April. A model just shipped to be it.</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Thu, 24 Sep 2026 16:06:22 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/i-built-the-classifier-tier-in-april-a-model-just-shipped-to-be-it-38e9</link>
      <guid>https://dev.to/michaelhairetis/i-built-the-classifier-tier-in-april-a-model-just-shipped-to-be-it-38e9</guid>
      <description>&lt;p&gt;TypeSafe AI came out of stealth with Jev, a model that does not generate text. You hand it state and a typed question, it returns a calibrated probability over a fixed answer set. Nothing to parse, no invalid answer possible.&lt;/p&gt;

&lt;p&gt;My reaction was not "clever idea." It was "I have had that in production since April."&lt;/p&gt;

&lt;p&gt;That reaction is partly wrong and worth unpacking.&lt;/p&gt;

&lt;p&gt;What I will not claim: using a model as a classifier is not novel and was not novel in April. It has been ordinary since function calling shipped, and calibrated-label training sits in a research lineage older than any of this. Arriving somewhere independently is not arriving first. And Jev is a different category from what I built: I implemented a pattern with a general-purpose agent, they trained a model whose entire job is that pattern.&lt;/p&gt;

&lt;p&gt;Three things their approach has that mine structurally cannot: probabilities instead of labels, many questions about one state in parallel for nearly nothing, and latency appropriate to something sitting in the hot path.&lt;/p&gt;

&lt;p&gt;What I will defend is the integration shape, and it is not about who was first.&lt;/p&gt;

&lt;p&gt;Both patterns in their launch writeup lead with the model. Route every request through the classifier. Send every tool call through the classifier. The classifier is the front door.&lt;/p&gt;

&lt;p&gt;Mine has three tiers and the classifier is the middle one:&lt;/p&gt;

&lt;p&gt;source: "mechanical" | "classifier" | "fallback"&lt;/p&gt;

&lt;p&gt;Mechanical runs first and settles genuinely unambiguous cases with no model call. A large share of real traffic is not ambiguous; someone typing an exact pipeline name needs a lookup, not a judgment call. The classifier handles only the residue. A dumb fallback sits underneath so unavailability degrades predictably.&lt;/p&gt;

&lt;p&gt;The reason is not elegance. The cheapest call is the one you do not make, and the second cheapest is the one whose answer you can predict without asking. Leading with the model means paying for judgment on inputs that required none.&lt;/p&gt;

&lt;p&gt;A second benefit I did not appreciate at first: attributing which tier answered. "The classifier decided this" and "a pattern matched this" are different bugs with different fixes. One undifferentiated tier gives you one undifferentiated failure.&lt;/p&gt;

&lt;p&gt;The genuinely new thing in Jev is not asking a model for a typed answer. It is that the classifier finally has purpose-built machinery behind it. What I would push back on is putting it at the front.&lt;/p&gt;

&lt;p&gt;One caveat that cuts against adopting it, and worth answering before rather than after: my cost argument rests on no per-call metering, which is exactly what let me put a judgment call where a per-token bill would have discouraged one. A metered classifier reintroduces per-call billing into a system built to avoid it. Probably trivial at my volumes. Probably trivial is a thing to check.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/jev-and-the-classifier-tier.html" rel="noopener noreferrer"&gt;https://openred.space/blog/jev-and-the-classifier-tier.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>software</category>
      <category>llm</category>
    </item>
    <item>
      <title>Two standards are reviving HTTP 402, and it changes who owns the customer</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Wed, 23 Sep 2026 11:06:41 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/two-standards-are-reviving-http-402-and-it-changes-who-owns-the-customer-3njn</link>
      <guid>https://dev.to/michaelhairetis/two-standards-are-reviving-http-402-and-it-changes-who-owns-the-customer-3njn</guid>
      <description>&lt;p&gt;HTTP 402 has sat unused in the spec since the early web. Two competing standards are now bringing it back, and the reason matters more than the trivia.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;x402&lt;/strong&gt;, from Coinbase. An agent requests a resource, the server answers 402 with payment instructions, the agent signs a stablecoin transaction and retries with proof attached. No account, no login, no stored card. Most traffic settles in USDC on Base or Solana. Coinbase and Cloudflare moved it under a foundation with Circle, Stripe and AWS involved, and it is one of the first extensions to Google's Agent Payments Protocol. It currently carries the most volume of the two.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;MPP&lt;/strong&gt;, the Machine Payments Protocol, from Stripe and Tempo. Spec at mpp.dev, launched March 2026. Two things make it notable. It is not stablecoin-only: Stripe offramps agent stablecoin payments into an ordinary Stripe balance, and fiat and card payments run through it via Shared Payment Tokens, with Visa publishing card specs and an SDK. And it adds a sessions primitive, where an agent authorises a spending limit once and then streams many small payments without settling each one separately.&lt;/p&gt;

&lt;p&gt;So the framing is not "stablecoins replace cards." Machine payments need three properties card rails were never designed for: small amounts, high frequency, and authority scoped and revocable per purchase rather than per account. Stablecoin rails got there first because they could. The card networks are arriving through the same protocols rather than being displaced.&lt;/p&gt;

&lt;p&gt;The honest state: in March 2026 x402 was processing on the order of tens of thousands of dollars a day, much of it testing, against an ecosystem valued in the billions. The protocol works and the facilitators are live. Widespread use is years away, not a question of whether.&lt;/p&gt;

&lt;p&gt;I wrote up what this does to the businesses in the middle. Short version: a delivery app's moat was the interface, not the logistics, and a driver fleet is a hard operational job rather than a defensible asset.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/the-interface-was-the-moat.html" rel="noopener noreferrer"&gt;https://openred.space/blog/the-interface-was-the-moat.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>uxdesign</category>
      <category>webdev</category>
    </item>
    <item>
      <title>I pointed Google's time-series foundation model at the stock market</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Wed, 23 Sep 2026 10:55:56 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/i-pointed-googles-time-series-foundation-model-at-the-stock-market-2nih</link>
      <guid>https://dev.to/michaelhairetis/i-pointed-googles-time-series-foundation-model-at-the-stock-market-2nih</guid>
      <description>&lt;p&gt;I spent two days evaluating Google's TimesFM 3.0 against 118,857 bars of Nasdaq futures tape, strictly causal and session confined. Short version: it is a competent, very small, very cheap component, and it produced no trading edge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Direction accuracy on price: 50.0%.&lt;/strong&gt; Exactly fifty. On error magnitude it came in 7.4% worse than assuming the next value equals the last one.&lt;/p&gt;

&lt;p&gt;Four framings. Coarser bars did not help. A reframing built deliberately to suit its known strengths did not help. Feeding it side information through its supported covariate channel made forecasts worse at every horizon.&lt;/p&gt;

&lt;p&gt;The most interesting failure: I built a probability signal from its own quantile output. When it was most confident it was most wrong. 24.7% hit rate against a 34.3% base rate. &lt;strong&gt;Anti-informative when confident is worse than useless&lt;/strong&gt;, because useless does not tempt you to size up.&lt;/p&gt;

&lt;p&gt;The reason is architectural rather than bad luck. It is stateless. Every predict() call sees only the array you hand it, nothing carries between calls, and there is no text interface, so there is no channel through which to tell it anything. It cannot represent "this signal is void while that condition holds." That is a state machine and there is nowhere in a numeric array to put one.&lt;/p&gt;

&lt;p&gt;To be fair to it: on a series with stable repeating structure it beat naive by 98.7% zero-shot with 100% direction accuracy. It runs in about 1.3 GB of VRAM, roughly 8% of a mid-range consumer GPU, at about 8 ms per series batched. And on a random walk it did slightly worse than doing nothing, at 47% direction, which is correct behaviour rather than a defect: a model that appeared to forecast a random walk would be inventing structure.&lt;/p&gt;

&lt;p&gt;"Foundation" describes how a model was trained and how it transfers. It does not describe competence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://michaelhairetis.medium.com/a-foundation-model-is-not-a-foundation-a3edf3383be9" rel="noopener noreferrer"&gt;https://michaelhairetis.medium.com/a-foundation-model-is-not-a-foundation-a3edf3383be9&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>timesfm</category>
      <category>llm</category>
      <category>tooling</category>
    </item>
    <item>
      <title>I've had plenty of these so I figure I'd write about them because I'm all of y'all have had at least one by now</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Tue, 22 Sep 2026 15:54:30 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/ive-had-plenty-of-these-so-i-figure-id-write-about-them-because-im-all-of-yall-have-had-at-1hij</link>
      <guid>https://dev.to/michaelhairetis/ive-had-plenty-of-these-so-i-figure-id-write-about-them-because-im-all-of-yall-have-had-at-1hij</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/michaelhairetis/the-ai-hangover-43ep" class="crayons-story__hidden-navigation-link"&gt;The AI Hangover&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/michaelhairetis" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4117615%2Ffb885c49-b01e-4edc-81fd-c62c8ae87c11.jpeg" alt="michaelhairetis profile" class="crayons-avatar__image" width="400" height="400"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/michaelhairetis" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Michael Hairetis
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Michael Hairetis
                
                
              
              &lt;div id="story-author-preview-content-4717674" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/michaelhairetis" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4117615%2Ffb885c49-b01e-4edc-81fd-c62c8ae87c11.jpeg" class="crayons-avatar__image" alt="" width="400" height="400"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Michael Hairetis&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/michaelhairetis/the-ai-hangover-43ep" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 22&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/michaelhairetis/the-ai-hangover-43ep" id="article-link-4717674"&gt;
          The AI Hangover
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/hangover"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;hangover&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/software"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;software&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/programming"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;programming&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/michaelhairetis/the-ai-hangover-43ep#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            3 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>ai</category>
      <category>discuss</category>
    </item>
    <item>
      <title>The AI Hangover</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Tue, 22 Sep 2026 15:53:19 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/the-ai-hangover-43ep</link>
      <guid>https://dev.to/michaelhairetis/the-ai-hangover-43ep</guid>
      <description>&lt;p&gt;You did not do more work. You did a different kind of work, and it costs more per hour.&lt;/p&gt;

&lt;p&gt;There is a state I end some days in now that I did not have two years ago.&lt;/p&gt;

&lt;p&gt;It arrives after a long session with a coding agent, and usually not a bad one. The opposite: a session where a great deal got built, the kind of afternoon that used to be a fortnight. At the end of it I am not tired the way a long day of programming used to make me tired. I am flattened. Food does not touch it. Coffee does not touch it. A walk does not touch it. The only thing that resolves it is sleep, and not a nap either.&lt;/p&gt;

&lt;p&gt;I have started calling it an AI hangover, and once I named it I found that most people doing this kind of work recognise it immediately and had not said it out loud.&lt;/p&gt;

&lt;p&gt;To be clear about what this is: an observation, from one person, repeated enough times to be a pattern. I am not making a claim about neurochemistry and I have not measured anything. What I can do is describe it precisely and argue about the cause, because the obvious explanation is wrong and the wrong explanation leads you to manage it badly.&lt;/p&gt;

&lt;p&gt;The obvious explanation, and why it fails&lt;br&gt;
The obvious explanation is volume. You produced four times as much, so of course you are wrecked.&lt;/p&gt;

&lt;p&gt;It fails on its own terms. The tiredness does not scale with output. I have had sessions that shipped an enormous amount and left me fine, and sessions that produced one stubborn module and left me unable to form sentences at dinner. If volume were the driver those would be the other way round.&lt;/p&gt;

&lt;p&gt;It also fails on kind. Ordinary programming fatigue is a depletion. You feel like you have been running and you slow down gradually, and you can usually feel it coming an hour out. This is not that. This has a distinct quality of overload rather than depletion, it arrives late and abruptly, and the thing that is exhausted is specifically judgment. I can still lift things. I can still hold a conversation. What I cannot do is decide anything.&lt;/p&gt;

&lt;p&gt;That last detail is the tell, and it points at the actual cause.&lt;/p&gt;

&lt;p&gt;You changed jobs without noticing&lt;br&gt;
Here is what actually changed when you started working this way.&lt;/p&gt;

&lt;p&gt;You stopped being an author and became a reviewer.&lt;/p&gt;

&lt;p&gt;That sounds like a smaller change than it is, because both activities are called programming and happen at the same desk. They are not the same work, and they do not cost the same.&lt;/p&gt;

&lt;p&gt;When you write code yourself, most of the minutes are production. You are typing, and typing is slow, and while your hands are moving your mind is running ahead assembling the next piece. Decisions are real but they are punctuation. They sit between long stretches of execution, and those stretches are where you recover.&lt;/p&gt;

&lt;p&gt;When you drive an agent, the production minutes are gone. The agent has them. What is left for you is the part that was always the expensive part: read this, understand what it does, decide whether it is right, decide whether it is right for this codebase, decide whether the thing it did instead of what you asked is better than what you asked, and decide all of that fast enough that you are not the bottleneck.&lt;/p&gt;

&lt;p&gt;Every single output is a judgment call, and there is nothing in between them.&lt;/p&gt;

&lt;p&gt;You did not take on more work. You took on a stream of work that is nothing but the costly part, with the cheap part removed, and then you did it for six hours because it was going well.&lt;/p&gt;

&lt;p&gt;Why reviewing costs more than writing&lt;br&gt;
Two reasons, and both get worse the better the agent gets.&lt;/p&gt;

&lt;p&gt;You have to build a model of code you did not write. When you write a function, the mental model arrives for free as a byproduct: you cannot type it without understanding it. When you read a function, the model has to be constructed deliberately, and construction is work. Any developer who has reviewed a large pull request knows that forty minutes of review is worse than forty minutes of writing, and nobody finds that surprising. What is new is doing it all day, at speed, as the entire job.&lt;/p&gt;

&lt;p&gt;You cannot fully trust the output, so you never stop watching. Not because agents are bad, but because they are good enough that errors are plausible rather than obvious. A broken thing announces itself. A subtly wrong thing does not, and the only defence is sustained attention. Sustained attention at a constant level, with no natural breaks, is a well-known way to wear a person out, and it is the state you are in from the first prompt to the last.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>hangover</category>
      <category>software</category>
      <category>programming</category>
    </item>
    <item>
      <title>TimesFM 3.0 is not an LLM, and the word 'foundation' is doing too much work</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Mon, 21 Sep 2026 13:44:46 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/timesfm-30-is-not-an-llm-and-the-word-foundation-is-doing-too-much-work-1nlk</link>
      <guid>https://dev.to/michaelhairetis/timesfm-30-is-not-an-llm-and-the-word-foundation-is-doing-too-much-work-1nlk</guid>
      <description>&lt;p&gt;I evaluated Google's TimesFM 3.0 time-series foundation model on financial data. Before the results, the thing that surprised me most: my own expectations were wrong, and I do not think I was alone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I expected from "foundation model"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Memory between calls. Learned behaviour that accumulates. Some way to pass domain knowledge or conditional strategy. The affordances the term carries now that language models have claimed it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it actually is&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A decoder-only transformer, pre-trained on a broad corpus of time series, used zero-shot. The forecaster exposes three public methods: from_pretrained, predict, predict_batch.&lt;/p&gt;

&lt;p&gt;That is the contract. predict() sees a numeric array. It is stateless, so nothing carries between calls. There is no text interface, no instructions, no reasoning step.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why the distinction matters practically&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If your domain model is a set of conditional rules, "signal X is void while regime Y holds", you need somewhere to express conditions. A numeric array is not that place. Those rules are state machines, and they are cheap and exact to write directly in code.&lt;/p&gt;

&lt;p&gt;Technically it IS a foundation model: pre-trained once, applied zero-shot across domains with no per-series fitting. That is genuine, and the deployment numbers back it up. But "foundation" describes how it was trained and how it transfers, not what it can do.&lt;/p&gt;

&lt;p&gt;Full series, including why it produced no usable signal on market data:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://michaelhairetis.medium.com/i-expected-memory-i-got-a-function-call-40b714cac7b6" rel="noopener noreferrer"&gt;https://michaelhairetis.medium.com/i-expected-memory-i-got-a-function-call-40b714cac7b6&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>datascience</category>
      <category>llm</category>
    </item>
    <item>
      <title>You cannot make the agent fast enough. Move the wait instead.</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Fri, 18 Sep 2026 12:26:32 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/you-cannot-make-the-agent-fast-enough-move-the-wait-instead-44ap</link>
      <guid>https://dev.to/michaelhairetis/you-cannot-make-the-agent-fast-enough-move-the-wait-instead-44ap</guid>
      <description>&lt;p&gt;I built a small learning app for my two kids on the same agent infrastructure that runs two other platforms. It introduced a constraint the other two never had.&lt;/p&gt;

&lt;p&gt;Somebody is waiting.&lt;/p&gt;

&lt;p&gt;The orchestration platform runs at four in the morning with nobody watching. If a call takes ninety seconds, it takes ninety seconds, and nothing in the system notices. The publishing platform has a person in front of it, but that person is me, reviewing on my own schedule.&lt;/p&gt;

&lt;p&gt;The third has a second-grader holding a pencil, looking at a screen. Her patience is not a performance budget I get to negotiate. It is a hard physical limit, and when I exceed it she wanders off and the product has failed in the only way that matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;I built it the obvious way first, and the obvious way was wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Generate content one question at a time. Each call is small, so it is fast and unlikely to come back malformed, and I prefetch the next question while the child works on the current one. In theory the waiting hides behind the thinking.&lt;/p&gt;

&lt;p&gt;In practice a child does not fill the gap the way a scheduler does. They answer, and then they sit. If the prefetch had not landed they watched a spinner, and they did that between every single question. Six items in a set, six opportunities to lose them.&lt;/p&gt;

&lt;p&gt;The architecture that looked responsive on paper was the least responsive thing I could have built, because it distributed the waiting across exactly the moments when the user had nothing to do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The instinct is to make the calls faster. That instinct is wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You usually cannot make an agent fast enough to be invisible, so stop trying and change when it runs instead.&lt;/p&gt;

&lt;p&gt;The app now plans three lessons at once and builds each one complete in a single call, every question and figure and answer finished and stored before the child sees any of it. Lesson one takes a minute or two. They wait for that, once. While they work through it, lessons two and three build in the background. From there, opening a lesson is a database read. I measured it at six milliseconds.&lt;/p&gt;

&lt;p&gt;The waiting did not shrink. It moved. One wait at the front, where a user will tolerate it because nothing has started yet, in exchange for zero waiting during the part where they are actually engaged.&lt;/p&gt;

&lt;p&gt;Two details make it hold up. &lt;strong&gt;The lesson has to be genuinely complete&lt;/strong&gt;, not a plan with placeholders, or the scheme collapses back into the original problem. And &lt;strong&gt;the background build has to survive a server restart&lt;/strong&gt;, because in-flight async tasks die silently and leave a lesson stuck at "building" forever.&lt;/p&gt;

&lt;p&gt;The full piece covers the applets the agent writes as working code, the positional bias I found in generated quizzes, and why my reader cannot read.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/a-seven-year-old-will-not-wait.html" rel="noopener noreferrer"&gt;https://openred.space/blog/a-seven-year-old-will-not-wait.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>ux</category>
      <category>performance</category>
      <category>programming</category>
    </item>
    <item>
      <title>Does your system read back what it writes?</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Wed, 16 Sep 2026 13:00:20 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/does-your-system-read-back-what-it-writes-29nb</link>
      <guid>https://dev.to/michaelhairetis/does-your-system-read-back-what-it-writes-29nb</guid>
      <description>&lt;p&gt;I had agents writing notes that nothing ever read. For months. Here is how I found out, and why I think it is a common shape of bug rather than a one-off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The setup.&lt;/strong&gt; An orchestration platform runs scheduled data jobs. Worker agents were instructed to bank durable technique in a per-job notes file: parsing traps, cadence, known-good access patterns. They complied. Some of those files reached 40 KB.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The check that took thirty seconds.&lt;/strong&gt; I grepped the codebase for the filename.&lt;/p&gt;

&lt;p&gt;Two hits. An archiving routine, and a UI file-lister.&lt;/p&gt;

&lt;p&gt;It was never injected into a prompt. Never read at plan time, never read at execution time. &lt;strong&gt;Write-only memory.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The tell I had already seen and misread.&lt;/strong&gt; Only a fraction of jobs maintained a notes file at all. I had assumed inconsistent agent behaviour. It was rational behaviour: nobody keeps a notebook that never comes back to them, and that turns out to be true of software too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second half of the same bug.&lt;/strong&gt; The plan handed to each worker was discarded when the run ended. So a job that had run seventy times re-derived its plan from zero every morning. Schema, paths, identity rules, all rebuilt nightly from nothing.&lt;/p&gt;

&lt;p&gt;Measured cost: a clean run needs five agent turns. Successful runs averaged 6.5, worst case 10. Runs that concluded "nothing new today" averaged 3.5, burning turns after the answer was already known.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three constraints worth stealing if you wire this up.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Bounded, not whole.&lt;/em&gt; My largest notes file was nearly double the size at which my prompt transport starts corrupting input. Injecting it wholesale would have reproduced a bug I had already fixed. The file now has a curated head that travels with every run and a tail that stays on disk at a path the agent knows.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Curated, not appended.&lt;/em&gt; Append-only is how you get bloat. Replace facts that changed, delete the superseded version, and if nothing durable changed, leave it alone. That last clause prevents churn.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Instructions survive, context yields.&lt;/em&gt; When plan plus notes would exceed the ceiling, the notes get trimmed and the plan never does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And one prompt-level detail that mattered more than the architecture:&lt;/strong&gt; tell the planner explicitly that "nothing new today" still means the plan worked. Without that, it reads a quiet outcome as underperformance and rewrites a plan that was fine.&lt;/p&gt;

&lt;p&gt;Go grep for the filename your agents write to. It takes thirty seconds and I would bet on the result.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/agents-have-amnesia.html" rel="noopener noreferrer"&gt;https://openred.space/blog/agents-have-amnesia.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>softwareengineering</category>
      <category>debugging</category>
    </item>
    <item>
      <title>Every editorial bug I shipped was an agent following orders exactly</title>
      <dc:creator>Michael Hairetis</dc:creator>
      <pubDate>Tue, 15 Sep 2026 14:54:19 +0000</pubDate>
      <link>https://dev.to/michaelhairetis/every-editorial-bug-i-shipped-was-an-agent-following-orders-exactly-3p1</link>
      <guid>https://dev.to/michaelhairetis/every-editorial-bug-i-shipped-was-an-agent-following-orders-exactly-3p1</guid>
      <description>&lt;p&gt;I run two platforms on the same agent infrastructure. One makes internal decisions, the other publishes prose to actual readers. The second one taught me something the first could not.&lt;/p&gt;

&lt;p&gt;For about a month I kept finding defects in published text. Internal shorthand in a reader-facing card. Database row identifiers appearing in a published sentence, literally &lt;code&gt;events/1042&lt;/code&gt; in a paragraph a subscriber could read. Terms of art used with no gloss. A section quietly editorialising instead of reporting.&lt;/p&gt;

&lt;p&gt;Every single time, my first instinct was that the agent had drifted, misunderstood, or needed a firmer instruction.&lt;/p&gt;

&lt;p&gt;Every single time, I was wrong.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agents were following their role definitions to the letter.&lt;/strong&gt; The problem was that each definition had been written for an internal audience and then quietly promoted to a reader-facing one.&lt;/p&gt;

&lt;p&gt;The clearest case: the public writer had an instruction to faithfully preserve the source's framing. Perfectly reasonable. But the sources were Treasury and Federal Reserve releases, and they use the internal vocabulary. So the agent was being obedient when it passed that vocabulary straight through to readers.&lt;/p&gt;

&lt;p&gt;It was not drifting. It was following an order that had become wrong the moment its output started being published.&lt;/p&gt;

&lt;p&gt;Another role had an explicit invariant instructing it to cite records by identifier rather than by name, because that was precise and useful when its only reader was me. When its output was later routed into a published summary, that same invariant produced &lt;code&gt;events/1042&lt;/code&gt; in the prose. &lt;strong&gt;The instruction never changed. The audience did.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So the debugging question is not "what did the agent get wrong."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It is: who does this role think it is writing for, and is that still who reads it?&lt;/p&gt;

&lt;p&gt;I now treat every role definition as an editorial brief with a named audience. When a role's output changes destination, the definition gets rewritten rather than patched.&lt;/p&gt;

&lt;p&gt;The failure mode nobody warns you about is not the agent going off-script. It is the agent following a script you wrote for a different reader and forgot to update.&lt;/p&gt;

&lt;p&gt;The full piece covers what happens when three agents read the same document with different jobs, why I deleted my entire document-extraction layer, and how agent failures needed a taxonomy rather than a retry.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openred.space/blog/calling-an-agent-for-sentences.html" rel="noopener noreferrer"&gt;https://openred.space/blog/calling-an-agent-for-sentences.html&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>agents</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
