<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dimitris Kyrkos </title>
    <description>The latest articles on DEV Community by Dimitris Kyrkos  (@cyclopt_dimitrisk).</description>
    <link>https://dev.to/cyclopt_dimitrisk</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3723233%2F0d42e922-0dff-4ae9-b8b8-f9fcc40130b6.png</url>
      <title>DEV Community: Dimitris Kyrkos </title>
      <link>https://dev.to/cyclopt_dimitrisk</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cyclopt_dimitrisk"/>
    <language>en</language>
    <item>
      <title>Your AI feature isn't a feature. It's a dependency you don't control.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Thu, 01 Oct 2026 06:34:06 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/your-ai-feature-isnt-a-feature-its-a-dependency-you-dont-control-33jc</link>
      <guid>https://dev.to/cyclopt_dimitrisk/your-ai-feature-isnt-a-feature-its-a-dependency-you-dont-control-33jc</guid>
      <description>&lt;p&gt;In the tutorial, the AI call always works. You import the SDK, paste in an API key, await a completion, and render the result. Ship it.&lt;/p&gt;

&lt;p&gt;In production, that same line of code is a network call to someone else's infrastructure, running someone else's capacity planning, under someone else's incident response. Sometimes it's slow. Sometimes it rate-limits you. Sometimes it's just down.&lt;/p&gt;

&lt;p&gt;The question isn't whether your AI provider will have an outage. It's what your users see when it does. If the answer is "a frozen screen" or "Something went wrong," the problem isn't the provider. It's the architecture.&lt;/p&gt;

&lt;p&gt;Here are the failure modes I see most often, and what graceful degradation looks like for each.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 1: the whole page waits on the model
&lt;/h2&gt;

&lt;p&gt;Illustrative scenario: a project management app adds an AI-generated summary at the top of every ticket. The summary is fetched server-side before the page renders. One afternoon the provider's latency jumps from 2 seconds to 40. Now every ticket page takes 40 seconds to load, including for users who never read the summary.&lt;/p&gt;

&lt;p&gt;The AI part was optional. The architecture made it mandatory.&lt;/p&gt;

&lt;p&gt;Fix: never put a model call on the critical rendering path. Load the core page first, fetch AI output asynchronously, and give it a hard timeout. If it doesn't arrive, the slot stays empty or shows a quiet "summary unavailable." The ticket still opens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 2: one provider, hardcoded
&lt;/h2&gt;

&lt;p&gt;If your code calls one vendor's SDK directly in twenty places, you have twenty single points of failure and no switch to flip.&lt;/p&gt;

&lt;p&gt;Put a thin routing layer between your app and the providers, and make fallback a first-class behavior rather than a try/catch afterthought:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;primary&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;callPrimaryModel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;timeoutMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;8000&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;secondary&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;callSecondaryProvider&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;timeoutMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;8000&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;local&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;call&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;callLocalModel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;timeoutMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;4000&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;];&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Task&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Result&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;chain&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isOpen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;withTimeout&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;timeoutMs&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recordSuccess&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;servedBy&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recordFailure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// caller must handle "no AI available"&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two details matter more than the loop itself. The circuit breaker stops you from hammering a provider that's clearly down (and paying the full timeout on every request). And &lt;code&gt;null&lt;/code&gt; is a legitimate return value, which forces every caller to decide what happens without AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 3: the fallback model gets the same job
&lt;/h2&gt;

&lt;p&gt;A small local model (served through something like Ollama or llama.cpp) is a great last line of defense, but it isn't a drop-in replacement for a frontier model. Prompts tuned for one often produce garbage on the other.&lt;/p&gt;

&lt;p&gt;Decide in advance which tasks are essential enough to fall back locally: classification, short extraction, simple rewrites. Give those their own prompts and test them. Everything else should degrade to "not available right now" rather than to a confidently worse answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 4: the core features depend on the AI path
&lt;/h2&gt;

&lt;p&gt;This is the one that turns a provider outage into your outage. Search that only works through embeddings. A form that can't be submitted until the AI validator responds. An onboarding flow that stalls because the welcome message is generated.&lt;/p&gt;

&lt;p&gt;Every AI feature should have a non-AI path that still lets the user finish their job: keyword search behind semantic search, rule-based validation behind AI validation, a static template behind the generated one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resilience is a day-one decision
&lt;/h2&gt;

&lt;p&gt;Graceful degradation is really just treating your AI provider like any other external dependency: timeouts, fallbacks, circuit breakers, and a clear answer for "what if it's gone?" The teams that struggle are the ones who treated the model as part of their own code instead of as a service they rent.&lt;/p&gt;

&lt;p&gt;Retrofitting this later means touching every call site. Designing it in from the start means one routing layer and a few decisions.&lt;/p&gt;

&lt;p&gt;What does your app actually do today when your AI provider goes down? Have you tested it, or are you assuming?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>webdev</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Half the AI agents in production are if-statements with a GPU bill</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Mon, 28 Sep 2026 07:45:38 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/half-the-ai-agents-in-production-are-if-statements-with-a-gpu-bill-4934</link>
      <guid>https://dev.to/cyclopt_dimitrisk/half-the-ai-agents-in-production-are-if-statements-with-a-gpu-bill-4934</guid>
      <description>&lt;p&gt;There's a new kind of technical debt, and it doesn't come from cutting corners. It comes from reaching for the most impressive tool in the room.&lt;/p&gt;

&lt;p&gt;Call it resume-driven AI engineering: picking an agent framework, a vector database, or a multi-model orchestration layer because it looks great on a CV, not because the problem needs it. The result works in the demo. It's also slower, more expensive, harder to debug, and nondeterministic in places where it didn't need to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  The demo vs. the pager
&lt;/h2&gt;

&lt;p&gt;In a tutorial, complexity is free. You spin up an agent, wire in a vector store, and watch it do something clever with ten sample documents.&lt;/p&gt;

&lt;p&gt;In production, every moving part has a cost: latency, token spend, a new failure mode, a new thing someone has to understand at 3 a.m. The question isn't "can an LLM do this?" (it usually can). It's "is an LLM the simplest thing that does this reliably?"&lt;/p&gt;

&lt;p&gt;Here are three places where the answer is often no. The scenarios are illustrative, but if you've been around AI projects for a while, they'll look familiar.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 1: a model call where a regex would do
&lt;/h2&gt;

&lt;p&gt;A team needs to pull invoice numbers out of incoming emails. Invoice numbers follow a fixed format: &lt;code&gt;INV-&lt;/code&gt; plus eight digits. They send every email to an LLM with a prompt asking it to extract the number.&lt;/p&gt;

&lt;p&gt;It works 98% of the time. The other 2%, the model "helpfully" reformats the number, or picks up a purchase order number instead. Each call costs money and adds a few hundred milliseconds.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;

&lt;span class="n"&gt;INVOICE_RE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\bINV-\d{8}\b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;extract_invoice_ids&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;INVOICE_RE&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deterministic, testable, effectively free, and it runs in microseconds. Keep the model for the messy cases the pattern can't handle, and route to it only when the regex finds nothing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 2: vector search where SQL would do
&lt;/h2&gt;

&lt;p&gt;"Show me all orders from customer 4417 in the last 30 days that are still unpaid."&lt;/p&gt;

&lt;p&gt;That's not a semantic question. It's a filter. Yet it's common to see this kind of query embedded, pushed through a vector store, and answered by an LLM summarizing the top-k chunks, which may or may not include every matching order.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4417&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'unpaid'&lt;/span&gt;
  &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;NOW&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;INTERVAL&lt;/span&gt; &lt;span class="s1"&gt;'30 days'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Exact, complete, indexed, auditable. Vector search is great when you're matching meaning ("tickets similar to this complaint"). It's the wrong tool when you're matching facts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 3: an autonomous agent where a decision tree would do
&lt;/h2&gt;

&lt;p&gt;A support workflow: if the customer is on the enterprise plan and the issue is billing, route to account management; if it's a bug, open a ticket; otherwise, send the FAQ link.&lt;/p&gt;

&lt;p&gt;That's four branches. Someone builds it as an autonomous agent with tool access, a planning loop, and a memory store. Now the routing is probabilistic, occasionally loops, and nobody can explain why ticket #8812 went to the wrong team.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;plan&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;enterprise&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;account_management&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;type&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bug&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;open_ticket&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;send_faq&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you can draw the logic on a whiteboard, you probably don't need an agent to rediscover it every request.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;A lot of "AI systems" are really ordinary software with an LLM bolted onto a step that didn't need one. The model isn't the problem. Using it as the default instead of the exception is.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical decision checklist
&lt;/h2&gt;

&lt;p&gt;Before adding a framework, a model call, or an agent, ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Is the input structured or the output fixed-format?&lt;/strong&gt; Start with parsing, regex, or schema validation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the question about facts or about meaning?&lt;/strong&gt; Facts go to SQL. Meaning can go to embeddings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can the logic be enumerated?&lt;/strong&gt; If yes, write the branches. Agents are for open-ended tasks where you genuinely can't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happens when it's wrong?&lt;/strong&gt; If the answer is "silent bad data," you want determinism.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who maintains this in a year?&lt;/strong&gt; Every framework is a dependency someone has to upgrade, understand, and debug.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this means "never use AI." Use it where ambiguity actually lives: unstructured text, fuzzy matching, generation. Just make it the tool you reach for on purpose, not by reflex.&lt;/p&gt;

&lt;p&gt;The best engineers aren't the ones with the most complex stack. They're the ones whose systems are still simple enough to understand when something breaks.&lt;/p&gt;

&lt;p&gt;What's the most over-engineered AI setup you've seen (or built) that could've been replaced with something boring?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>architecture</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Your model doesn't need more training. It needs a better search index.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Thu, 24 Sep 2026 09:40:25 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/your-model-doesnt-need-more-training-it-needs-a-better-search-index-3mca</link>
      <guid>https://dev.to/cyclopt_dimitrisk/your-model-doesnt-need-more-training-it-needs-a-better-search-index-3mca</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;Every team that plugs an LLM into its business hits the same moment. Someone asks the model about an internal pricing rule, a product SKU, or a clause in the standard contract, and it answers confidently and wrongly.&lt;/p&gt;

&lt;p&gt;The reflex that follows is almost universal: "The model doesn't know our business. Let's fine-tune it."&lt;/p&gt;

&lt;p&gt;So the team spends weeks exporting tickets and docs, cleaning them, formatting them into JSONL, and paying for training runs. The new model sounds more like the company. It uses the right jargon. And it still makes things up.&lt;/p&gt;

&lt;p&gt;That's not bad luck. It's a category error.&lt;/p&gt;

&lt;h2&gt;
  
  
  What fine-tuning is actually good at
&lt;/h2&gt;

&lt;p&gt;Fine-tuning adjusts a model's weights so it behaves differently. That's powerful when the thing you want to change is behavior:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A specific tone or writing style (support replies that sound like your brand, not like a chatbot)&lt;/li&gt;
&lt;li&gt;A strict output format (always return this JSON schema, always produce this report layout)&lt;/li&gt;
&lt;li&gt;A narrow, repetitive task where a smaller tuned model can replace a bigger general one&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Notice what all of those have in common. They're about how the model answers, not what facts it knows.&lt;/p&gt;

&lt;p&gt;When the goal is "the model should know our refund policy, our product catalog, and last quarter's changes," you're asking weights to act like a database. They're a bad one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 1: it still hallucinates, just in your accent
&lt;/h2&gt;

&lt;p&gt;A fine-tuned model doesn't gain a reliable sense of what it knows and what it doesn't. Training on your documents shifts probabilities toward your vocabulary, but when a question lands in a gap, the model does what it always does: it produces the most plausible-sounding continuation.&lt;/p&gt;

&lt;p&gt;Illustrative scenario: a support assistant tuned on two years of tickets is asked about a warranty extension introduced last month. It has never seen it. It answers anyway, in perfect company voice, with the terms of the old warranty. The answer is more convincing than a generic model's would have been, which makes it more dangerous, not less.&lt;/p&gt;

&lt;p&gt;Fluency in your domain is not the same as accuracy about your domain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 2: every fact change is a training run
&lt;/h2&gt;

&lt;p&gt;Business knowledge isn't static. Prices change, policies get revised, products get deprecated, a regulation lands and three documents get rewritten.&lt;/p&gt;

&lt;p&gt;If that knowledge lives in the weights, every update means rebuilding the dataset, retraining, re-evaluating, and redeploying. In practice, teams don't do that weekly. So the model drifts out of date, quietly, and nobody knows exactly which facts are stale.&lt;/p&gt;

&lt;p&gt;Compare that with a retrieval setup: a document changes, you re-index it, and the next query sees the new version. The update cycle is minutes, not a project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Failure mode 3: you can't show your work
&lt;/h2&gt;

&lt;p&gt;When an answer comes from retrieval, you can point to the exact chunk of the exact document it was grounded in. You can show it to the user. You can log it. When someone disputes an answer, you can check whether the source was wrong or the model misread it.&lt;/p&gt;

&lt;p&gt;When an answer comes from fine-tuned weights, there is no source. The knowledge is smeared across billions of parameters. You can't cite it, you can't audit it, and you can't tell a compliance team where a claim came from.&lt;/p&gt;

&lt;p&gt;For anything customer-facing, regulated, or contractual, that alone should end the discussion.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to build first instead
&lt;/h2&gt;

&lt;p&gt;Before training custom weights, invest in the boring part: retrieval.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user question
   -&amp;gt; query rewriting / expansion
   -&amp;gt; hybrid search (keyword + embeddings) over a clean, deduplicated index
   -&amp;gt; reranking
   -&amp;gt; top-k chunks + source metadata into the prompt
   -&amp;gt; answer with citations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Most of the quality comes from things that have nothing to do with the model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Clean source documents (no three conflicting versions of the same policy)&lt;/li&gt;
&lt;li&gt;Sensible chunking that keeps related context together&lt;/li&gt;
&lt;li&gt;Metadata (dates, owners, product, region) you can filter on&lt;/li&gt;
&lt;li&gt;Hybrid search, because embeddings alone miss exact identifiers like SKUs and error codes&lt;/li&gt;
&lt;li&gt;An evaluation set of real questions with known correct answers, so you can measure whether changes help&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is glamorous. All of it pays off more than a training run for knowledge-heavy use cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  When fine-tuning does make sense
&lt;/h2&gt;

&lt;p&gt;This isn't "never fine-tune." It's "fine-tune for the right reason, and usually later":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retrieval is solid, answers are grounded, but the output format or tone keeps drifting. Tune for behavior.&lt;/li&gt;
&lt;li&gt;You need a smaller, cheaper model to handle a narrow, high-volume task a big model already does well. Tune for cost.&lt;/li&gt;
&lt;li&gt;The model struggles to use retrieved context well in your domain (for example, dense technical or legal text). Tune on examples of reading and citing context, not on the facts themselves.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In all three cases, facts still come from the index. The weights handle behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;Fine-tuning teaches a model how to talk. Retrieval tells it what's true right now. Most "the model doesn't understand our business" problems are the second kind wearing the costume of the first.&lt;/p&gt;

&lt;p&gt;What pushed your team toward fine-tuning (or away from it), and did it actually fix the knowledge problem you were trying to solve?&lt;/p&gt;




</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>architecture</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Your LLM has no memory. Your application had better have one.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Mon, 21 Sep 2026 10:05:47 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/your-llm-has-no-memory-your-application-had-better-have-one-38mf</link>
      <guid>https://dev.to/cyclopt_dimitrisk/your-llm-has-no-memory-your-application-had-better-have-one-38mf</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;Every LLM tutorial has the same shape. You send a list of messages, you get a reply, you append the reply, you send the list again. It works, it demos well, and it quietly teaches you the wrong mental model.&lt;/p&gt;

&lt;p&gt;That loop is not state management. It is the absence of it, dressed up as a feature.&lt;/p&gt;

&lt;p&gt;The APIs behind LLM platforms are stateless. Each request is independent of the last one. The model does not "remember" your conversation, it re-reads whatever you hand it, every single time. Meanwhile the thing you are building, a support flow, a code migration, an approval process, a multi-step agent, is stateful by nature. Somebody started it, it is halfway done, and it has to survive interruptions, retries and crashes.&lt;/p&gt;

&lt;p&gt;That gap between a stateless processor and a stateful process is where production systems break. Here are the three places I see it break first.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. "Just send the history" stops scaling
&lt;/h2&gt;

&lt;p&gt;The tutorial answer to "how does the model know what happened earlier?" is to resend the whole transcript. It is fine for ten turns. At a hundred turns you are paying for the same tokens over and over, latency creeps up, and you start hitting context limits at the worst possible moment.&lt;/p&gt;

&lt;p&gt;The usual patch is summarization: compress old turns into a paragraph and carry on. That helps with size but introduces a subtler problem. The summary is now the model's opinion of what mattered, and nothing in your system can verify it.&lt;/p&gt;

&lt;p&gt;Imagine a multi-step refund workflow (illustrative scenario, not a real case). At turn 4 the user confirms the order ID. At turn 30 the summary says "user wants a refund for a recent order." The ID got compressed away, and the model now confidently picks the wrong one.&lt;/p&gt;

&lt;p&gt;The fix is to stop treating the transcript as the state. Facts that the process depends on belong in a structured object your code owns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@dataclass&lt;/span&gt;
&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;RefundState&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;order_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
    &lt;span class="n"&gt;amount_confirmed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;
    &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;collect_order&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;   &lt;span class="c1"&gt;# explicit position in the workflow
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;build_prompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;RefundState&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recent_turns&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Message&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="c1"&gt;# The model gets the current state as input, plus only the turns it needs.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;system_prompt&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="nf"&gt;state_message&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;recent_turns&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;:]]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the context you send is small, deterministic and reviewable. The model reads state, it does not store it.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The user interrupts halfway through
&lt;/h2&gt;

&lt;p&gt;Real users do not wait politely for a long-running process to finish. They close the tab, change their mind, or send "actually, cancel that" while three tool calls are still in flight.&lt;/p&gt;

&lt;p&gt;If the only record of progress is "whatever the model said last," you cannot answer basic questions. Which steps already ran? Which are safe to abandon? Does the half-finished action need to be rolled back?&lt;/p&gt;

&lt;p&gt;A workflow that can be interrupted needs explicit steps with explicit statuses, persisted outside the model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Step&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Enum&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;PENDING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;RUNNING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;running&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;DONE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;done&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;CANCELLED&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cancelled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="c1"&gt;# Persisted per run, updated by your code, never inferred from model output.
&lt;/span&gt;&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reserve_inventory&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DONE&lt;/span&gt;
&lt;span class="n"&gt;run&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;steps&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;charge_card&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Step&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;RUNNING&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the interrupt arrives, your code reads the run record, decides what a cancellation means at this point, and tells the model the outcome. The model is not asked to work out where it was.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The API call fails in the middle of a transaction
&lt;/h2&gt;

&lt;p&gt;Retries are where stateless design bites hardest. If a request times out after your tool executed but before you saw the response, resending it can execute the tool twice. Charge the card twice. Send the email twice. Open two tickets.&lt;/p&gt;

&lt;p&gt;This is an old distributed systems problem, and the old answers apply: idempotency keys, an append-only event log, and recovery by replay.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;execute_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;run_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;ToolCall&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;run_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;step_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="nf"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_result&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;                      &lt;span class="c1"&gt;# already ran, do not run again
&lt;/span&gt;    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;idempotency_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;store&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;save_result&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that none of this depends on the model behaving well. That is the point. A model can be asked to "remember not to repeat itself," and it will still eventually repeat itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern underneath
&lt;/h2&gt;

&lt;p&gt;All three failures come from the same mistake: letting the conversation double as the system of record. Once you separate the two, the design gets much less mysterious.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;State&lt;/strong&gt; lives in your application: structured, persisted, versioned, recoverable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The model&lt;/strong&gt; is a stateless processor. It receives the relevant state, produces a proposed next action, and your code validates and applies it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The transcript&lt;/strong&gt; is a log for humans and for debugging, not the source of truth.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An LLM workflow is really a state machine with a language model on top of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build it or lean on the platform?
&lt;/h2&gt;

&lt;p&gt;Some providers now offer conversation or thread objects that manage history for you. They are convenient for prototypes and for simple chat. Before relying on them for a business process, ask a few questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can you inspect and export the state in a form your own code can reason about?&lt;/li&gt;
&lt;li&gt;Can you resume, replay or roll back a run after a failure?&lt;/li&gt;
&lt;li&gt;Can you move to another model or provider without losing the process?&lt;/li&gt;
&lt;li&gt;Who is accountable when the stored history and reality disagree?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the answers are "not really," keep the state in your own store and treat the platform's memory as a cache at best. Managed history is fine as an optimization. It should not be where correctness lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing thought
&lt;/h2&gt;

&lt;p&gt;If you rely on the model to remember the state of the conversation, your system will eventually break, and it will break in the least reproducible way available. Keep your application code as the source of truth and let the model do what it is good at: processing what you hand it, one request at a time.&lt;/p&gt;

&lt;p&gt;Where does your team keep workflow state for multi-step LLM interactions: your own database, an event log, or the provider's thread objects? And what made you pick it?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>programming</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Two-second latency isn't an AI problem. It's an architecture problem your stack was never built to hide.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Fri, 18 Sep 2026 06:25:25 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/two-second-latency-isnt-an-ai-problem-its-an-architecture-problem-your-stack-was-never-built-to-32mj</link>
      <guid>https://dev.to/cyclopt_dimitrisk/two-second-latency-isnt-an-ai-problem-its-an-architecture-problem-your-stack-was-never-built-to-32mj</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;There's a gap between the demo and the deployment that nobody puts in the tutorial.&lt;/p&gt;

&lt;p&gt;In the demo, you call the model, you get a response, you render it. Done. In production, that same call sits in your request path for one, two, sometimes four seconds, and your users are staring at whatever you left on screen while they wait.&lt;/p&gt;

&lt;p&gt;We spend weeks shaving milliseconds off database queries and tuning cache layers to get a page from 180ms to 90ms. Then we bolt an LLM call onto the same request/response cycle and it takes two full seconds to resolve. Nobody notices the discrepancy until users start bouncing.&lt;/p&gt;

&lt;p&gt;Here's the actual gap: traditional backend work happens in a latency band humans don't perceive. LLM inference happens in a band they very much do. Same request/response pattern, completely different experience on the other end of it. If you don't change how you architect for it, users feel the friction immediately, no matter how good the model's output is.&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure mode 1: the silent request
&lt;/h3&gt;

&lt;p&gt;Picture a support-ticket triage feature. User submits a ticket, the backend calls an LLM to classify and route it, and the endpoint doesn't return until the model finishes. On a slow model day, that's a three-second white screen with no feedback. The user assumes the button didn't register and clicks again. Now you've got two calls in flight for the same ticket. This is a made-up example, but if you've shipped an LLM feature behind a synchronous POST, you've lived some version of it.&lt;/p&gt;

&lt;p&gt;The fix isn't a faster model, it's not treating the call as instant in the first place.&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure mode 2: streaming as an afterthought
&lt;/h3&gt;

&lt;p&gt;Streaming gets treated as a nice-to-have polish pass instead of the default. Teams ship the blocking version first, get it working end to end, and plan to "add streaming later." Later rarely comes, because by then the blocking version is load-bearing and touching it means retesting the whole flow. Token-by-token streaming isn't a UX nicety here, it's the difference between "the interface feels instant" and "the interface feels broken."&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure mode 3: heavy work still living in the request path
&lt;/h3&gt;

&lt;p&gt;Some LLM work genuinely doesn't need to block the response at all: summarizing a document after upload, generating suggested tags, running a background analysis pass. These still get built as synchronous calls because that's the default pattern every other endpoint in the codebase already uses. Moving that work to an async job with a webhook or polling endpoint is often a bigger UX win than any amount of prompt optimization, because the user was never meant to wait on it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Failure mode 4: no optimistic path forward
&lt;/h3&gt;

&lt;p&gt;Even with streaming and async jobs in place, there's often no way for the user to keep moving while the model works. Optimistic UI, showing the next step as available, prefilling a draft state, letting the user act on a provisional result, keeps people in the workflow instead of parked on a spinner watching a progress indicator that means nothing to them.&lt;/p&gt;

&lt;h3&gt;
  
  
  The reframe
&lt;/h3&gt;

&lt;p&gt;None of this is really an "AI latency" problem. It's a request/response architecture problem that AI just made visible, because AI is the first thing in most stacks slow enough that users notice the assumption baked into synchronous request handling. The fix isn't specific to LLMs: stream what you can, push what doesn't need to block into the background, and give the user something to do while they wait.&lt;/p&gt;

&lt;p&gt;When you're deciding which technique to reach for: if the output needs to be seen as it's produced, stream it. If it doesn't need to block the next user action at all, take it out of the request path entirely. If it has to block but you can predict the shape of the result, use optimistic UI to bridge the gap.&lt;/p&gt;

&lt;p&gt;We hit this directly while building the analysis pipeline behind Cyclopt: the moment an LLM call entered the loop, the whole interaction model around it had to change, not just the backend call itself.&lt;/p&gt;

&lt;p&gt;How is your team handling this? Streaming everything by default, pushing more to background jobs, or still treating the LLM call like any other backend request?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Your LLM Isn't Bad At Math. It Was Never Doing Math In The First Place.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Tue, 15 Sep 2026 06:31:43 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/your-llm-isnt-bad-at-math-it-was-never-doing-math-in-the-first-place-3j67</link>
      <guid>https://dev.to/cyclopt_dimitrisk/your-llm-isnt-bad-at-math-it-was-never-doing-math-in-the-first-place-3j67</guid>
      <description>&lt;h2&gt;
  
  
  Your LLM Isn't Bad At Math. It Was Never Doing Math In The First Place.
&lt;/h2&gt;

&lt;p&gt;In a tutorial, an LLM call looks like a function: pass in text, get back an answer, move on. It's easy to start treating the model like it's evaluating your business logic the same way a function would, deterministically, the same input always producing the same output.&lt;/p&gt;

&lt;p&gt;In production, that assumption breaks in a specific, predictable way, and it's worth naming precisely why.&lt;/p&gt;

&lt;p&gt;Traditional software runs on deterministic logic: if input A meets condition B, output C happens, every time, with a proof you could write on a whiteboard. An LLM runs on statistical probability: it predicts the most likely next token given the patterns in its training data. Those are two different kinds of "correct." The friction shows up exactly where a team asks the second kind to do the first kind's job.&lt;/p&gt;

&lt;p&gt;Below are the places that friction actually shows up in a live system. These are illustrative scenarios based on the shape this failure takes, not one specific incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Asking a model to do arithmetic is asking it to guess what arithmetic usually looks like
&lt;/h2&gt;

&lt;p&gt;Say a pipeline extracts line items from a batch of invoices, then asks the same model call to also compute the subtotal, apply tax, and return a total. On short, simple invoices, it's right almost every time, because short simple math is heavily represented in training data and easy to pattern-match. On a 40-line invoice with mixed tax rates and a rounding rule, it starts being right most of the time, which is a different and much worse thing than right. Nothing throws an error. The number is just plausible instead of correct, and "plausible instead of correct" is invisible until someone reconciles the books.&lt;/p&gt;

&lt;h2&gt;
  
  
  A statistically-applied business rule is not the same rule as a logically-applied one
&lt;/h2&gt;

&lt;p&gt;Say the business rule is "refunds are allowed within 30 days of purchase." Put that rule in a prompt and ask the model to decide eligibility case by case, and it will get the easy cases right: a purchase from six months ago, obviously no; a purchase from yesterday, obviously yes. Where it gets interesting is the boundary: day 29, day 30, day 31, across time zones, with a purchase timestamp in one format and today's date passed in another. A deterministic date comparison gets this right every single time by construction. A model is producing its best guess at what "a refund decision near the boundary" looks like, based on how those decisions were phrased in its training data, and that's a meaningfully different operation even when it happens to output the correct answer nine times out of ten.&lt;/p&gt;

&lt;h2&gt;
  
  
  The output still looks like a normal response, which is exactly the problem
&lt;/h2&gt;

&lt;p&gt;This is what makes the failure mode dangerous rather than just annoying: a wrong statistical answer doesn't look different from a right one. It's the same JSON shape, the same confident tone, the same absence of an exception. There's no signal in the response itself that tells you this was a guess rather than a computation. You find out later, from a customer service escalation or a finance reconciliation, not from anything in your logs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the line actually goes
&lt;/h2&gt;

&lt;p&gt;None of this means don't use the model near your business logic. It means be precise about which half of the job you're handing it.&lt;/p&gt;

&lt;p&gt;What the model is good at: turning unstructured input, an email, a contract clause, a support ticket, a scanned form, into structured data. That's genuinely ambiguous work: language is fuzzy, humans phrase the same request ten different ways, and a statistical model that's seen millions of phrasings is the right tool for mapping "hey can I send this back, it's been like a month" into structured fields like intent and stated days since purchase.&lt;/p&gt;

&lt;p&gt;What the model should never be the last word on: the validation, the calculation, and the rule enforcement that runs on that structured data once you have it. That's deterministic code's job, on purpose, with a strict schema at the boundary so a malformed or out-of-range extraction fails loudly instead of quietly flowing downstream as a confident-looking guess.&lt;/p&gt;

&lt;p&gt;We lean on this split in how Cyclopt's automated checks work: the model interprets unstructured signals in a codebase or a pull request, but the actual rule enforcement, the pass or fail line, runs through deterministic logic against a defined schema, not through the model re-deciding the rule each time. The interpretation layer changes. The rule layer doesn't get to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;Statistical vs. logical correctness isn't a model-quality problem that gets solved by a better model. It's an architecture decision, and better models make it easier to ignore, not less necessary to make. Let the model handle the ambiguity. Let your code handle the rules.&lt;/p&gt;

&lt;p&gt;Where do you actually draw that line in a system you've shipped? What's the one thing you learned the hard way should never have been the model's call to make?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>programming</category>
      <category>discuss</category>
    </item>
    <item>
      <title>AI Model Upgrades Aren't Patches. They're Silent Breaking Changes Wearing a Version Bump.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Fri, 11 Sep 2026 06:15:02 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/ai-model-upgrades-arent-patches-theyre-silent-breaking-changes-wearing-a-version-bump-md8</link>
      <guid>https://dev.to/cyclopt_dimitrisk/ai-model-upgrades-arent-patches-theyre-silent-breaking-changes-wearing-a-version-bump-md8</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;AI Model Upgrades Aren't Patches. They're Silent Breaking Changes Wearing a Version Bump.&lt;/p&gt;

&lt;p&gt;In a tutorial, calling an LLM API looks simple: send a prompt, get a response, parse it, move on. The model is treated like any other stable dependency, something that behaves today the way it behaved last week.&lt;/p&gt;

&lt;p&gt;In production, that assumption quietly stops being true, and nothing in your stack tells you when it happens.&lt;/p&gt;

&lt;p&gt;When a database vendor ships a patch, your queries keep returning the same shape of data they always did. Compatibility is the whole point of a patch. When an AI provider updates a model, even a version bump billed as a minor optimization can change how that model interprets the exact same system prompt it handled fine yesterday. Nobody signs off on that change on your end. Nobody flags it in your changelog. It just starts behaving differently, silently, the next time your endpoint routes to the new weights.&lt;/p&gt;

&lt;p&gt;Below are the places that drift actually shows up. These are illustrative scenarios based on the shape this class of failure takes, not one specific incident.&lt;/p&gt;

&lt;h2&gt;
  
  
  Structured output schemas quietly reshape themselves
&lt;/h2&gt;

&lt;p&gt;Say a pipeline asks a model to return JSON with keys like &lt;code&gt;status&lt;/code&gt;, &lt;code&gt;confidence&lt;/code&gt;, and &lt;code&gt;flagged_reason&lt;/code&gt;. That's worked reliably for months against a pinned prompt. Then the underlying model gets swapped for a newer checkpoint behind the same generic endpoint, and it starts returning &lt;code&gt;reason_flagged&lt;/code&gt; instead of &lt;code&gt;flagged_reason&lt;/code&gt; on a subset of edge-case inputs, close enough to pass a casual glance, different enough to break every downstream consumer expecting the old key. No exception gets thrown. The field is just missing, and whatever reads it either crashes on a null or, worse, silently treats the absence as "nothing flagged."&lt;/p&gt;

&lt;h2&gt;
  
  
  Guardrails degrade exactly where you can't see them
&lt;/h2&gt;

&lt;p&gt;Validation logic built around a model's known quirks, how it phrases refusals, how it handles ambiguous instructions, where it tends to hedge, is built against a specific behavioral fingerprint. Swap the model and that fingerprint shifts. A regex or classifier tuned to catch a model's old refusal pattern stops catching the new one. The guardrail doesn't fail loudly. It just quietly stops doing its job on the exact edge cases it existed to catch, the ones rare enough that nobody's manually reviewing that output anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  "It still returns 200" is why nobody notices
&lt;/h2&gt;

&lt;p&gt;This is the core of why these failures are dangerous: nothing about them looks like an error. The API call succeeds. The response is well-formed enough to pass a basic type check. The status code is fine. The only thing wrong is that the content is subtly different from what every downstream system was built to expect, and that difference doesn't surface as a stack trace. It surfaces three weeks later as a support ticket about corrupted records, or a data quality review that turns up a cluster of malformed rows nobody can explain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Auto-updating endpoints turn every deploy into an uncontrolled experiment
&lt;/h2&gt;

&lt;p&gt;Pointing at a generic, "always latest" model endpoint feels convenient right up until the provider ships an update on its own schedule instead of yours. At that point every request your system sends is effectively running against a dependency you never tested, on a timeline you don't control, with no diff to review before it goes live. Nobody would deploy a database migration that way. Plenty of teams are doing the equivalent with their model layer without noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;This isn't really an AI reliability problem. It's a dependency management problem that happens to involve a model, and it deserves the same discipline any other production dependency gets: pin the version, test before you upgrade, and monitor for drift instead of waiting for it to show up as bad data.&lt;/p&gt;

&lt;p&gt;We ended up building a version of this discipline into how Cyclopt's automated checks handle LLM-based analysis, mostly because the alternative is finding out about drift from a bug report instead of a test run.&lt;/p&gt;

&lt;p&gt;Concretely, that means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pin exact model versions.&lt;/strong&gt; Point at a specific checkpoint, not a generic "latest" alias, so an upgrade is something you opt into instead of something that arrives unannounced.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build regression tests that run before any model version change.&lt;/strong&gt; Treat a model swap the same way you'd treat a major dependency bump: a known set of inputs, known expected shapes, run before the new version touches production traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor output schemas in real time.&lt;/strong&gt; Track the actual shape of what's coming back, not just whether the call succeeded, so schema drift shows up as an alert instead of a data quality incident.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stability in the AI ecosystem isn't something you get by default. It's something you have to actively engineer, the same way you'd engineer it for any other dependency you don't control the release schedule of.&lt;/p&gt;

&lt;p&gt;Has a silent model upgrade ever broken something in your pipeline before you knew what was happening? What caught it: a test, a monitor, or a support ticket?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>programming</category>
      <category>discuss</category>
    </item>
    <item>
      <title>AI Didn't Kill the Need for System Design. It Just Made Bad System Design Easier to Ship.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Tue, 08 Sep 2026 09:10:19 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/ai-didnt-kill-the-need-for-system-design-it-just-made-bad-system-design-easier-to-ship-44fg</link>
      <guid>https://dev.to/cyclopt_dimitrisk/ai-didnt-kill-the-need-for-system-design-it-just-made-bad-system-design-easier-to-ship-44fg</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;AI Didn't Kill the Need for System Design. It Just Made Bad System Design Easier to Ship.&lt;/p&gt;

&lt;p&gt;Every "no-code AI" pitch makes the same promise: describe what you want, and the AI builds it. No engineers required.&lt;/p&gt;

&lt;p&gt;Here's the gap nobody puts in the demo video: getting a working prototype out of a prompt and getting a production system that survives real users are two completely different problems. The first is a syntax problem. AI is genuinely good at that part. The second is a system design problem, and system design doesn't get easier just because a model is writing the implementation. If anything, it gets more dangerous, because the code now looks finished before anyone has checked whether the underlying architecture is sound.&lt;/p&gt;

&lt;p&gt;Below are four places where that gap shows up hardest. The scenarios are illustrative, not real incidents, but they're the shape of thing that happens constantly once a "no-code" app leaves the demo and meets actual users.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data normalization is invisible until it corrupts something
&lt;/h2&gt;

&lt;p&gt;Picture a no-code AI tool asked to "build a customer database with orders and shipping addresses." It'll happily generate tables. What it won't do on its own is ask whether an address belongs to a customer or to an order, whether a customer can have multiple addresses, or what happens when someone updates an address mid-order. Get that wrong and you don't get an error message, you get silently wrong data: an order ships to an address the customer changed three weeks ago, and nobody notices until support tickets pile up. Schema design is a series of judgment calls about what the business actually means by its own data, and an AI has no way to know that unless someone tells it, explicitly, up front.&lt;/p&gt;

&lt;h2&gt;
  
  
  State persistence across sessions is where "it works in the demo" quietly stops being true
&lt;/h2&gt;

&lt;p&gt;A prompt-generated checkout flow will work perfectly in a five-minute demo where one person clicks through it once. Put it in front of real traffic and now you've got a user who adds an item to their cart on their phone, closes the tab, comes back on their laptop an hour later, and expects the cart to still be there. Or two tabs open at once, both trying to update the same session state. None of that is a coding problem in the "write more code" sense. It's a decision about where state lives, how it's synchronized, and what "the same session" even means across devices and time, decisions that have to be made before a single line gets written, not discovered after the fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  API rate limits and error handling are a design decision wearing a bug report's clothes
&lt;/h2&gt;

&lt;p&gt;When a no-code AI app calls a third-party payment or shipping API, what happens when that API returns a 429, or times out, or comes back with a response the app didn't expect? An AI generating happy-path code by default will wire up the success case and leave the rest as an unhandled exception. That's not a bug you patch later, it's an architectural gap: does the system retry, queue, degrade gracefully, or fail loudly? Every one of those is a different design, and none of them are implied by "connect to the payment API."&lt;/p&gt;

&lt;h2&gt;
  
  
  Security boundaries don't announce themselves in a prompt
&lt;/h2&gt;

&lt;p&gt;"Let users manage their own profile" sounds simple enough to generate. Whether that same code accidentally lets user A read or edit user B's profile by changing an ID in a request is an access control question, and access control is exactly the kind of thing that looks correct in a demo (because the demo only ever has one user) and falls apart the moment two accounts exist. This is architecture, not implementation detail, and it's usually the last thing anyone checks because it's the thing that's hardest to notice is missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;No-code AI isn't really removing the need for engineering. It's removing the friction of writing syntax and quietly reassigning the hard part, the architectural judgment, to whoever is willing to notice it's still required. The teams that get burned aren't the ones who used AI to generate code. They're the ones who mistook "the AI wrote it and it ran" for "the AI designed it and it's sound."&lt;/p&gt;

&lt;p&gt;We've run into a version of this while building Cyclopt's Companion: code generated by an AI assistant will often pass a lint check and even a basic security scan while the schema or state model underneath it is quietly wrong. Catching that isn't a code-generation problem, it's a "does someone with architectural judgment ever look at this" problem, and that's a different tool and a different habit than "prompt better."&lt;/p&gt;

&lt;p&gt;So here's the actual decision teams need to make: use AI to generate implementation once the architecture is already decided by someone who understands the domain, or let AI-generated architecture ship and find out what broke in production. Those are very different bets.&lt;/p&gt;

&lt;p&gt;Where has this bitten you, or your team, hardest? Data model drift, session bugs, or something in access control that only showed up after the fact?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>programming</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Your AI-generated tests aren't testing your code. They're testing the AI's blind spots.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Fri, 04 Sep 2026 06:54:42 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/your-ai-generated-tests-arent-testing-your-code-theyre-testing-the-ais-blind-spots-46mo</link>
      <guid>https://dev.to/cyclopt_dimitrisk/your-ai-generated-tests-arent-testing-your-code-theyre-testing-the-ais-blind-spots-46mo</guid>
      <description>&lt;h3&gt;
  
  
  Intro
&lt;/h3&gt;

&lt;p&gt;There's a pitch behind every "AI writes your tests too" workflow: more coverage, less manual toil, a safety net that used to take a sprint now takes minutes.&lt;/p&gt;

&lt;p&gt;The pitch skips over what that safety net is actually made of. When the same model writes the implementation and the test suite, you haven't added a second, independent check. You've asked one reviewer to grade its own homework and handed you the green checkmark as if someone else had signed off.&lt;/p&gt;

&lt;h3&gt;
  
  
  The blind spot loop
&lt;/h3&gt;

&lt;p&gt;A model reasons about a function once, forms an implicit set of assumptions (input shapes, timezone handling, what counts as "empty"), and writes the implementation against those assumptions. Ask the same model to write tests for that function, and it doesn't re-derive correct behavior from scratch. It writes tests against the same assumptions it just used to write the code. If it assumed dates always arrive as ISO strings in UTC, the implementation assumes that, and the tests assume it too. The suite goes green. The assumption is still wrong.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tests that pass for the wrong reason
&lt;/h3&gt;

&lt;p&gt;(Illustrative, not a specific case, but recognizable to anyone who's shipped an AI-generated suite.) Picture a discount-calculation function where the model assumes quantities are always positive integers. The implementation skips a negative-quantity check. The generated tests exercise 1, 5, and 100, because those are the "normal" values a model reaching for plausible test data will reach for. Nothing ever asks what happens at -1 or 0, because neither pass, the code or the tests, ever considered them worth asking about. Coverage tooling reports 100% on this function. The bug ships anyway.&lt;/p&gt;

&lt;h3&gt;
  
  
  Coverage becomes a false signal
&lt;/h3&gt;

&lt;p&gt;High line or branch coverage from an AI-authored suite tells you the code paths were exercised, not that the right inputs exercised them. A suite can hit every line of a function and still never send it a null, an empty array, a duplicate key, or a value at a type boundary, if the author, human or model, never imagined those as possibilities in the first place. Coverage percentage was never built to detect a shared blind spot. It just counts what got tried.&lt;/p&gt;

&lt;h3&gt;
  
  
  What to keep human, and what to hand to AI
&lt;/h3&gt;

&lt;p&gt;The fix isn't "stop using AI for tests." It's separating the two jobs testing actually does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Defining correct behavior (the assertions) should come from a human who understands the spec or the business rule the function is supposed to honor, decided independently of however the implementation happens to work.&lt;/li&gt;
&lt;li&gt;Generating volume and variety (mock data, randomized inputs, edge-case permutations) is exactly what AI is good at, and doing it well doesn't require the model to have written the implementation.&lt;/li&gt;
&lt;li&gt;Human review time is best spent on boundary conditions specifically: zero, negative, empty, duplicate, malformed, concurrent, because that's statistically where both human and AI implementations tend to fail first.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't really a testing problem. It's a correlated-error problem wearing a green checkmark. Two independent reviewers catch different mistakes because they're independent. One reviewer checking its own work twice catches the same mistakes it already missed, twice.&lt;/p&gt;

&lt;p&gt;Where's your line? Do you let AI touch your test assertions at all, or only the scaffolding, mocks, fixtures, input generation, around them?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>software</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Semantic caching isn't a cost-saving hack. It's an admission that most "AI features" are FAQ bots in disguise.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Tue, 01 Sep 2026 09:10:45 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/semantic-caching-isnt-a-cost-saving-hack-its-an-admission-that-most-ai-features-are-faq-bots-93j</link>
      <guid>https://dev.to/cyclopt_dimitrisk/semantic-caching-isnt-a-cost-saving-hack-its-an-admission-that-most-ai-features-are-faq-bots-93j</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;There's a pitch behind every new AI-powered feature: it understands anything a user throws at it. Open-ended, flexible, genuinely intelligent.&lt;/p&gt;

&lt;p&gt;The pitch is half true. What it leaves out is that in most production systems, "anything a user throws at it" collapses into the same twenty questions asked twenty different ways. A support bot doesn't get asked infinite variations of physics, it gets asked "what's your return policy," "how do I get a refund," and "can I send this back" on a loop, forever.&lt;/p&gt;

&lt;p&gt;That gap, between the imagined variety of user queries and the actual repetition underneath it, is where the token bill lives.&lt;/p&gt;

&lt;h3&gt;
  
  
  The naive approach: every request hits the model
&lt;/h3&gt;

&lt;p&gt;Say a team ships a support assistant. Every user message goes straight to a frontier model: embed the question, retrieve context, generate an answer, return it. It works. It's also expensive in a very specific way, identical intent, paid for from scratch, every single time.&lt;/p&gt;

&lt;p&gt;(Illustrative, not a specific case, but recognizable to anyone who's watched an LLM API bill by day one versus day thirty.)&lt;/p&gt;

&lt;h3&gt;
  
  
  Exact-match caching almost works, then doesn't
&lt;/h3&gt;

&lt;p&gt;The obvious fix is a cache: hash the incoming query, check for a hit, skip the model call if found. This works exactly once, for the exact same string. "How do I reset my password" and "forgot my password, help" are the same question and two different cache keys. Exact-match caching is a solution for a problem language doesn't actually have, users asking things identically.&lt;/p&gt;

&lt;h3&gt;
  
  
  Embeddings turn paraphrase into a solvable problem
&lt;/h3&gt;

&lt;p&gt;Semantic caching replaces the hash with a vector. Embed the incoming query, run a nearest-neighbor search against previously answered queries, and if the closest match clears a similarity threshold, serve the cached response instead of calling the model at all. Below the threshold, call the model as normal and add the new query/response pair to the cache for next time.&lt;/p&gt;

&lt;p&gt;This is the actual mechanism at work: not "the AI understood the question was similar," but a distance calculation in vector space with a cutoff you chose.&lt;/p&gt;

&lt;h3&gt;
  
  
  A stale cache is worse than no cache
&lt;/h3&gt;

&lt;p&gt;Here's the failure mode nobody budgets for. A support answer gets cached in March. The return policy changes in April. The cache doesn't know that, it just has a vector close enough to keep matching, and it will confidently serve outdated information for as long as the entry lives. A cache without an expiration policy isn't a cost optimization, it's a bug that pays for itself to keep running.&lt;/p&gt;

&lt;h3&gt;
  
  
  Threshold tuning is the whole game
&lt;/h3&gt;

&lt;p&gt;Set the similarity threshold too loose and you'll serve March's return policy answer to an April refund question, because the two "look" adjacent in vector space even though the real answer changed. Set it too tight and you get single-digit hit rates, because no two users phrase anything identically. There's no universal number here, only a number you have to measure against your own traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Build it yourself, or don't
&lt;/h3&gt;

&lt;p&gt;If you're at the scale where cache hit rate is a line item on your infra bill, and your domain has stable, well-defined intents (support, FAQ, internal tooling), rolling your own semantic cache with something like pgvector or Redis is a weekend project, not a research problem. If your queries are genuinely open-ended and low-repeat (creative generation, one-off analysis), you're optimizing for a hit rate that doesn't exist, and the caching layer is wasted engineering effort.&lt;/p&gt;

&lt;p&gt;Semantic caching isn't really an AI technique. It's ordinary distributed-systems caching (TTLs, invalidation, cache keys) with an embedding standing in for the key. The AI part was never the hard part.&lt;/p&gt;

&lt;p&gt;We ran into a version of this question building Cyclopt Companion's analyzers: a one-line diff shouldn't necessarily trigger a full re-analysis from scratch, and figuring out what counts as "close enough to skip" turned out to be the actual engineering problem, not the analysis itself.&lt;/p&gt;

&lt;p&gt;What's your current cache hit rate on user-facing LLM calls, and have you actually measured it, or is it a guess?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>programming</category>
      <category>discuss</category>
    </item>
    <item>
      <title>AI autocomplete isn't a productivity tool. It's a judgment test you take every few seconds.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Fri, 28 Aug 2026 06:52:03 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/ai-autocomplete-isnt-a-productivity-tool-its-a-judgment-test-you-take-every-few-seconds-5anl</link>
      <guid>https://dev.to/cyclopt_dimitrisk/ai-autocomplete-isnt-a-productivity-tool-its-a-judgment-test-you-take-every-few-seconds-5anl</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;There's a pitch behind every AI coding assistant: it makes you faster. Fewer keystrokes, less boilerplate, more shipped features per sprint.&lt;/p&gt;

&lt;p&gt;The pitch is half true. What it leaves out is the gap between a tutorial demo and a real codebase under real pressure. In a demo, every suggestion is correct because the demo was built to make the suggestion look correct. In production, the assistant doesn't know your architecture, your team's conventions, or the ticket you're actually trying to close. It just knows what tends to come next in code that looks like yours.&lt;/p&gt;

&lt;p&gt;That gap is where the noise lives.&lt;/p&gt;

&lt;h3&gt;
  
  
  The instant-accept trap
&lt;/h3&gt;

&lt;p&gt;Say a developer is mid-flow, wiring up a new endpoint. The assistant suggests a validation helper that looks reasonable, so they hit tab. It compiles, tests pass, they move on.&lt;/p&gt;

&lt;p&gt;Three weeks later a teammate finds two nearly identical validation helpers in the codebase: one written by a human eight months ago, one autocompleted last sprint. Nobody meant to duplicate logic. The suggestion was locally correct and globally redundant, and nothing about "correct code that compiles" caught that.&lt;/p&gt;

&lt;p&gt;(This is an illustrative scenario, not a specific incident, but most teams running Copilot or similar tools for more than a few months will recognize the shape of it.)&lt;/p&gt;

&lt;h3&gt;
  
  
  Architecture creep, one suggestion at a time
&lt;/h3&gt;

&lt;p&gt;No single autocompleted line breaks your architecture. That's exactly the problem. An assistant trained on generic patterns will happily suggest a new abstraction, a new dependency, a new way of doing something you already do three other ways elsewhere in the codebase, because it has no visibility into "elsewhere." Accept enough of these one at a time and the codebase drifts into a dozen small dialects of the same idea, none of them wrong in isolation.&lt;/p&gt;

&lt;h3&gt;
  
  
  The review tax
&lt;/h3&gt;

&lt;p&gt;The real cost isn't the code that's obviously bad, that gets caught. It's the code that's plausible enough to pass a quick glance and wrong enough to need real review time later. If you accept every suggestion without evaluating it against the code you already have, you're not saving time, you're deferring the thinking to code review, or worse, to whoever debugs it in production. Teams that measure this honestly often find they're spending more time reviewing and pruning generated code than they would have spent writing the smaller, more deliberate version themselves.&lt;/p&gt;

&lt;h3&gt;
  
  
  It's not an AI problem, it's a taste problem
&lt;/h3&gt;

&lt;p&gt;Strip away the tooling and this isn't new. Junior engineers have always generated more code than senior engineers, because judgment about what not to write is a skill that takes time to build. AI assistants didn't invent that gap, they just made it faster to fall into, because the suggestion arrives before you've had time to ask whether you need it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Autocomplete on, or on-demand only?
&lt;/h3&gt;

&lt;p&gt;There's no universal right answer here, but there is a decision worth making deliberately instead of by default:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Always-on autocomplete works if you already have strong instincts for when to reject a suggestion, and you treat every accepted line as your own code, not the tool's.&lt;/li&gt;
&lt;li&gt;On-demand AI for specific tasks (boilerplate, tests, migrations) forces a moment of intent before you invoke it, which some teams find keeps architecture decisions in human hands.&lt;/li&gt;
&lt;li&gt;Research and planning only keeps AI out of the diff entirely and uses it to think faster, not to type faster. (We lean on this mode a lot while building Cyclopt Companion, mostly to keep the assistant out of the actual diff until a human has decided the shape of the change.)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are wrong. They're trade-offs between speed and friction, and the right one depends on how much you trust your own review discipline in the moment.&lt;/p&gt;

&lt;p&gt;So: do you keep autocomplete on at all times, use AI only on-demand for specific tasks, or limit it to research and planning? And has your answer changed since you started using these tools?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>discuss</category>
    </item>
    <item>
      <title>AI-Generated Code Doesn't Have a Readability Problem. It Has a Style-Enforcement Problem Wearing an AI Hat.</title>
      <dc:creator>Dimitris Kyrkos </dc:creator>
      <pubDate>Tue, 25 Aug 2026 13:18:22 +0000</pubDate>
      <link>https://dev.to/cyclopt_dimitrisk/ai-generated-code-doesnt-have-a-readability-problem-it-has-a-style-enforcement-problem-wearing-an-5age</link>
      <guid>https://dev.to/cyclopt_dimitrisk/ai-generated-code-doesnt-have-a-readability-problem-it-has-a-style-enforcement-problem-wearing-an-5age</guid>
      <description>&lt;h2&gt;
  
  
  Intro
&lt;/h2&gt;

&lt;p&gt;Ask an AI assistant to "handle the checkout flow" and it will hand you a working function in about four seconds. Validate the cart, call three different APIs, format a receipt, log the analytics event, catch whatever errors show up, all in one block, three levels of nesting deep. It runs. It passes the one test anyone bothered to write. And it is, structurally, a small readability crisis waiting for whoever opens that file next.&lt;/p&gt;

&lt;p&gt;That gap, between code that works and code a human can actually follow six months later, is where most of the real cost of AI-assisted development quietly accumulates. Here's where it actually shows up. (Illustrative composites drawn from common patterns, not specific incidents.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The function that does six things because nobody told it not to
&lt;/h2&gt;

&lt;p&gt;A generated function rarely sets out to do too much. It just keeps absorbing responsibility one prompt at a time: first it validates input, then someone asks it to also handle a retry, then to also log the outcome, then to also format the response for the frontend. Each addition looks reasonable in isolation. Read top to bottom six weeks later, it's a single function holding four unrelated jobs, and splitting it back apart means first reverse-engineering which lines belong to which job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Names that describe the prompt, not the domain
&lt;/h2&gt;

&lt;p&gt;Generated code tends to name things after the immediate task rather than the concept it represents: &lt;code&gt;data2&lt;/code&gt;, &lt;code&gt;tempResult&lt;/code&gt;, &lt;code&gt;processedItems&lt;/code&gt;, &lt;code&gt;handleStuff&lt;/code&gt;. None of these are wrong in the sense of breaking anything. They're wrong in the sense that a teammate reading the code six months from now has to run it mentally just to figure out what &lt;code&gt;tempResult&lt;/code&gt; actually holds, instead of the name just telling them.&lt;/p&gt;

&lt;h2&gt;
  
  
  A slightly different helper function, invented every time
&lt;/h2&gt;

&lt;p&gt;Ask a model to format a date in one file and it writes &lt;code&gt;formatDate&lt;/code&gt;. Ask it again in another file, in the same session even, and it might write &lt;code&gt;toDateString&lt;/code&gt;, doing almost the same thing with a slightly different edge case handled. Nothing is technically duplicated, so no linter flags it, but the codebase slowly fills with near-identical helpers that all do roughly the same job slightly differently, because each generation has no memory of what already exists two files over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Style that resets at every file boundary
&lt;/h2&gt;

&lt;p&gt;File A uses early-return guard clauses. File B nests every conditional three deep because that's what was statistically nearby in training for that particular pattern. Neither is "wrong" on its own, a reviewer skimming one file at a time won't necessarily flag it, but navigating the codebase starts to mean re-learning the local dialect every time you open a new file, because there isn't one.&lt;/p&gt;

&lt;p&gt;None of this is a code-quality problem in the traditional sense, where someone wrote something sloppy and a linter catches it. It's a missing style contract problem: nothing external is constraining the model toward the conventions your team already agreed on, so by default it reaches for whatever's statistically common across its training data, which is rarely what's locally correct for your codebase.&lt;/p&gt;

&lt;p&gt;So what do you actually do about it? One option is writing an exhaustive style guide and hoping every prompt includes enough of it, which doesn't scale past a few files and quietly rots the moment someone forgets to paste it. The other is treating readability as something enforced at the commit gate, the same way you'd enforce tests passing, independent of whether a human or a model wrote the line. Prompts don't have memory. Commits do. That's really the whole argument for moving readability enforcement from "hope the prompt was good" to "verify the commit is."&lt;/p&gt;

&lt;p&gt;How is your team actually enforcing readability on AI-generated code right now, a style guide baked into the prompt, a linter, code review, or honestly, nothing yet?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>software</category>
      <category>discuss</category>
    </item>
  </channel>
</rss>
