<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Danil Galeev</title>
    <description>The latest articles on DEV Community by Danil Galeev (@danilgaleev).</description>
    <link>https://dev.to/danilgaleev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4082182%2F23591c06-ef30-4c83-8711-aa1f37797bb6.png</url>
      <title>DEV Community: Danil Galeev</title>
      <link>https://dev.to/danilgaleev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/danilgaleev"/>
    <language>en</language>
    <item>
      <title>The agent finished. Who turns off the VM?</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Thu, 08 Oct 2026 19:24:22 +0000</pubDate>
      <link>https://dev.to/danilgaleev/the-agent-finished-who-turns-off-the-vm-5h9i</link>
      <guid>https://dev.to/danilgaleev/the-agent-finished-who-turns-off-the-vm-5h9i</guid>
      <description>&lt;p&gt;Suppose a coding agent opens a pull request at 6 p.m. The tests pass, a preview is running, and the reviewer has already logged off.&lt;/p&gt;

&lt;p&gt;The agent's task is complete. The machine still has something to do: serve that preview until someone looks at it. Or perhaps it doesn't. Perhaps the preview can be rebuilt tomorrow and the workspace should shut down now.&lt;/p&gt;

&lt;p&gt;Either choice is reasonable. Leaving it undecided is how a short coding task turns into a machine nobody remembers starting.&lt;/p&gt;

&lt;p&gt;Web Dev Cody's &lt;a href="https://www.youtube.com/watch?v=UA_HKbvG7k8" rel="noopener noreferrer"&gt;Coder walkthrough&lt;/a&gt; shows agents working in remote environments, including an AWS EC2 workspace, with changes and previews available for review. The video is sponsored by Coder. It prompted a narrower question for me: what should stop when an agent finishes?&lt;/p&gt;

&lt;p&gt;There are three separate answers: the command, the agent session, and the workspace. A limit on one doesn't automatically control the others.&lt;/p&gt;

&lt;h2&gt;
  
  
  A test timeout only covers the test
&lt;/h2&gt;

&lt;p&gt;Here is a small guardrail for a Node.js project. It requires Bash, GNU &lt;code&gt;timeout&lt;/code&gt;, installed dependencies, and a working &lt;code&gt;npm test&lt;/code&gt; script. Run it from the project directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="nb"&gt;timeout&lt;/span&gt; &lt;span class="nt"&gt;--kill-after&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;30s 10m npm &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;test.log 2&amp;gt;&amp;amp;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'Tests passed. Output: test.log\n'&lt;/span&gt;
&lt;span class="k"&gt;else
  &lt;/span&gt;&lt;span class="nv"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$?&lt;/span&gt;
  &lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'Tests failed or timed out (exit %s). Output: test.log\n'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;exit&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At ten minutes, &lt;code&gt;timeout&lt;/code&gt; sends a termination signal. If the command still hasn't exited thirty seconds later, it sends a kill signal. A timeout normally returns &lt;code&gt;124&lt;/code&gt;; a forced kill can return &lt;code&gt;137&lt;/code&gt;. Otherwise, the command's exit status is preserved. &lt;a href="https://www.gnu.org/software/coreutils/manual/html_node/timeout-invocation.html" rel="noopener noreferrer"&gt;GNU documentation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now imagine the agent reads the failure, changes something, and runs the tests again. Each attempt has a limit. The task as a whole can still continue indefinitely.&lt;/p&gt;

&lt;p&gt;That needs a separate deadline or retry budget, enforced by whatever launches and supervises the agent. Putting a timeout inside a script the agent can rewrite is useful for ordinary failures, but it isn't a reliable boundary for the agent itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  An active workspace may never become idle
&lt;/h2&gt;

&lt;p&gt;An inactivity timer sounds like an answer to the overnight-machine problem. Its usefulness depends on what counts as activity.&lt;/p&gt;

&lt;p&gt;Coder's &lt;a href="https://coder.com/docs/user-guides/workspace-scheduling" rel="noopener noreferrer"&gt;scheduling documentation&lt;/a&gt; says that a working agent can extend the workspace's shutdown deadline. This prevents a legitimate task from being interrupted. It also means that an agent stuck doing work may keep its environment alive.&lt;/p&gt;

&lt;p&gt;For the hypothetical 6 p.m. pull request, my policy would be: end the agent session when it hands over the result, save the review evidence, and retain the preview only until an explicit expiry time. A longer review window should be a deliberate extension.&lt;/p&gt;

&lt;p&gt;The supervisor should also handle the less tidy cases: the agent crashes, loses its connection, or keeps retrying without producing a result. A deadline that depends on the agent successfully reporting completion won't cover those failures.&lt;/p&gt;

&lt;p&gt;This is a policy to implement and test, not a claim that every agent platform provides the same controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the evidence before removing the workspace
&lt;/h2&gt;

&lt;p&gt;Shutdown is awkward when the only copy of a change or its test output lives on the machine being stopped.&lt;/p&gt;

&lt;p&gt;For a task using a Git worktree, the setup might look like this. Run it from an existing checkout, with a new branch name and an unused destination:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;task_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"issue-142"&lt;/span&gt;
&lt;span class="nv"&gt;task_dir&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"../agent-&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;task_id&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

git worktree add &lt;span class="nt"&gt;-b&lt;/span&gt; &lt;span class="s2"&gt;"agent/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;task_id&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$task_dir&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; HEAD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This starts from committed &lt;code&gt;HEAD&lt;/code&gt;; it doesn't copy uncommitted edits. A worktree separates working files while sharing repository data. It doesn't isolate credentials, processes, or cloud resources. &lt;a href="https://git-scm.com/docs/git-worktree" rel="noopener noreferrer"&gt;Git documentation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Before cleanup, preserve the commit remotely and attach the test result to the pull request. Include the command that ran and any failures. If a reviewer needs the preview, record which commit it serves and when it expires.&lt;/p&gt;

&lt;p&gt;After review, inspect the worktree from the original checkout. Only remove it once its work and useful logs are preserved:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git &lt;span class="nt"&gt;-C&lt;/span&gt; ../agent-issue-142 status &lt;span class="nt"&gt;--short&lt;/span&gt;

&lt;span class="c"&gt;# Run only after checking the status and preserving the work.&lt;/span&gt;
git worktree remove ../agent-issue-142
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without &lt;code&gt;--force&lt;/code&gt;, Git refuses removal if uncommitted changes or untracked files remain. That includes the &lt;code&gt;test.log&lt;/code&gt; from our earlier example. A clean status doesn't prove the commits were pushed, so check that separately.&lt;/p&gt;

&lt;p&gt;Removing the worktree leaves the branch and cloud resources in place. Even stopping an EC2 instance is different from terminating it. The infrastructure needs its own cleanup step and retention rules. &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/Stop_Start.html" rel="noopener noreferrer"&gt;EC2 documentation&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Try one failed task before running ten agents
&lt;/h2&gt;

&lt;p&gt;Before expanding this setup, deliberately run a task that cannot finish. Let its deadline expire. Then check whether the agent stopped, whether the logs survived, and which resources remain active.&lt;/p&gt;

&lt;p&gt;Repeat with a successful task whose reviewer is unavailable until the next day. The preview should remain available for the promised window without requiring the agent to keep working.&lt;/p&gt;

&lt;p&gt;Those two cases would tell me more about readiness than another successful demo. I want to be able to close the laptop knowing what will still be running tomorrow—and why.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>cloud</category>
      <category>automation</category>
    </item>
    <item>
      <title>The AI feature nobody is shipping: a system that asks you questions</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Sun, 04 Oct 2026 10:42:22 +0000</pubDate>
      <link>https://dev.to/danilgaleev/the-ai-feature-nobody-is-shipping-a-system-that-asks-you-questions-22mg</link>
      <guid>https://dev.to/danilgaleev/the-ai-feature-nobody-is-shipping-a-system-that-asks-you-questions-22mg</guid>
      <description>&lt;h1&gt;
  
  
  The AI feature nobody is shipping: a system that asks you questions
&lt;/h1&gt;

&lt;p&gt;Open any AI product and you get the same shape: you type a task, it executes, you get a result. Confident and fast. It's also missing the one thing that would have made it correct, which is a question.&lt;/p&gt;

&lt;p&gt;I've spent the last year building agents for reporting, CRM, and support workflows. What keeps breaking isn't the model. It's the assumption that the person in front of it already knows how to describe what they want.&lt;/p&gt;

&lt;h2&gt;
  
  
  The execution trap
&lt;/h2&gt;

&lt;p&gt;We've optimized hard for execution. Give the model a clear task and it will do it, often better than the person who asked. That's real, and it's also where most teams stop thinking.&lt;/p&gt;

&lt;p&gt;The trouble starts the moment the request is underspecified, which is most of the time. A real example from my own life. I asked an assistant to sort the email from my kids' schools. Three kids, three schools, a flood of messages about events, forms, and payments. The system said "task received" and started categorizing. It never asked who actually manages the schedule. It never asked whether the information matters differently for a younger child and a teenager. It never asked who picks the kids up on which day.&lt;/p&gt;

&lt;p&gt;I got back exactly the task I typed. Not the task I had.&lt;/p&gt;

&lt;p&gt;The person who actually owns that calendar is not me. They read those emails first, they know the schedule, they handle the forms. Any system that wants to help with school logistics has to know that, and it can't know it unless it asks. Instead it treated my typed sentence as the whole context and produced a tidy, useless summary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why nobody ships the question
&lt;/h2&gt;

&lt;p&gt;Asking is harder than answering, for reasons that are mostly economic.&lt;/p&gt;

&lt;p&gt;An assistant that asks four questions before acting feels slower than one that answers instantly. Product teams optimize for the demo, and the demo rewards a fast, confident answer. A model that replies "before I do this, tell me who owns the schedule" looks like a worse product in a 30-second clip, even when it's the better one.&lt;/p&gt;

&lt;p&gt;There's a cost angle too. Every clarifying turn spends tokens and time. At the top of the market that's real money: serious agent work runs into tens of thousands of dollars a month, and a single hour of a strong model working a hard task is priced like a junior contractor. Asking questions multiplies the number of turns, so the incentive is to guess and move on.&lt;/p&gt;

&lt;p&gt;There's a safety angle vendors rarely say out loud. A system that interrogates you is a system that collects your context. Once it knows who owns the calendar, where the kids go, and what you care about, it holds a profile of your life. That's the data that makes the product valuable, and the data that makes it a liability.&lt;/p&gt;

&lt;p&gt;So the current generation of models does the easy half well. It executes, and it rarely asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the asking model wins
&lt;/h2&gt;

&lt;p&gt;The lab that ships a model that genuinely interrogates you takes the market, and I don't think it's close.&lt;/p&gt;

&lt;p&gt;Execution is turning into a commodity. Every serious model can now take a clear instruction and carry it out. What none of them do well is build the picture of your world that makes the instruction meaningful. The model that asks good questions, and remembers the answers, compounds. Each conversation makes the next one better. You can't copy that by releasing a slightly smarter model.&lt;/p&gt;

&lt;p&gt;If even a fraction of the people already using AI assistants were asked a few good questions and had the answers stored, you'd have hundreds of millions of filled-in profiles no competitor can replicate. The value sits in the memory of who the user is, not in the model weights.&lt;/p&gt;

&lt;p&gt;Today's models still wait for you to volunteer context. They treat a short prompt as a complete specification. The gap between "it did what I said" and "it understood what I meant" is where the next product cycle lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to build it
&lt;/h2&gt;

&lt;p&gt;If you're building an agent, the question step isn't a nice-to-have. It's a node in your graph, and it deserves the same care as your tool calls.&lt;/p&gt;

&lt;p&gt;A rough shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;intake ──&amp;gt; assess_context
             │
             ├── context sufficient ──&amp;gt; plan ──&amp;gt; execute ──&amp;gt; respond
             │
             └── context missing ──&amp;gt; ask ──&amp;gt; (wait for answer) ──&amp;gt; assess_context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two rules I'd put on that node:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Ask before acting when the stakes are non-trivial.&lt;/strong&gt; A password reset is fine to just do. Anything touching money, access, schedules, or other people should trigger a question first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persist what you learn.&lt;/strong&gt; A question you ask once and forget is worse than no question at all. The answer belongs in durable memory, keyed to the user, and it should change how the next request is handled.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The loop is the easy part. Deciding what to ask is the hard part, and it's a product problem, not a model problem. The best questions resolve the most ambiguity with the least friction: who owns this, what does "done" look like, what should I never do without asking. Get those three right and the rest follows.&lt;/p&gt;

&lt;h2&gt;
  
  
  The wider shift
&lt;/h2&gt;

&lt;p&gt;The asking model is one piece of a larger change in how I build with AI.&lt;/p&gt;

&lt;p&gt;Stop chasing the word "agent." It's too abstract, and it pushes you to optimize one small piece of a system instead of seeing the whole thing. "An agent that spawned 8,500 sub-agents" sounds impressive and means nothing. I expect the term to fade, the way "prompt engineering" is already fading. Teaching someone to write prompts today is like teaching them to change the screen resolution on an operating system from a decade ago. The models are being built to read your context themselves.&lt;/p&gt;

&lt;p&gt;Rebuild the process around the model instead of bolting the model onto the old process. If you're adding an AI module inside a CRM that exists only because software used to be hard to change, you're automating a container that may not need to exist. The useful question isn't how to add AI to this workflow. It's why this workflow has this shape at all.&lt;/p&gt;

&lt;p&gt;Pick one primary system and go deep, two at most. Scattering across five half-learned tools is how you end up fluent in none of them.&lt;/p&gt;

&lt;p&gt;Access to the strong models will stay expensive and limited. The best systems already sit with a handful of corporations and states. That's not a reason to wait. Build the memory layer, the part that's yours, while the model layer keeps getting rented from someone else.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I keep coming back to
&lt;/h2&gt;

&lt;p&gt;If I had to bet on one feature for the next year, it isn't a bigger context window or a faster model. It's an assistant that asks me three good questions before it touches anything, and remembers the answers forever.&lt;/p&gt;

&lt;p&gt;The model that asks is the model that understands. Everything else is execution, and execution is getting cheap.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;I build AI agents for reporting, CRM, and support workflows, with human review where it counts. Based in Potsdam, Germany.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I built an L1 ticket deflector that knows when to stop: LangGraph + human-in-the-loop</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Sat, 03 Oct 2026 22:57:52 +0000</pubDate>
      <link>https://dev.to/madebyexpert/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the-loop-4fd3</link>
      <guid>https://dev.to/madebyexpert/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the-loop-4fd3</guid>
      <description>&lt;h1&gt;
  
  
  I built an L1 ticket deflector that knows when to stop: LangGraph + human-in-the-loop
&lt;/h1&gt;

&lt;p&gt;The same tickets arrive at every IT service desk: password resets, VPN that won't connect, "how do I install 7-Zip?", a printer that's offline again. A widely cited industry range puts routine requests at 50–70% of L1 volume — work a script could handle, except nobody trusts a script with the tickets that actually matter.&lt;/p&gt;

&lt;p&gt;Building an agent that answers everything is easy. Building one that knows when to hand a ticket to a human is not. I built the second kind and open-sourced it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/zedxter/l1-ticket-deflector" rel="noopener noreferrer"&gt;github.com/zedxter/l1-ticket-deflector&lt;/a&gt; (MIT)&lt;/p&gt;

&lt;h2&gt;
  
  
  The design decision that shapes everything
&lt;/h2&gt;

&lt;p&gt;Most "AI helpdesk" demos optimize for deflection rate. That's the wrong target. A high deflection rate is trivial if you let the model guess.&lt;/p&gt;

&lt;p&gt;The constraint I optimized for instead: never let the agent act on a sensitive request. Anything touching access, privileges, security, or money goes to a human with prepared context. The agent resolves the boring 30% and escalates the rest — with the relevant KB article already attached, so the human doesn't start from zero.&lt;/p&gt;

&lt;p&gt;I called it human-in-the-loop, but the honest framing is narrower: the agent has a hard boundary, and that boundary is the part worth testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The graph
&lt;/h2&gt;

&lt;p&gt;The production version is a LangGraph state machine. Three outcomes, no cleverness:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;classify ──&amp;gt; no match / low confidence ──&amp;gt; queue (general L1 queue)
   │
   ├── sensitivity = low        ──&amp;gt; auto_resolve ──&amp;gt; notify_user
   └── sensitivity = high/crit  ──&amp;gt; human_review ──&amp;gt; escalate ──&amp;gt; notify_user
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The classifier returns a KB article ID and a confidence score. Two gates decide the path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Confidence gate&lt;/strong&gt; — below &lt;code&gt;0.55&lt;/code&gt;, the ticket goes to the general queue. No guessing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sensitivity gate&lt;/strong&gt; — high/critical tickets go to &lt;code&gt;service_desk_l2&lt;/code&gt; or &lt;code&gt;security_ops&lt;/code&gt;, never to auto-resolve.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Each KB article carries a &lt;code&gt;sensitivity&lt;/code&gt; label. That's the trick: the model doesn't decide what's dangerous. The knowledge base does, and it's a human-authored field. The LLM only picks the article.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route_after_classify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;article_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;CONFIDENCE_THRESHOLD&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;queue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sensitivity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;AUTO_RESOLVE_SENSITIVITY&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;auto_resolve&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;human_review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the entire safety story in six lines. Sensitive categories (access grants, license purchases, phishing reports, ransomware) can never reach &lt;code&gt;auto_resolve&lt;/code&gt;, no matter what the model says.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two versions, on purpose
&lt;/h2&gt;

&lt;p&gt;The repo ships two implementations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;demo/&lt;/code&gt;&lt;/strong&gt; — a zero-dependency Python script. Keyword matching plus stemming, no LLM. It runs on a Raspberry Pi. Its job is to make the mechanics visible and the numbers checkable.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;graph/&lt;/code&gt;&lt;/strong&gt; — the LangGraph agent. RAG over the KB, structured output via Pydantic, tool-calls into the ITSM, &lt;code&gt;MOCK_MODE=1&lt;/code&gt; so you can run the logic without an API key.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I kept the offline version because a demo you can't run is just a claim. Anyone can clone the repo and reproduce the numbers in ten seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers (and why I'm suspicious of them)
&lt;/h2&gt;

&lt;p&gt;On 45 labeled tickets, the offline classifier scored:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Auto-resolved&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Escalated to human&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No match → queue&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decision accuracy&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Routing accuracy&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sample deflection&lt;/td&gt;
&lt;td&gt;55.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;100% accuracy is a red flag, not a trophy. It means the test set is too small and too clean. 45 tickets, hand-labeled, written by the same person who wrote the KB — of course it scores well. I'm publishing it because it's honest about what it is: a sanity check, not a benchmark. The real test is a client's messy backlog, and that number doesn't exist yet.&lt;/p&gt;

&lt;p&gt;The deflection rate is the one I'd defend. 55.6% on this sample, but the ROI model uses a conservative &lt;strong&gt;30%&lt;/strong&gt;, because market benchmarks for first-year deflection sit at 20–40% and I'd rather be wrong in the client's favor.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ROI model
&lt;/h2&gt;

&lt;p&gt;Conservative assumptions for a ~500-person company in DACH:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;L1 tickets / month&lt;/td&gt;
&lt;td&gt;~1,200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Deflection&lt;/td&gt;
&lt;td&gt;30%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Avg handling time&lt;/td&gt;
&lt;td&gt;19 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IT rate (fully loaded)&lt;/td&gt;
&lt;td&gt;€45 / hour&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time saved&lt;/td&gt;
&lt;td&gt;~114 h/month (~0.7 FTE)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Savings&lt;/td&gt;
&lt;td&gt;~€5,130 / month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retainer&lt;/td&gt;
&lt;td&gt;€4,000 / month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ROI&lt;/td&gt;
&lt;td&gt;1.3× / month (15× / year)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The retainer is deliberately below the savings. A model where the vendor captures all the value is a model nobody signs. 1.3×/month is not a spectacular number — it's a defensible one, and defensible is what gets a pilot approved.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's fake and what's real
&lt;/h2&gt;

&lt;p&gt;I want to be precise about this, because "proof of work" usually means "screenshot of a demo."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real:&lt;/strong&gt; the classifier, the graph, the routing logic, the KB structure, the ROI math. Clone it, run it, read every line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stubbed:&lt;/strong&gt; the tool-calls. &lt;code&gt;itsm_create_ticket&lt;/code&gt;, &lt;code&gt;directory_reset_password&lt;/code&gt;, and &lt;code&gt;mdm_push_install&lt;/code&gt; return strings, not real ITSM actions. Wiring them to Jira Service Management or Zammad over MCP is a pilot task, not a demo task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not built:&lt;/strong&gt; the Slack/Teams approval UI, telemetry, multilingual KB. They're on the roadmap and I'm not pretending otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell anyone building this
&lt;/h2&gt;

&lt;p&gt;Three things I'd carry into the next version:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Put the safety boundary in data, not in the prompt.&lt;/strong&gt; A &lt;code&gt;sensitivity&lt;/code&gt; field on each KB article is auditable. "The model is instructed to be careful" is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship the runnable offline version.&lt;/strong&gt; It costs an afternoon and buys all the credibility.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Publish the suspicious number.&lt;/strong&gt; 100% accuracy invites the right question — "on how many tickets?" — and answering it honestly is worth more than a flattering metric.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The repo is MIT-licensed and the whole point is that you can pull it apart. If you run an IT desk and want to compare deflection on your own tickets, that's the interesting experiment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/zedxter/l1-ticket-deflector" rel="noopener noreferrer"&gt;github.com/zedxter/l1-ticket-deflector&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;I build AI agents for reporting, CRM, and support workflows — with human review where it counts. Based in Potsdam, Germany.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>langgraph</category>
      <category>llm</category>
      <category>python</category>
    </item>
    <item>
      <title>Updating Xbox controller firmware from Linux: the USB hotplug trap</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Tue, 29 Sep 2026 18:46:57 +0000</pubDate>
      <link>https://dev.to/danilgaleev/updating-xbox-controller-firmware-from-linux-the-usb-hotplug-trap-47l4</link>
      <guid>https://dev.to/danilgaleev/updating-xbox-controller-firmware-from-linux-the-usb-hotplug-trap-47l4</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fciy9t23ij6g214u0hxeb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fciy9t23ij6g214u0hxeb.jpg" alt="Xbox controller connected to a Linux host and Windows VM" width="799" height="336"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I wanted to update two Xbox controllers without booting my laptop into Windows. Both eventually went from &lt;strong&gt;5.9.2709.0 to 5.23.6.0&lt;/strong&gt;, with Xbox Accessories confirming &lt;strong&gt;Updated!&lt;/strong&gt; and then &lt;strong&gt;No update available&lt;/strong&gt; for each one.&lt;/p&gt;

&lt;p&gt;The update reset the USB connection, and getting the controller back into the VM took most of the work.&lt;/p&gt;

&lt;p&gt;The scripts and setup guide are in &lt;a href="https://github.com/zedxter/xbox-firmware-linux" rel="noopener noreferrer"&gt;zedxter/xbox-firmware-linux&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This uses Microsoft's official updater inside a Windows VM hosted on Linux. It does not implement a native Linux firmware flasher or distribute firmware files.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tested setup and limits
&lt;/h2&gt;

&lt;p&gt;The successful session used openSUSE Tumbleweed with SELinux enforcing, KVM, rootful Docker, Windows 10 22H2 and two Microsoft controllers with USB ID &lt;code&gt;045e:0b12&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I turned the helpers from this session into configurable scripts. The portable version has tests for event filtering, generated configuration and cleanup protections, but &lt;strong&gt;the complete refactored workflow has not yet been retested on hardware&lt;/strong&gt;. Windows 11 is the default for a new setup; that guest path also remains unvalidated. Firmware 5.23.6.0 was the version Accessories installed on these two controllers; other models may receive a different version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the controller disappeared after USB reset
&lt;/h2&gt;

&lt;p&gt;A firmware update can reset a controller's USB connection. The device address can change, and a bootloader may use a different product ID.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.qemu.org/docs/master/system/devices/usb.html" rel="noopener noreferrer"&gt;QEMU supports selecting a physical USB port&lt;/a&gt;, so I used a configuration shaped like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;-device usb-host,hostbus=3,hostport=1,id=xbox-controller
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The physical port &lt;code&gt;3-1&lt;/code&gt; is an example. The scripts detect the port you select and generate the arguments.&lt;/p&gt;

&lt;p&gt;After a reset, Linux could see the controller, and a fresh libusb context inside the container could see it too. The running QEMU process still could not reacquire it.&lt;/p&gt;

&lt;p&gt;The device node was accessible and had the correct SELinux label. QEMU's cached USB object was misleading: its existence did not prove a working connection.&lt;/p&gt;

&lt;p&gt;The evidence pointed toward missing hotplug events. &lt;a href="https://github.com/libusb/libusb/blob/master/libusb/os/linux_udev.c" rel="noopener noreferrer"&gt;libusb's Linux backend&lt;/a&gt; consumes processed udev notifications. A USB bind mount provides device access, but does not automatically provide those host events inside a separate network namespace.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/systemd/systemd/blob/main/src/libsystemd/sd-device/device-monitor.c" rel="noopener noreferrer"&gt;systemd's device monitor&lt;/a&gt; can disable its udev subscription when &lt;code&gt;/run/udev/control&lt;/code&gt; is absent and &lt;code&gt;/dev&lt;/code&gt; is not devtmpfs.&lt;/p&gt;

&lt;p&gt;I added a monitor marker and a small relay. It forwards original root-origin udev packets only for Microsoft USB device add/remove events on the selected physical port. It excludes USB interfaces and unrelated ports. The repository's architecture notes explain the checks and access boundaries.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: prepare the host
&lt;/h2&gt;

&lt;p&gt;You need a Linux x86-64 host with KVM, local rootful Docker Engine, Compose v2, Python 3.10+, systemd/udev, and a USB data cable. Allow room for a 4 GB Windows guest and preferably at least 80 GB of free storage. On SELinux hosts, the helpers also use &lt;code&gt;chcon&lt;/code&gt; and &lt;code&gt;restorecon&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Disconnect other wired Xbox controllers: preparation temporarily unloads their shared &lt;code&gt;xpad&lt;/code&gt; driver. Review the scripts before running them as root.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/zedxter/xbox-firmware-linux.git
&lt;span class="nb"&gt;cd &lt;/span&gt;xbox-firmware-linux
python3 scripts/configure.py &lt;span class="nt"&gt;--list&lt;/span&gt;

&lt;span class="c"&gt;# Replace 3-1 with the controller's actual physical port.&lt;/span&gt;
python3 scripts/configure.py &lt;span class="nt"&gt;--port&lt;/span&gt; 3-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Configuration creates a local Compose file, a random Windows password in an ignored &lt;code&gt;.env&lt;/code&gt;, and storage for the VM. No passwords or VM disks are included in the repository. Windows licensing remains your responsibility.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: start Windows and the relay
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;docker compose &lt;span class="nt"&gt;-f&lt;/span&gt; compose.local.json create
&lt;span class="nb"&gt;sudo &lt;/span&gt;python3 scripts/host_session.py prepare
&lt;span class="nb"&gt;sudo &lt;/span&gt;docker compose &lt;span class="nt"&gt;-f&lt;/span&gt; compose.local.json start
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If preparation fails, keep the container stopped and run &lt;code&gt;sudo python3 scripts/host_session.py cleanup&lt;/code&gt; before fixing the error and retrying.&lt;/p&gt;

&lt;p&gt;In another terminal, from the same directory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;systemd-inhibit &lt;span class="nt"&gt;--what&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;sleep&lt;/span&gt;:idle &lt;span class="nt"&gt;--mode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;block &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--who&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;XboxFirmwareUpdate &lt;span class="nt"&gt;--why&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'Xbox controller firmware update'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  python3 &lt;span class="nt"&gt;-u&lt;/span&gt; scripts/relay.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep that terminal open. Look for &lt;code&gt;READY&lt;/code&gt; and then &lt;code&gt;FORWARDED add&lt;/code&gt;. Keep the laptop on power with its lid open.&lt;/p&gt;

&lt;p&gt;Open &lt;code&gt;http://127.0.0.1:8006&lt;/code&gt; and let Windows install. The setup builds on &lt;a href="https://github.com/dockur/windows" rel="noopener noreferrer"&gt;dockur/windows&lt;/a&gt;; the repository pins the image used in the original session. The console is bound to localhost.&lt;/p&gt;

&lt;p&gt;The relay filters its events. The VM still has privileged USB access, so run it as a trusted maintenance tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: prepare Accessories, then test reconnect
&lt;/h2&gt;

&lt;p&gt;Install available Windows updates, finish guest restarts, and install Xbox Accessories from the Microsoft Store.&lt;/p&gt;

&lt;p&gt;We encountered three different blockers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Windows OS update required:&lt;/strong&gt; updating Windows resolved it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gaming Runtime Services sign-in failed:&lt;/strong&gt; installing the official Xbox app and signing in resolved our case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Device Manager reported a working controller, but Accessories did not show it:&lt;/strong&gt; uninstalling the device without deleting the driver package, followed by a Windows restart, restored discovery.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I only needed these fixes when the corresponding error appeared.&lt;/p&gt;

&lt;p&gt;Before starting any firmware update, unplug and reconnect the controller once to the &lt;strong&gt;same physical port&lt;/strong&gt;. Accessories must find it again without a Docker restart, and the relay should report an add event.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;python3 scripts/status.py
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that pre-flash test fails, stop and fix discovery first. A guest restart preserves the relay; restarting the Docker container requires starting a new relay.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: update one controller at a time
&lt;/h2&gt;

&lt;p&gt;Start the available update in Xbox Accessories. Do not disconnect USB, stop the relay, suspend Linux or restart Windows/Docker while updating.&lt;/p&gt;

&lt;p&gt;Wait for the explicit success screen. Then reopen the controller details and verify the full firmware version and update status.&lt;/p&gt;

&lt;p&gt;The first controller needed a Windows restart after the success screen before Accessories would read its new version. The second reported its new version immediately. Both were confirmed at &lt;strong&gt;5.23.6.0&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Only after verifying the first controller should you swap in the second, using the same cable and port. Repeat the reconnect test and update.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5: finish the session
&lt;/h2&gt;

&lt;p&gt;After every update is finished, shut Windows down through its Start menu. Wait for the container to exit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;docker inspect &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s1"&gt;'{{.State.Status}}'&lt;/span&gt; xbox-firmware-linux
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If Windows is fully shut down but Docker remains running, stop it with &lt;code&gt;sudo docker compose -f compose.local.json stop&lt;/code&gt;. Never use this to interrupt firmware updating: Docker can force-stop after its grace period.&lt;/p&gt;

&lt;p&gt;The relay exits with its container process. Restore the host settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;python3 scripts/host_session.py cleanup
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The cleanup command refuses a running VM. It removes matching temporary rules, restores the USB label and restores the previous &lt;code&gt;xpad&lt;/code&gt; loaded state. Reconnect the cable for Linux. The Windows disk is retained for later use and should remain private.&lt;/p&gt;

&lt;h2&gt;
  
  
  Firmware did not solve every Bluetooth problem
&lt;/h2&gt;

&lt;p&gt;After updating, the first controller could pair but still failed to reconnect after being turned off and on. A clean reboot and fresh pairing did not resolve it.&lt;/p&gt;

&lt;p&gt;On our MediaTek adapter with BlueZ 5.87, testing this option made reconnection work:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[LE]&lt;/span&gt;
&lt;span class="py"&gt;CentralAddressResolution&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;0&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a separate, adapter-wide workaround, not a universal Xbox setting. It does not delete pairing keys or disable encryption. Only the first controller's Bluetooth reconnect was independently verified. The repository includes &lt;a href="https://github.com/zedxter/xbox-firmware-linux/blob/main/docs/bluetooth.md" rel="noopener noreferrer"&gt;instructions and rollback guidance&lt;/a&gt;; the scripts do not apply it automatically.&lt;/p&gt;

&lt;p&gt;For general pairing cleanup after firmware changes, consult &lt;a href="https://github.com/atar-axis/xpadneo/blob/master/docs/TROUBLESHOOTING.md" rel="noopener noreferrer"&gt;xpadneo's troubleshooting guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Share reproducible results
&lt;/h2&gt;

&lt;p&gt;If you try the toolkit, report your controller model, host and guest versions, whether the pre-flash reconnect test passed, and the firmware version Accessories confirmed. Remove serial numbers, Bluetooth addresses and keys from logs.&lt;/p&gt;

&lt;p&gt;I would test USB rediscovery before clicking Update again. If the controller does not return to Accessories after a simple cable reconnect, a firmware reset is unlikely to go better.&lt;/p&gt;

</description>
      <category>linux</category>
      <category>docker</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Agent Memory's Real Failure Is Currency, Not Retrieval</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Wed, 23 Sep 2026 08:14:32 +0000</pubDate>
      <link>https://dev.to/madebyexpert/agent-memorys-real-failure-is-currency-not-retrieval-40h9</link>
      <guid>https://dev.to/madebyexpert/agent-memorys-real-failure-is-currency-not-retrieval-40h9</guid>
      <description>&lt;p&gt;Retrieval returns the correct record. The record is stale. And because the model states fresh and dead facts with identical confidence, nothing in the output tells you which one you got.&lt;/p&gt;

&lt;p&gt;Your agent's memory isn't broken because it can't find the fact. It's broken because the fact it found stopped being true.&lt;/p&gt;

&lt;h2&gt;
  
  
  That is not a retrieval problem
&lt;/h2&gt;

&lt;p&gt;It's a state-management problem — the same one distributed systems have always had: validity windows, invalidation, cache coherency. A memory layer with none of those is a cache that never expires.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the numbers say
&lt;/h2&gt;

&lt;p&gt;Memory is now a first-class layer with its own benchmarks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LoCoMo 92.5&lt;/li&gt;
&lt;li&gt;LongMemEval 94.4&lt;/li&gt;
&lt;li&gt;BEAM 1M 64.1&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;...at roughly 6,900 tokens per query on average. The biggest gains are in temporal reasoning (+29.6) and multi-hop (+23.1). And the open problems named in the same report are cross-session identity, temporal abstraction at scale — and staleness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stale facts get surfaced. They just don't get acted on.
&lt;/h2&gt;

&lt;p&gt;On the STALE benchmark — a separate 2026 study on stale-fact handling:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a system surfaced the updated fact about 77% of the time, but marked it as worth acting on only ~3%;&lt;/li&gt;
&lt;li&gt;one frontier model caught a stale fact 92% of the time when asked directly — and 30% when the question silently depended on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fact wasn't missing. It was present, and discounted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern that works
&lt;/h2&gt;

&lt;p&gt;Zep's Graphiti shows the shape of the fix: store facts with validity dates, mark contradicting facts invalid instead of overwriting them, and surface only the current version.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule
&lt;/h2&gt;

&lt;p&gt;Memory is not "accumulate more." A fact needs a bounded window of validity — otherwise you're shipping an agent that is confidently wrong on a schedule.&lt;/p&gt;

&lt;p&gt;How are you handling staleness in your agent's memory layer?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>database</category>
      <category>programming</category>
    </item>
    <item>
      <title>A weekly reporting workflow that stops when its sources are incomplete</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Wed, 16 Sep 2026 13:18:03 +0000</pubDate>
      <link>https://dev.to/madebyexpert/a-weekly-reporting-workflow-that-stops-when-its-sources-are-incomplete-5abl</link>
      <guid>https://dev.to/madebyexpert/a-weekly-reporting-workflow-that-stops-when-its-sources-are-incomplete-5abl</guid>
      <description>&lt;p&gt;A weekly report can contain plausible numbers and still be unsuitable for a decision. A missing support export or an old CRM snapshot changes what the report can honestly say.&lt;/p&gt;

&lt;p&gt;I built a &lt;a href="https://madeby.expert/demo.html" rel="noopener noreferrer"&gt;small reporting demo for MadeBy.Expert&lt;/a&gt; to make those conditions visible. It uses synthetic billing, CRM and support snapshots, deterministic calculations and a template narrative. There are no live integrations or language model calls in this version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define the numbers before automating the report
&lt;/h2&gt;

&lt;p&gt;The sample distinguishes monthly recurring revenue from revenue collected during the week. Open pipeline is the unweighted value of open opportunities, not a revenue forecast. Support metrics distinguish unresolved tickets at the snapshot from tickets resolved within the reporting period.&lt;/p&gt;

&lt;p&gt;Each metric includes a source reference and a definition. Those details give the reviewer something concrete to check against the input.&lt;/p&gt;

&lt;p&gt;For a client workflow, I would agree on these definitions before connecting the source systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  A missing source is a report state
&lt;/h2&gt;

&lt;p&gt;The browser demo lets you switch between complete sources, missing support data and stale CRM data. The report keeps the metrics it can calculate, omits the unavailable ones and records a blocker. It does not substitute zero for a missing value.&lt;/p&gt;

&lt;p&gt;The final state comes from that blocker list:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;blockers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;blocked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;draft&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For this example, all three sources are required for approval. A snapshot is stale if it was captured before the reporting period ended or more than 24 hours before the report's &lt;code&gt;asOf&lt;/code&gt; time. A future timestamp is rejected as invalid input.&lt;/p&gt;

&lt;p&gt;Those are choices for this workflow. A daily operations report might need a tighter freshness limit; another report might be useful with an explicitly optional source. The thresholds need agreement with the person using the result.&lt;/p&gt;

&lt;p&gt;Here is the actual test covering all three sources and both failure modes. It uses Node's built-in test runner and strict assertions. &lt;code&gt;fixture()&lt;/code&gt; loads the synthetic input afresh for each case:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;missing/stale sources are unavailable and block approval&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;kind&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;billing&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;crm&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;support&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;mode&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;missing&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;stale&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fixture&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;mode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;missing&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;delete&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
      &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;capturedAt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;2026-09-12T08:00:00Z&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;buildReport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nx"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;equal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;blocked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nx"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;blockers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startsWith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)));&lt;/span&gt;
      &lt;span class="nx"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;equal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;metrics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sourceId&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nf"&gt;fixture&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nx"&gt;sources&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;kind&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="nx"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;throws&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;approveReport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Demo reviewer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sr"&gt;/Blocked/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The last assertion matters: reporting the problem in the output is insufficient if the approval function still accepts it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Approval applies to one version of the report
&lt;/h2&gt;

&lt;p&gt;The implementation computes a SHA-256 digest of the report content, which includes a digest of the input. A reviewer supplies the report digest when approving a local export. The approval function recomputes it and checks that the report is still a draft without blockers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;approveReport&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;report&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;reviewer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toISOString&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;stored&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;report&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;hash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;stored&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;digest&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;stored&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;fail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Approval digest does not match the current report&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;report&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;draft&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;report&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;blockers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;fail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Blocked reports cannot be approved&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;reviewer&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="nx"&gt;reviewer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="nx"&gt;reviewer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
    &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;[\r\n]&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;reviewer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;fail&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;A named reviewer is required&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;report&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;approved&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;approval&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;reviewer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;reviewer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt; &lt;span class="na"&gt;at&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;digest&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here, &lt;code&gt;hash&lt;/code&gt; is SHA-256 over &lt;code&gt;JSON.stringify(value)&lt;/code&gt; and &lt;code&gt;fail&lt;/code&gt; throws an error. This is an excerpt from the local prototype, not a standalone approval service.&lt;/p&gt;

&lt;p&gt;Rebuilding the report after changing the input produces a different digest, so the previous digest cannot approve the new report. The tests also alter a calculated metric directly and check that approval rejects the modified content.&lt;/p&gt;

&lt;p&gt;There are limits to this approach. The digest is a consistency check, not a signature or proof of who reviewed the report. The reviewer name is supplied as text; there is no authentication or authorization layer. The JSON serialization is also specific to this implementation, not a canonical format for independent producers.&lt;/p&gt;

&lt;p&gt;A deployed approval service would need to establish the reviewer's identity and permissions, store the review event, and bind the delivery action to the approved version. Those parts are not implemented here.&lt;/p&gt;

&lt;p&gt;The browser page shows the data conditions and report. The local runnable example handles approval and export; neither sends a report automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where AI could fit later
&lt;/h2&gt;

&lt;p&gt;A language model might help draft commentary from validated facts. That is a possible extension, not an implemented feature of this demo. I would keep metric calculation, source checks and the approval boundary explicit regardless of how the narrative is produced.&lt;/p&gt;

&lt;p&gt;The current prototype makes no claim about client time savings, live connector reliability or model quality. Those need measurement against an actual workflow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://madeby.expert/demo.html" rel="noopener noreferrer"&gt;Try the demo&lt;/a&gt;, including the missing-source and stale-source cases.&lt;/p&gt;

&lt;p&gt;If your team assembles a similar report manually, &lt;a href="https://madeby.expert/#contact" rel="noopener noreferrer"&gt;describe the workflow&lt;/a&gt;: which tools supply the data, how often the report is needed, and who checks it. We can discuss whether a small paid pilot with clear success criteria fits.&lt;/p&gt;

</description>
      <category>automation</category>
      <category>architecture</category>
      <category>javascript</category>
      <category>testing</category>
    </item>
    <item>
      <title>Typed Contracts Between AI Agents: The Interface That Actually Breaks</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Tue, 15 Sep 2026 05:44:04 +0000</pubDate>
      <link>https://dev.to/madebyexpert/typed-contracts-between-ai-agents-the-interface-that-actually-breaks-547j</link>
      <guid>https://dev.to/madebyexpert/typed-contracts-between-ai-agents-the-interface-that-actually-breaks-547j</guid>
      <description>&lt;p&gt;Most multi-agent failures I've seen don't happen &lt;em&gt;inside&lt;/em&gt; an agent. They happen &lt;strong&gt;between&lt;/strong&gt; two of them.&lt;/p&gt;

&lt;p&gt;When the handoff between two agents is treated as loosely structured JSON or implicit shared state, the interface is underspecified — and underspecified interfaces fail quietly.&lt;/p&gt;

&lt;p&gt;A production handoff should be treated like any other API contract:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an explicit input schema;&lt;/li&gt;
&lt;li&gt;an explicit output schema;&lt;/li&gt;
&lt;li&gt;required and optional fields;&lt;/li&gt;
&lt;li&gt;semantic constraints, not only valid JSON;&lt;/li&gt;
&lt;li&gt;a clear rejection and fallback path.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A concrete example
&lt;/h2&gt;

&lt;p&gt;A research agent should not hand an unstructured object to a scoring agent. It should emit something the scorer can actually validate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;ResearchResult&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;subject_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;
    &lt;span class="n"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;EvidenceItem&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;min_length&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;source_count&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ge&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;confidence&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ge&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;le&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scorer runs only after that result passes validation. That gate is the important part.&lt;/p&gt;

&lt;h2&gt;
  
  
  Probabilistic output, consequential action
&lt;/h2&gt;

&lt;p&gt;An LLM output is probabilistic. A database write, an external message, or a business decision is an action. Those two steps should not be directly coupled.&lt;/p&gt;

&lt;p&gt;Between them, put a deterministic gate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Parse the model output.&lt;/li&gt;
&lt;li&gt;Validate its schema and business rules.&lt;/li&gt;
&lt;li&gt;Reject, retry, route to review, or stop the branch when validation fails.&lt;/li&gt;
&lt;li&gt;Execute the action only on an accepted contract.&lt;/li&gt;
&lt;li&gt;Record the handoff version and validation result.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Why it changes failure behavior
&lt;/h2&gt;

&lt;p&gt;Without a contract, a missing field becomes a null value, the null becomes a misleading score, and the score becomes a bad CRM update. The symptom shows up several steps away from the cause, and by then the trail is cold.&lt;/p&gt;

&lt;p&gt;With contracts, the same failure is localized: &lt;code&gt;ResearchAgent -&amp;gt; ScoringAgent&lt;/code&gt;, field &lt;code&gt;source_count&lt;/code&gt;, validation rule failed.&lt;/p&gt;

&lt;p&gt;That is more than defensive programming. It is architecture for software systems whose internal steps are non-deterministic.&lt;/p&gt;

&lt;p&gt;It gives staff engineers independently testable boundaries, helps principal engineers reduce coupling between teams and agents, and gives architects and CTOs a clearer way to reason about compatibility, ownership, and recovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Every probabilistic step that feeds a consequential action needs an explicit validation gate. Typed contracts are one strong way to define that boundary.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model can be creative inside the boundary. The boundary should not be.&lt;/p&gt;

&lt;p&gt;Where have you found the most valuable contract in an agent workflow — between agents, between an agent and a tool, or before the final action?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>python</category>
      <category>programming</category>
    </item>
    <item>
      <title>Building AI Agents That Actually Work: What Nobody Tells You</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Fri, 04 Sep 2026 09:11:40 +0000</pubDate>
      <link>https://dev.to/madebyexpert/building-ai-agents-that-actually-work-what-nobody-tells-you-2f88</link>
      <guid>https://dev.to/madebyexpert/building-ai-agents-that-actually-work-what-nobody-tells-you-2f88</guid>
      <description>&lt;p&gt;I've been building AI agents for a while now. Not chatbots. Not RAG demos. Real agents that take actions, make decisions, and run autonomously.&lt;/p&gt;

&lt;p&gt;Here's what nobody tells you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap between demo and production
&lt;/h2&gt;

&lt;p&gt;Every AI agent framework shows you a 5-line demo that works perfectly. Then you deploy it, and it fails in ways you didn't imagine.&lt;/p&gt;

&lt;p&gt;The reason is simple: a demo is a happy path. Production is a graph of failure states.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Observability is not optional
&lt;/h2&gt;

&lt;p&gt;Your agent is only as good as your ability to see what it's doing.&lt;/p&gt;

&lt;p&gt;When an agent makes a wrong decision, you need to know exactly why. Was it a bad prompt? A hallucinated tool call? Missing context from a previous step?&lt;/p&gt;

&lt;p&gt;Log every thought. Audit every action. If you can't replay an agent's decision process, you can't trust it.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Tools are the bottleneck, not the LLM
&lt;/h2&gt;

&lt;p&gt;Most people think the LLM is the hard part. It's not. The hard part is the tools.&lt;/p&gt;

&lt;p&gt;Your agent needs to call APIs, read databases, write files, send emails. Each of those is a failure point. Network timeout. Auth expired. Schema changed. Rate limited.&lt;/p&gt;

&lt;p&gt;Build your tool layer like you build a distributed system. Retries. Circuit breakers. Timeouts. Graceful degradation.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The context window is a trap
&lt;/h2&gt;

&lt;p&gt;Long-running agents accumulate context. The more they do, the more context they carry. Eventually, the context window fills with noise, and the agent starts making bad decisions.&lt;/p&gt;

&lt;p&gt;Strategies that actually work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Summarization&lt;/strong&gt; — compress old context into summaries, don't carry raw history&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Structured memory&lt;/strong&gt; — separate short-term (current task) from long-term (learned patterns)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forgetting&lt;/strong&gt; — actively prune irrelevant context. If it didn't matter in the last 10 steps, it probably won't matter now&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Determinism is underrated
&lt;/h2&gt;

&lt;p&gt;Everyone wants creative agents. What you actually want is predictable agents.&lt;/p&gt;

&lt;p&gt;A creative agent that hallucinates a solution is useless. A predictable agent that follows a known pattern is valuable.&lt;/p&gt;

&lt;p&gt;Design for determinism first. Add creativity as a controlled parameter.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The orchestration layer matters more than the model
&lt;/h2&gt;

&lt;p&gt;You can swap GPT-4 for Claude or Gemini and your agent still works — if your orchestration is solid.&lt;/p&gt;

&lt;p&gt;Good orchestration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Clear state machine&lt;/li&gt;
&lt;li&gt;Explicit error handling&lt;/li&gt;
&lt;li&gt;Human-in-the-loop for critical decisions&lt;/li&gt;
&lt;li&gt;Audit trail for every action&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The uncomfortable truth
&lt;/h2&gt;

&lt;p&gt;Building AI agents that actually work is not an AI problem. It's a software engineering problem.&lt;/p&gt;

&lt;p&gt;The LLM is the easiest part. Everything around it — tools, observability, state management, error handling, orchestration — that's where the real work is.&lt;/p&gt;

&lt;p&gt;And that's also where the real value is.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>engineering</category>
    </item>
    <item>
      <title>Your Agent Isn't Losing Memory. It's Rotting.</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Fri, 28 Aug 2026 18:33:00 +0000</pubDate>
      <link>https://dev.to/madebyexpert/your-agent-isnt-losing-memory-its-rotting-2edd</link>
      <guid>https://dev.to/madebyexpert/your-agent-isnt-losing-memory-its-rotting-2edd</guid>
      <description>&lt;p&gt;Agents don't forget because the model is bad. They forget because we built it into them. Six hours and forty messages in, what an agent needs is gone — or still there but wrong, drowned in noise.&lt;/p&gt;

&lt;p&gt;That's context rot.&lt;/p&gt;

&lt;p&gt;I've spent the last month inside a team of always-on AI agents that run for weeks. Different jobs, one codebase, one shared vault of lessons. They review each other's PRs, publish, plan. And they rot. I watched every flavor of it, then built the countermeasures.&lt;/p&gt;

&lt;h2&gt;
  
  
  What context rot is
&lt;/h2&gt;

&lt;p&gt;Four failures that compound.&lt;/p&gt;

&lt;p&gt;Compression loss. When a session runs out of window, a summary replaces raw history. It sounds fine until the agent needs the detail it summarized away. I watched an agent repeat a plan that contradicted a decision two screens earlier. Nobody noticed, because the contradiction lived at a depth that was gone.&lt;/p&gt;

&lt;p&gt;Drift. Each turn pulls the agent toward the recent and the loud. Given ten pieces of context, it weights the last five. Over hundreds of turns it quietly solves a different problem than the one you assigned.&lt;/p&gt;

&lt;p&gt;Priority blurring. "We decided X" and "we looked at X" feel the same in the tokens. Decisions stop being binding and become suggestions.&lt;/p&gt;

&lt;p&gt;Noise accretion. Stale output and dead ends pile up. The signal-to-noise ratio falls, and a degraded ratio looks like a dumber model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part nobody warns you about
&lt;/h2&gt;

&lt;p&gt;The cheap fix is "just enlarge the window." I built that reflex out of myself. Longer windows don't fix rot, they defer it — later in the timeline, larger cost per turn, and past some point the model does visibly worse than with a lean window. Remembering everything is a form of forgetting: the important thing drowns.&lt;/p&gt;

&lt;p&gt;And rot is per-profile, not per-agent. In a team sharing a codebase, what agent A rotted away still matters to agent B. So the fix must live outside any single session. Memory in one agent's head dies with that session.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually works
&lt;/h2&gt;

&lt;p&gt;None of this is exotic.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Summarize to executable facts, not prose. Emit decisions and constraints as explicit, fixed-schema entries the agent can act on. A good summary answers "what is still binding" before "what happened."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Move memory out of the session into a semantic index. Our biggest win. Keep lessons and decisions in a plain markdown vault, index with embeddings, and search returns meaning, not exact strings. "Offer-notification" and "pet-booking" are unrelated strings and near-identical problems; string search never connects them, the graph does — in milliseconds, on one SQLite file and a small local embedding model. No vector DB, no cluster, no bill. Don't ask one long-lived brain to do all the retrieval.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Turn knowledge into skills, not notes. A lesson read once and never referenced is a lesson that rots. The durable form is a skill: a procedure loaded on demand behind a trigger, pulled into context exactly when it applies and out the rest of the time. That's the difference between a memory of what to do and a memory of when to use it. The trigger is the part to get right.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Layer memory by half-life. Instant preferences live in a compact always-on store; entity facts in a structured store you can query; long-lived conventions in a vault you load on purpose. Each level has its own audit cadence, so the always-on layer never bloats and the deep layer never goes stale silently. Cap and prune the always-on store on a schedule: an always-on memory that grows forever is context rot in slow motion.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Patterns to rot-proof an agent
&lt;/h2&gt;

&lt;p&gt;Whatever you pick, these hold.&lt;/p&gt;

&lt;p&gt;Make forgetting visible. If a session drops something, drop it loudly: what entered context, what left it, why. You can't manage decay you can't see.&lt;/p&gt;

&lt;p&gt;Bind decisions harder than facts. A decision needs a source, a status, an owner. A degraded decision is an incident, not a nuisance.&lt;/p&gt;

&lt;p&gt;Retrieve, don't carry. Pull the relevant slice on demand instead of carrying a fat context everywhere. The leaner the working set, the slower the rot.&lt;/p&gt;

&lt;p&gt;One source of truth for shared rules. When agents cooperate, conventions belong in one place anyone can read, not duplicated into private memory where they rot out of sync and you end up fighting three versions of the same rule.&lt;/p&gt;

&lt;p&gt;Treat memory as code. Schema changes, evals, tests. The vault is source, the index is a produced artifact: rebuild it on a cron and never hand-edit it. A knowledge graph you can't rebuild from source is a liability.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable take
&lt;/h2&gt;

&lt;p&gt;You never beat context rot. You manage it. Design so the cost of rotting is contained and observable, and so the source of truth survives the session because it lives outside it.&lt;/p&gt;

&lt;p&gt;Our agents now run week-long cycles with far less decay than the first version at hour three. The difference was never a better model. It was admitting that an agent's context is a perishable working surface, not a mind, and engineering it that way.&lt;/p&gt;

&lt;p&gt;I keep getting asked which framework solves this. The answer stays boring: write good compression, index for meaning, encode lessons as triggered skills, cap the always-on memory, and keep long-lived truth somewhere the agent can query and rebuild instead of carry. The agent will look cooler carrying everything. It will also rot faster.&lt;/p&gt;

&lt;p&gt;So: how do you decide what an agent is allowed to forget, and how do you make sure it tells you when it does?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>agents</category>
    </item>
    <item>
      <title>When Should You NOT Use an Agent?</title>
      <dc:creator>Danil Galeev</dc:creator>
      <pubDate>Thu, 27 Aug 2026 19:21:36 +0000</pubDate>
      <link>https://dev.to/madebyexpert/when-should-you-not-use-an-agent-4bk2</link>
      <guid>https://dev.to/madebyexpert/when-should-you-not-use-an-agent-4bk2</guid>
      <description>&lt;p&gt;Everyone is asking "should we use agents?" The real question is "when should we NOT?"&lt;/p&gt;

&lt;p&gt;I keep seeing teams bolt an agent on because it's the hot thing — then discover they reinvented a state machine with worse debugging. Agents don't solve a problem by existing. They are a mechanism for &lt;em&gt;deferring decisions to a runtime&lt;/em&gt;. When your inputs, tools, and failure modes are well-understood, that deferral buys you nothing but nondeterminism.&lt;/p&gt;

&lt;p&gt;The architecture question is not "LLM or not." It is: &lt;strong&gt;where does the judgment boundary sit?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Three places the boundary lands
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A single agent is a program.&lt;/strong&gt; For a well-scoped task with a known toolset, you don't need a loop at all. You need a deterministic pipeline — with the LLM as one component, not the orchestrator.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The moment you need a loop, you're building a runtime.&lt;/strong&gt; A runtime is a different beast. It has observability, tool permissions, credit and rate limits, and a way to explain what it did after the fact. That's not "more AI." That's distributed systems with a language model as the cognitive layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The expensive failure is capability you never signed up for.&lt;/strong&gt; Agents surface things you didn't design for: open-ended tool calls, emergent side-effects, scale that hits budgets or audit. This is where maturity shows — not in the cleverness of the model, but in the &lt;em&gt;constraints&lt;/em&gt; around it: entitlements, approval, observability, evals that cover the failure path, not just the happy path.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tell
&lt;/h2&gt;

&lt;p&gt;Most of what teams call "agent architecture" is a decision-making boundary placed in the wrong spot, papered over with more layers. If you reach for a framework, an orchestrator, a runtime — stop and ask what you're actually deferring, and whether that deferral is working or just making the system harder to debug.&lt;/p&gt;

&lt;p&gt;A few heuristics I use before building:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Can the task be expressed as steps with a known order?&lt;/strong&gt; Pipeline with the LLM as a step — not an agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does the agent invoke known tools with expected outputs?&lt;/strong&gt; A single governed agent, thin on top of a deterministic core.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do you need to route mid-task to unexpected states?&lt;/strong&gt; Only then a real runtime, and only if you accept owning its observability, limits, and audit trail as first-class work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Is the failure non-recoverable?&lt;/strong&gt; Then don't put an agent in the loop at all. A wrong tool call inside a runtime can hurt you faster than a slower deterministic path ever will.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Constrain before you automate. The agent looks cooler; it's also much harder to explain three months from now.&lt;/p&gt;

&lt;p&gt;So: when you're scoping a new system, what makes you reach for a deterministic pipeline instead of an agent — or the other way? I'd like the rules teams actually run with, not the ones they present in talks.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>agents</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
