<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: T. Alam</title>
    <description>The latest articles on DEV Community by T. Alam (@timalam01).</description>
    <link>https://dev.to/timalam01</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4010534%2F995f9bf0-d513-4e6b-8bb6-6ce69de16b13.jpeg</url>
      <title>DEV Community: T. Alam</title>
      <link>https://dev.to/timalam01</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/timalam01"/>
    <language>en</language>
    <item>
      <title>The Azure Bug That Cost Us an Afternoon Was One String Not Matching Another String</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Fri, 25 Sep 2026 19:42:26 +0000</pubDate>
      <link>https://dev.to/timalam01/the-azure-bug-that-cost-us-an-afternoon-was-one-string-not-matching-another-string-578m</link>
      <guid>https://dev.to/timalam01/the-azure-bug-that-cost-us-an-afternoon-was-one-string-not-matching-another-string-578m</guid>
      <description>&lt;p&gt;We migrated a feature from calling OpenAI's API directly to going through Azure AI instead — same underlying model family, just routed through our Azure subscription for reasons that had nothing to do with the model itself. Should've been a config change. Instead it turned into an afternoon of checking API keys, checking network rules, checking whether our Azure subscription had actually provisioned correctly, before someone finally found the actual problem: a string that didn't match another string, in a place none of us thought to look first.&lt;/p&gt;

&lt;p&gt;Writing this up because we're fairly sure we won't be the last team to lose an afternoon to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually different about Azure
&lt;/h2&gt;

&lt;p&gt;Every other provider we'd connected before this — OpenAI, Anthropic, Gemini, Hugging Face — you get an API key and you call model names directly. gpt-4o means gpt-4o. Azure AI doesn't work that way. You have to have already deployed a model inside your own Azure resource before anything can call it, and Azure doesn't let you address that model by its public name. You address it by whatever name you (or whoever set up the resource) gave that deployment when it was created in the Azure portal.&lt;/p&gt;

&lt;p&gt;That one fact is responsible for basically all of the confusion in this story.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Two things have to exist before this works at all: an Azure resource, and a model actually deployed inside it.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup, which looked completely normal
&lt;/h2&gt;

&lt;p&gt;In the DNotifier portal: Projects → AI Studio → Providers → Azure AI. We plugged in the resource endpoint (something like &lt;code&gt;https://your-resource.openai.azure.com&lt;/code&gt;), an API key, and the deployment list. Installed the SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; @dnotifier-realtime/dnotifier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Took our existing OpenAI-calling code, swapped the provider field, and shipped it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_API_KEY&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;USER_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;azure_ai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-4o&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// &amp;lt;-- this is the line that broke everything&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Summarize this quarter's key metrics.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;saveHistory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looks completely reasonable. It is completely reasonable, if your Azure deployment happens to be named &lt;code&gt;gpt-4o&lt;/code&gt;. Ours wasn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually failed, and why nothing pointed at the real cause
&lt;/h2&gt;

&lt;p&gt;Every call came back with a 404-ish deployment error. Not "invalid API key," not "network unreachable" — a deployment error, which in hindsight is exactly the right error and in the moment told us almost nothing, because none of us had internalized yet that &lt;code&gt;model&lt;/code&gt; here doesn't mean what it means for every other provider.&lt;/p&gt;

&lt;p&gt;We checked the API key. Fine. We checked the resource endpoint URL for typos. Fine. We checked whether the Azure resource had finished provisioning. Fine. Somewhere around the twenty-minute mark, someone actually opened the Azure portal and looked at the resource's deployment list instead of staring at our own code one more time.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Azure Portal → your-resource → Deployments

Deployment name          Model                Status
chatapp-gpt4o-prod        gpt-4o               Succeeded
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;The error looks like an integration problem. It's almost always a naming mismatch, checked against the Azure portal's deployment list.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There it was. Whoever set up this Azure resource — months earlier, a different engineer, long since moved to a different team — had named the deployment &lt;code&gt;chatapp-gpt4o-prod&lt;/code&gt;, not &lt;code&gt;gpt-4o&lt;/code&gt;. Reasonable choice on their part, especially in an org with more than one deployment floating around. Just not the string our code was sending.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix, once we knew what to look for
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;USER_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;azure_ai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chatapp-gpt4o-prod&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// the actual deployment name, not the model's public name&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Summarize this quarter's key metrics.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;saveHistory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One line. Worked immediately, no other changes needed anywhere. The fix took about two minutes once someone knew to look at deployment names specifically. Finding that out took the other twenty-plus.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule we now follow, no exceptions
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;model&lt;/code&gt; field in a &lt;code&gt;sendAI()&lt;/code&gt; call against Azure AI is whatever your deployment is named in the Azure (or Microsoft Foundry) portal, full stop. It can match the underlying model's public name — Azure's default deployment name sometimes does — but "sometimes" is exactly the trap. We stopped assuming and started checking the actual deployment list every single time we point new code at an Azure resource, especially one we didn't personally set up.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Azure Portal → [your resource] → Model deployments

Whatever appears here, verbatim, is what belongs in the `model` field.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Not the model's public name. Not an abbreviation. Not what you assume it "should" be called based on every other provider's convention.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One resource, several model families — the decision is which lane a task actually belongs in, not which model sounds most impressive.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  It's not just the name — deployment type is its own decision
&lt;/h2&gt;

&lt;p&gt;Once we'd fixed the immediate bug, we ran into the second layer of this: Azure's model catalog spans OpenAI's GPT family, Claude, Llama, Mistral, DeepSeek, and others, all inside the same Foundry resource, and picking a model family is only half the decision. Deployment type — Standard, Global Standard, Data Zone Standard, Provisioned Throughput — behaves differently on pricing, data residency, and latency consistency under load. A task with hard data-residency requirements might need Data Zone Standard regardless of which model family you picked. A high-volume, latency-sensitive feature might justify Provisioned Throughput even for a mid-tier model.&lt;/p&gt;

&lt;p&gt;We'd initially deployed our most capable available model for a document-summarization tool, assuming quality would scale with size. Running the same representative prompts through DNotifier's Prompt Testing Studio against a smaller, cheaper deployment from the same model family showed statistically indistinguishable output quality for our specific documents — the task just didn't need the bigger model. Switching cut per-request cost substantially with zero measurable drop in quality, and we wouldn't have made that switch with any real confidence without the side-by-side comparison sitting in front of us.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// same task, two deployments, tested against the same real prompts&lt;/span&gt;
&lt;span class="c1"&gt;// before committing production traffic to either one&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;resultA&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;azure_ai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chatapp-gpt4o-prod&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;testPrompt&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;resultB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;azure_ai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;chatapp-gpt4o-mini-prod&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;testPrompt&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="c1"&gt;// compare quality, latency, and cost before picking one for real traffic&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What we'd tell anyone setting this up fresh
&lt;/h2&gt;

&lt;p&gt;Before you write a single line pointing at an Azure deployment: open the portal, find the actual deployment name, and use that string exactly. Don't assume it matches the model's public name just because every other provider you've touched works that way — Azure is genuinely different here, not being difficult for no reason, just structured around your own resource rather than a shared public catalog. And once it's working, don't stop at "which model family" — test the actual deployment type and tier against real prompts before you commit production traffic to it. Twenty minutes of confusion over a naming mismatch is annoying. Overpaying for capability a task never needed is the quieter, more expensive version of the same mistake.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>We Built a Multi-Agent Support System That Refuses to Auto-Send Refunds. Here's the Whole Thing.</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Thu, 24 Sep 2026 17:21:58 +0000</pubDate>
      <link>https://dev.to/timalam01/we-built-a-multi-agent-support-system-that-refuses-to-auto-send-refunds-heres-the-whole-thing-2oa</link>
      <guid>https://dev.to/timalam01/we-built-a-multi-agent-support-system-that-refuses-to-auto-send-refunds-heres-the-whole-thing-2oa</guid>
      <description>&lt;p&gt;One AI agent answering every support ticket is a good demo. We know because we built that first, showed it around internally, and got exactly the reaction you'd expect — "cool, when's it live." Then we actually thought about what "live" meant. Billing questions need real invoice data. Technical questions need product docs and sometimes an actual diagnostic call. And anything involving a refund above a certain amount needs a human to look at it before it goes out, because we were not comfortable with a language model unilaterally deciding to give someone their money back.&lt;/p&gt;

&lt;p&gt;Cramming all of that into one system prompt and one model call gets you something mediocre at everything instead of good at anything. So we didn't. Here's the actual build — three specialized agents, a shared session, and a human approval gate that we do not let anything skip.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we're building
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customer message
       │
       ▼
 [router agent] ──reads session memory, classifies──┐
       │                                              │
   ┌───┴────┬─────────────┐                           │
   ▼         ▼             ▼                          │
[billing] [technical] [escalation] ──always flags for human review
   │         │             │
   └─────────┴─────────────┘
              │
              ▼
     draft response + needsHumanApproval flag
              │
     ┌────────┴────────┐
     ▼                  ▼
 auto-send          held for review
 (routine)          (sensitive)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One shared session, three specialists, one gate that anything sensitive has to pass through before a customer ever sees it.&lt;/p&gt;

&lt;p&gt;One ticket, one shared session, three specialized agents, and a human who signs off before anything sensitive ships.&lt;/p&gt;

&lt;h2&gt;
  
  
  Defining the agents
&lt;/h2&gt;

&lt;p&gt;DNotifier's agent primitive is DNotifier.defineAgent({ name, model, run(ctx) }). The ctx your run function gets handed gives you the shared workflow state, this step's input, and a sendAI() scoped to whatever model you put on that specific agent.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@dnotifier-realtime/dnotifier&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;routerAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;router-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-4o-mini&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// cheap and fast — this runs on every single message&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;classification&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Classify this support message into one category:
        billing, technical, or escalation. Message: "&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"
        Respond with just the category word.`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;classification&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;billingAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;billing-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-4o&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;draftResponse&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;technicalAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;technical-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-4o&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;draftResponse&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;escalationAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;escalation-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-4o&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;draft&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Draft a careful, empathetic response to this escalated
        issue, and flag it for human review before sending:
        "&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;draftResponse&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;draft&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;needsHumanApproval&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;draft&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model choice per agent isn't an afterthought — it's the whole point. The router runs on gpt-4o-mini because deciding "is this a billing question or a technical one" doesn't need a flagship model, and it's going to run on literally every incoming message, so it had better be cheap. The specialists run on gpt-4o because they're doing the reasoning a customer will actually judge. If we ran everything through the expensive model just because the escalation path needed it, we'd be paying flagship prices for "what are your support hours" questions all day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tying it into a workflow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;supportWorkflow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Workflow&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;customer-support-pipeline&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Routes and resolves inbound support tickets across three specialist agents&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;observability&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;routerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;billing&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;billingAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;category&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;technical&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;technicalAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;escalationAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;category&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;needsHumanApproval&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;needsHumanApproval&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;supportWorkflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;registerAgents&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nx"&gt;routerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;billingAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;technicalAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;escalationAgent&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;observability: true is one line and it's saved us more debugging time than almost anything else in this system. The first time someone from support Slacked us "the bot gave a weird answer on ticket #4471," we didn't reconstruct anything from a transcript. We opened the run, saw exactly which agent handled it, on which model, with what input, and had an answer back in under a minute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it against a real ticket
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;appId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_APP_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_SECRET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;support-system&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;transport&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ws&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;WebSocketImpl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;runWorkflow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;supportWorkflow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;I was charged twice for my subscription this month.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ticket-88213&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;// { category: 'billing', response: '...', needsHumanApproval: false }&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Routine billing question, routes to the billing agent, resolves without anyone in the loop. Now the same customer, same ticket, a follow-up:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;followUp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;runWorkflow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;supportWorkflow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;This is the third time this has happened and I want to cancel and get a full refund for the year.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ticket-88213&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// { category: 'escalation', response: '...', needsHumanApproval: true }&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Because it's the same sessionId, the escalation agent isn't starting from zero — it already has the prior turn's context about the duplicate charge, without our application code having to reassemble and re-pass that history manually. That part genuinely surprised us the first time we tested it; we expected to have to wire up our own history-stitching logic and didn't need to.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part we actually care about most: the approval gate
&lt;/h2&gt;

&lt;p&gt;needsHumanApproval coming back true is a hard stop, not a suggestion. This is where the drafted response gets pushed into whatever review surface support already uses — for us that's a Slack channel with an approve/reject button — and the "send to customer" function only fires once a human actually clicks approve.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;needsHumanApproval&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;notifyReviewQueue&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;ticketId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ticket-88213&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;draft&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="c1"&gt;// held here. nothing goes out until a human approves it.&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendToCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We'll be blunt about why this matters more than it might look on the page: a system that can draft a refund confirmation and a system that can send one without anyone checking it first are very different systems from a risk standpoint. "Which of our agents can take an irreversible action without a human in the loop" is a question we wanted a deliberate answer to, not one we backed into by accident because the demo worked and nobody circled back to add the guardrail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why three agents instead of one really good prompt
&lt;/h2&gt;

&lt;p&gt;We got this pushback internally more than once, and it's a fair question — doesn't one sufficiently detailed prompt handle all of this? For a while it kind of does. It falls apart for a few concrete reasons once you're past the prototype stage.&lt;/p&gt;

&lt;p&gt;The knowledge base gets muddier fast. A billing agent grounded only in invoice and pricing docs gives sharper answers than one generalist searching across billing, technical, and policy docs at once — less noise for the retrieval step to sort through.&lt;/p&gt;

&lt;p&gt;Cost stops making sense. Routing "what's your refund policy" through the same expensive model handling genuine escalations is money spent for zero quality gain on the easy majority of traffic.&lt;/p&gt;

&lt;p&gt;Debugging gets genuinely harder. One sprawling prompt that goes wrong means untangling a single enormous set of instructions. Three focused agents means the observability dashboard tells you exactly which one misfired, every time.&lt;/p&gt;

&lt;p&gt;And the approval logic gets fuzzy. "Escalate to a human if the refund's over $500" is a clean, auditable rule sitting in a dedicated escalation agent. It's a rule that's a lot easier to lose track of buried inside one long prompt that's also handling billing lookups and troubleshooting steps.&lt;/p&gt;

&lt;p&gt;None of this is specific to support tickets, either. The same router-then-specialists-then-shared-state shape, with an optional human gate bolted on, shows up in content review pipelines and internal ops tooling we've built since. Support was just the clearest place to build it first and see if the pattern actually held up under real traffic. It did.&lt;/p&gt;

&lt;p&gt;If you're staring down a single monolithic support prompt that's slowly turning into a maintenance headache, this is the refactor. It took us about a day to build the version above, and the approval gate alone was worth it the first week — it caught a draft response none of us would have wanted to send automatically.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Everyone's Talking About MCP Servers. Here's What Actually Happens When You Wire One Into an Agent</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Wed, 23 Sep 2026 17:33:25 +0000</pubDate>
      <link>https://dev.to/timalam01/everyones-talking-about-mcp-servers-heres-what-actually-happens-when-you-wire-one-into-an-agent-14gb</link>
      <guid>https://dev.to/timalam01/everyones-talking-about-mcp-servers-heres-what-actually-happens-when-you-wire-one-into-an-agent-14gb</guid>
      <description>&lt;p&gt;If you've been anywhere near agent frameworks in the last year, you've seen "MCP" mentioned enough times that it's started to feel like one of those terms everyone nods along to without necessarily having built anything with it. We were in that camp for a while too. Then we actually needed an agent that could check a customer's real account data instead of answering from general knowledge, and MCP stopped being a buzzword and started being the thing that saved us from writing the same integration glue code for the fourth time.&lt;/p&gt;

&lt;p&gt;This post is what MCP actually is, why it's grown as fast as it has, and the actual code for wiring an MCP server into a DNotifier agent — not the conceptual version, the one we run.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-line version
&lt;/h2&gt;

&lt;p&gt;The Model Context Protocol is an open standard, built by Anthropic, that gives a model a consistent way to talk to external tools and data sources — a filesystem, a database, an internal API — without every single application inventing its own bespoke integration for each one. Think of it as a shared plug shape: build a server once, and any MCP-aware model or platform can use it, instead of writing custom glue per tool per model.&lt;/p&gt;

&lt;p&gt;That sounds like a small thing until you've been the person writing the fourth version of "describe this tool to the model in the system prompt, parse whatever comes back, hope the format doesn't drift."&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's not just Anthropic's pet project anymore
&lt;/h2&gt;

&lt;p&gt;We didn't take this seriously until we looked at the actual numbers. As of a recent count, there are roughly 15,930 public MCP servers spread across the major registries, with close to 9,700 of those in the official one alone. Monthly SDK downloads are past 97 million. And — this is the part that actually got our attention — Microsoft, Google, and AWS have all built native MCP support into their own platforms (Copilot Studio, Gemini, Bedrock AgentCore). That's an unusual amount of cross-vendor agreement for anything in this space. Companies that compete on basically everything else agreed on this one plug shape.&lt;/p&gt;

&lt;p&gt;The protocol also shipped what its maintainers called the biggest revision in its history: the transport layer moved to stateless-by-default, and three older features got put on a 12-month deprecation clock. The major SDKs had compatible releases out within about three weeks of that announcement. That's a protocol being taken seriously enough to move fast on breaking changes, not something coasting on hype.&lt;/p&gt;

&lt;p&gt;The honest caveat, because we'd be doing you a disservice not to mention it: independent security scans have found real, exploitable flaws in a meaningful chunk of public MCP servers — enough that the NSA and CISA jointly put out security guidance on it. Fast growth plus low friction to publish a server means you get both real adoption and real sloppiness in the same ecosystem. Whatever server you're about to connect to a production agent deserves the same scrutiny you'd give any other third-party dependency touching your data. "It implements a protocol Anthropic built" is not the same thing as "it's trustworthy."&lt;/p&gt;

&lt;h2&gt;
  
  
  How this actually fits into DNotifier
&lt;/h2&gt;

&lt;p&gt;DNotifier's AI Foundation layer treats MCP servers as a first-class connection type, right alongside your models and your enterprise data sources. You connect a server once, in the dashboard, and it becomes available to whatever agents need it — the same mental model as connecting an API key.&lt;/p&gt;

&lt;p&gt;One agent, two channels. The model side and the tool side are both configured once, in the dashboard — not re-wired into every prompt.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// the agent's own call doesn't change based on which MCP tools are attached —&lt;/span&gt;
&lt;span class="c1"&gt;// tool access is a dashboard-level connection, not something you re-describe&lt;/span&gt;
&lt;span class="c1"&gt;// on every request&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Check the latest invoice for this account and summarize it.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="c1"&gt;// behind the scenes: the model reasons about the request, decides it needs&lt;/span&gt;
&lt;span class="c1"&gt;// the invoice data, and DNotifier routes that lookup through the connected&lt;/span&gt;
&lt;span class="c1"&gt;// MCP server — none of that plumbing lives in this call&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's genuinely the whole client-side shape of it. The heavy lifting — deciding a tool call is warranted, structuring the call, interpreting what comes back — is the model's job and the protocol's job, not something we're writing per feature anymore.&lt;/p&gt;

&lt;h2&gt;
  
  
  A real example: grounding a support agent in a live account database
&lt;/h2&gt;

&lt;p&gt;Here's roughly what we had before MCP, for an agent that needed to check account state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// the old way — a hand-rolled function, described in the prompt,&lt;/span&gt;
&lt;span class="c1"&gt;// parsed manually, and repeated for every new data source&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;checkAccountStatus&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;accountId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SELECT * FROM accounts WHERE id = $1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;accountId&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;row&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;systemPrompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`
You can call checkAccountStatus(accountId) to look up account data.
If the user asks about their account, call this function and use the result.
Respond in this format: ...
`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="c1"&gt;// then hand-parse whatever the model decides to output, hope it matches&lt;/span&gt;
&lt;span class="c1"&gt;// the format you asked for, repeat this whole dance for the next tool&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That works, technically, for one tool. It gets worse fast once you have four or five data sources, each with its own quirks, each needing its own parsing logic, each one more surface area for the model to accidentally hallucinate a call it never actually makes.&lt;/p&gt;

&lt;p&gt;With a database MCP server connected once through DNotifier, the same capability looks like this from the application's side:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;customerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;sessionId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Why was I charged twice this month?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="c1"&gt;// the model recognizes it needs account/invoice data, calls the MCP&lt;/span&gt;
&lt;span class="c1"&gt;// server's exposed tool for it, gets back structured data, and reasons&lt;/span&gt;
&lt;span class="c1"&gt;// over the actual result — no bespoke parsing code on our end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The integration code didn't grow. When we added a second MCP server a few weeks later — this one for an internal ticketing system — the application-side code for the agent's sendAI() call didn't change at all. That's the part that made this worth adopting for us: the cost of adding tool number five is the same as tool number one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does this only matter if you're using Claude?
&lt;/h2&gt;

&lt;p&gt;No, but it's worth being straight about why Claude specifically tends to be smoother here. Anthropic didn't just publish the spec and walk away — Claude's own training and tool-use behavior were shaped by being the model built alongside this exact protocol. In practice that shows up as Claude being noticeably more fluent about deciding when a tool call is warranted and structuring it correctly, with less prompt-engineering scaffolding required to hold its hand through the process. We've run MCP-connected tools against other providers through DNotifier too, and it works — it's just that Claude needed less coaxing to use them well, which tracks given who built the protocol.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we'd actually check before connecting a new server
&lt;/h2&gt;

&lt;p&gt;Not a formal checklist, just what we've started doing by habit: who maintains it, and how actively. Is it in the official registry or some unlisted repo somebody linked in a Discord. What permissions is it actually requesting versus what the task genuinely needs. If any of those answers feel shaky, we don't connect it to anything touching real customer data, full stop — there are enough servers out there now that "just use a different one" is usually a real option, not a hypothetical.&lt;/p&gt;

&lt;p&gt;If you're still hand-rolling tool descriptions into system prompts for every integration you add, this is worth the hour it takes to try. The protocol did the annoying part already. What's left is mostly picking servers you trust and wiring them in once.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>We Ran a Model Locally with Ollama and DNotifier. It Failed Silently, and Then We Accidentally Exposed It to the Internet.</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Tue, 22 Sep 2026 17:35:19 +0000</pubDate>
      <link>https://dev.to/timalam01/we-ran-a-model-locally-with-ollama-and-dnotifier-it-failed-silently-and-then-we-accidentally-2h8e</link>
      <guid>https://dev.to/timalam01/we-ran-a-model-locally-with-ollama-and-dnotifier-it-failed-silently-and-then-we-accidentally-2h8e</guid>
      <description>&lt;p&gt;Here's a fun way to spend an afternoon: get an integration working perfectly on your laptop, ship the exact same config to your actual deployment, and watch every single call fail with an error message that tells you absolutely nothing useful. That's what happened the first time we hooked Ollama up to DNotifier, and the fix — once we understood it — took about four minutes. Understanding it took considerably longer.&lt;/p&gt;

&lt;p&gt;This is the writeup of what actually went wrong, why, and then the second mistake we made fixing it, which was arguably worse than the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why bother with local models at all
&lt;/h2&gt;

&lt;p&gt;Every other provider we'd connected to DNotifier up to this point — OpenAI, Claude, Gemini, Bedrock — is the same basic shape: you get an API key, you make a network call to somebody else's cloud, you get a bill. Ollama isn't that. It's an open-source runtime, MIT-licensed, built on llama.cpp, that runs models on hardware you control. Your laptop, a workstation, a box in your own data center, whatever. No account to create on DNotifier's side for it, no key to paste, no usage-based bill from a model vendor.&lt;/p&gt;

&lt;p&gt;We didn't reach for it because it's cooler (though it kind of is). We reached for it because we had one specific workflow step where the data genuinely could not leave our own infrastructure — some internal document analysis where "send this to a third-party API" wasn't a decision we got to make casually. Ollama was the obvious answer for that one step. Everything else stayed on cloud providers.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Every cloud provider routes to someone else's infrastructure. Ollama routes to yours.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting it up looks almost insultingly simple
&lt;/h2&gt;

&lt;p&gt;On the machine running Ollama:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama pull llama3.2
ollama list
&lt;span class="c"&gt;# NAME               ID              SIZE      MODIFIED&lt;/span&gt;
&lt;span class="c"&gt;# llama3.2:latest     a80c4f17acd5    2.0 GB    3 minutes ago&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then in the DNotifier portal, under Projects → AI Studio → Providers → Ollama, you configure exactly one thing: the base URL DNotifier should call. That's it. No auth field, because Ollama doesn't require one by default. Provider string in code is just "ollama":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_API_KEY&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;senderId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ollama&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;llama3.2&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// has to match what `ollama list` actually shows&lt;/span&gt;
  &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Summarize this doc.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;saveHistory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// more on this below&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We copy-pasted &lt;a href="http://127.0.0.1:11434" rel="noopener noreferrer"&gt;http://127.0.0.1:11434&lt;/a&gt; into the base URL field, because that's what curl had been happily talking to on the dev machine for the last hour. Saved it. Ran the exact same code against the deployed environment.&lt;/p&gt;

&lt;p&gt;Every single call failed. Not a helpful failure either — nothing that said "hey, this address is unreachable from here." Just generic connection errors that looked, at a glance, like they could be an auth problem, a firewall thing, a DNS thing, basically anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mistake, once we actually understood it
&lt;/h2&gt;

&lt;p&gt;127.0.0.1 is a loopback address. By definition, it means "this exact machine, and nothing else." When we tested with curl on our own laptop, 127.0.0.1:11434 correctly meant our laptop. When DNotifier's cloud AI runtime tried to call that same address, 127.0.0.1 from its perspective means itself — some server sitting in DNotifier's infrastructure, definitely not our laptop, and definitely not running Ollama.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The single most common Ollama setup failure, illustrated: pasting a loopback URL into a portal that runs somewhere else entirely.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;It's such an obvious mistake in hindsight that it's almost embarrassing to write up. But it's also, apparently, the single most common way this integration goes wrong for basically everyone, based on how directly DNotifier's own docs call it out: "That address only works if the DNotifier worker that runs AI can reach that address. Cloud workers cannot see your laptop loopback."&lt;/p&gt;

&lt;p&gt;There are three real fixes, and we want to be honest that we tried them roughly in order of "least amount of work" before landing on the right one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run Ollama on a host that's actually network-reachable from wherever DNotifier's AI runtime executes.&lt;/li&gt;
&lt;li&gt;Use a tunnel or VPN.&lt;/li&gt;
&lt;li&gt;Use a self-hosted DNotifier deployment, if that's available to you.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We ended up standing up a small VM inside our own cloud VPC, running Ollama there instead of on anyone's laptop, and pointing the portal at that address. Same three lines of code as above, just a different base URL. The whole "fix" was maybe four minutes once we knew what we were actually looking at — the twenty-plus minutes before that were spent checking API keys and network settings that had nothing to do with the actual problem.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// what actually worked, pointed at the internal VM instead of a loopback address&lt;/span&gt;
&lt;span class="c1"&gt;// base URL configured in the portal: http://10.0.4.17:11434&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Then we made the second mistake
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Solving "unreachable" by making an endpoint public solves reachability and opens a new problem in the same step, unless you also solve for who else can reach it.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Getting Ollama reachable from DNotifier's side solves exactly one problem: reachability. It does not solve the problem of who else can reach it. We learned this the annoying way.&lt;/p&gt;

&lt;p&gt;To get past the loopback issue quickly on a proof-of-concept, someone on the team opened the VM's inference port directly to the public internet, no auth in front of it, with the reasoning "it's just a demo, we'll tear it down soon." The demo ran longer than "soon" implies demos ever do. A routine security review a few weeks later flagged it as an active finding: an unauthenticated model-inference endpoint, reachable by anyone who found the IP, with no way to tell legitimate traffic from anything else. Somebody could, in theory, have been running their own workloads on our compute and we'd have had no way to know.&lt;/p&gt;

&lt;p&gt;The fix for that was also small — a reverse proxy in front of the Ollama endpoint that checks a credential before forwarding anything through — but it should never have been necessary in the first place if we'd thought about reachability and security as two separate questions from the start, instead of assuming solving one solved the other.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="c1"&gt;# rough shape of what we put in front of it — nginx checking a&lt;/span&gt;
&lt;span class="c1"&gt;# shared secret header before forwarding to Ollama at all&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/ollama/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;if&lt;/span&gt; &lt;span class="s"&gt;(&lt;/span&gt;&lt;span class="nv"&gt;$http_x_internal_key&lt;/span&gt; &lt;span class="s"&gt;!=&lt;/span&gt; &lt;span class="s"&gt;"REDACTED")&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kn"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://127.0.0.1:11434/&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(That's illustrative, not a copy-paste production config — the actual setup uses a proper secrets manager for the key, not a string sitting in an nginx conf file. Don't do that part the way this snippet implies.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The three ways to close that gap, and which we'd actually recommend
&lt;/h2&gt;

&lt;p&gt;A tunnel or VPN, done right, keeps the endpoint reachable only to traffic that's already authenticated at the network layer — closest in spirit to the original "only reachable from trusted places" property a loopback address has, just extended specifically to include DNotifier's runtime. A public host needs its own explicit protection layered on top, since nothing about a public IP restricts who can hit it — that's the reverse-proxy-with-a-credential pattern above, or an IP allowlist if DNotifier's outbound traffic comes from a known, stable range. A self-hosted DNotifier deployment shifts the boundary again, potentially keeping the entire call path inside infrastructure you already control end to end.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;None of these are exotic — they're the same patterns any team uses to expose an internal service safely, applied here to a model-inference endpoint.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;If we were starting over, we'd go straight to the VPN/tunnel option instead of "public host plus proxy," mostly because it means the endpoint is never actually internet-facing at any point — one less thing to get wrong later.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we'd tell past-us
&lt;/h2&gt;

&lt;p&gt;Two things, really. First: if you're setting this up and it's failing silently, check the base URL before you check anything else — 127.0.0.1 anywhere in that config is almost always the actual bug, no matter what the error message seems to be pointing at. Second: reachability and security are not the same problem, and solving the first one doesn't get you the second one for free. Budget the extra hour for the proxy or the VPN from the start — it's a lot cheaper than fixing it after a security review flags it for you.&lt;/p&gt;

&lt;p&gt;We still run Ollama for that one document-analysis step where it genuinely earns its place, mixed into the same DNotifier Workflow as our cloud-provider steps. Same defineAgent shape, same ctx.state, no special-casing in the application code for the fact that one step never leaves our own hardware. That part, at least, worked exactly the way it was supposed to the first time.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>We Stopped Hardcoding a Model Provider Into Every Feature. Here's the Router That Fixed It</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Mon, 21 Sep 2026 17:23:33 +0000</pubDate>
      <link>https://dev.to/timalam01/we-stopped-hardcoding-a-model-provider-into-every-feature-heres-the-router-that-fixed-it-375n</link>
      <guid>https://dev.to/timalam01/we-stopped-hardcoding-a-model-provider-into-every-feature-heres-the-router-that-fixed-it-375n</guid>
      <description>&lt;p&gt;A couple of months ago we had four different features in production, each one calling a different model provider directly, each one with its own little pile of error handling, its own retry logic, its own idea of what a "timeout" should mean. OpenAI for the chat widget. A cheaper Hugging Face model for the classification step nobody thought about until the AWS bill showed up. Claude for anything that touched a document. It worked, technically. It also meant that every time we wanted to try a new model, or a cheaper one, or route around an outage, someone had to go change code in four places and redeploy four things.&lt;/p&gt;

&lt;p&gt;The fix wasn't "pick one provider and stick with it forever." It was building a router — one small agent whose entire job is deciding where a request should go, sitting in front of everything else. This post is that router, built with DNotifier's actual defineAgent and Workflow primitives, not a diagram with boxes and arrows that doesn't compile.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with "just pick the best model"
&lt;/h2&gt;

&lt;p&gt;Here's the thing nobody tells you until you've shipped a few of these systems: most of your traffic doesn't need your best model. We looked at three weeks of real requests hitting one of our internal tools and something like 80% of them were dead simple — short factual lookups, basic rewrites, one-line classifications. The other 20% actually needed real reasoning. Running everything through the expensive model because some requests need it is just burning money on the easy majority.&lt;/p&gt;

&lt;p&gt;So the router's job isn't "find the best model." It's "figure out what this specific request actually needs, then send it to the cheapest thing that can handle it."&lt;/p&gt;

&lt;h2&gt;
  
  
  What we're building
&lt;/h2&gt;

&lt;p&gt;One classifier agent that looks at an incoming request and decides how hard it is. Two (or more) answer agents behind it, each pointed at a different model. A workflow that ties them together and logs every routing decision so you can actually see what happened later, instead of guessing.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;request → [classifier agent] → "simple" or "complex"
                                      │
                    ┌─────────────────┴─────────────────┐
                    ▼                                     ▼
           [small/cheap model]                   [larger/capable model]
                    │                                     │
                    └─────────────────┬─────────────────┘
                                       ▼
                                   response
                                (logged + observable)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The routing decision is itself a small, cheap model call — not a guess encoded in application logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: the classifier
&lt;/h2&gt;

&lt;p&gt;This is the part people over-engineer. You don't need a big model to decide if a question is simple — you need a small, fast one that's good at exactly one narrow job.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@dnotifier-realtime/dnotifier&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;routerAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;complexity-router&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;huggingface/meta-llama/Llama-3.1-8B-Instruct:fastest&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;classification&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Classify this request as "simple" or "complex".
        Simple: factual lookups, short rewrites, basic classification.
        Complex: multi-step reasoning, long-document analysis, nuanced judgment calls.
        Respond with exactly one word.

        Request: "&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;classification&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;toLowerCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the &lt;code&gt;:fastest&lt;/code&gt; hint on the model string. That's us telling DNotifier we care more about latency than shaving fractions of a cent on this particular call — it's an 8B model running one classification, it should come back almost instantly. We'll flip that priority for the expensive path in a second.&lt;/p&gt;

&lt;p&gt;One thing worth being honest about: this is a soft classifier, not a guarantee. Occasionally it'll call something "simple" that really wasn't. That's fine — we'll get to why that's a recoverable problem later, not a fatal one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: the two answer paths
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;simpleAnswerAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;simple-answer-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;huggingface/meta-llama/Llama-3.1-8B-Instruct:fastest&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;complexAnswerAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;complex-answer-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;huggingface/openai/gpt-oss-120b:cheapest&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;See the difference in routing hints — &lt;code&gt;:fastest&lt;/code&gt; on the small model, &lt;code&gt;:cheapest&lt;/code&gt; on the big one. That's not a copy-paste mistake. On the small model, latency is what matters, and the cost difference between backing providers is basically noise. On the bigger, pricier call, the cost spread between backing providers is actually worth optimizing for, so we tell DNotifier to go find the cheapest route to that model instead of the fastest one. Small detail, but it adds up once you're running this at any real volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: wire it into a workflow
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;routingWorkflow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Workflow&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;complexity-based-model-router&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Classifies request complexity and routes to an appropriately sized model&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;observability&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;routerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;complex&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;complexAnswerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;simpleAnswerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;complexity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;routingWorkflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;registerAgents&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nx"&gt;routerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;simpleAnswerAgent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;complexAnswerAgent&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;observability: true&lt;/code&gt; is doing more work here than its one line suggests. Once this is live, every routing decision — what came in, what it got classified as, which model actually answered — shows up in the dashboard as it happens. The first time a teammate asked "why did the bot give a weak answer to this," we didn't have to reproduce anything. We opened the run and looked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: actually run it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;appId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_APP_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_SECRET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;routing-system&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;transport&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ws&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;WebSocketImpl&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;WebSocket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;runWorkflow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;routingWorkflow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;What year did the Berlin Wall fall?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;→&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;outcome&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;// simple → 1989&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same workflow, harder question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outcome2&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;notifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;runWorkflow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;routingWorkflow&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Given these three quarterly reports, what's driving the margin decline and is it structural or seasonal?&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outcome2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;complexity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;→&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;outcome2&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="c1"&gt;// complex → [a real, reasoned answer from the larger model]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nobody flipped a switch between those two calls. The router looked at the second question, decided it actually needed reasoning depth, and sent it somewhere else — automatically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this stops being a toy example
&lt;/h2&gt;

&lt;p&gt;The version above routes between two Hugging Face model sizes because that's a clean example, but nothing about &lt;code&gt;defineAgent&lt;/code&gt; ties you to one provider. This is the part that actually changed how we think about this stuff: because every agent is just a name, a model string, and a run function, the exact same pattern routes across any combination of providers you've got connected.&lt;/p&gt;

&lt;p&gt;We eventually extended this same shape to include a local model running through Ollama, for one specific category of request where the data genuinely couldn't leave our own infrastructure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;localOnlyAgent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;defineAgent&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;local-only-agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;provider&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ollama&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;llama3.2&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;text&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same &lt;code&gt;defineAgent&lt;/code&gt; shape. Same workflow wiring. The only thing that changed is which agent gets picked, and now the router isn't just choosing between "cheap" and "expensive" — it's choosing between "cloud" and "stays on our own hardware," based on whatever classification rule you want to write into the router's prompt.&lt;/p&gt;

&lt;p&gt;The same pattern, generalized past two model sizes into any mix of providers you've got connected — cloud or local.&lt;/p&gt;

&lt;p&gt;We'll get into the specifics of running Ollama in production (and the very dumb networking mistake we made the first time) in the next post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why we bothered with the extra agent instead of an if/else in application code
&lt;/h2&gt;

&lt;p&gt;We got asked this internally more than once — why not just check &lt;code&gt;input.length &amp;gt; 200&lt;/code&gt; or some heuristic in plain code and skip the model call entirely? Two honest reasons. First, a length check is a terrible proxy for complexity — "what's 2+2" and a three-sentence question requiring real synthesis can be the same length. Second, the classifier call is cheap enough (one small, fast model call) that the cost of running it on every single request, even the ones that end up routed to the expensive model anyway, is trivial next to what it saves on the 80% that don't need the expensive model at all.&lt;/p&gt;

&lt;p&gt;And when the router does misclassify something — it happens, rarely — it's a soft failure, not a crash. A complex question that gets routed to the simple model comes back with a weaker answer, not an error. We watch the classification and the answer quality together over time through the observability dashboard, and when the pattern shows the router being too aggressive about sending things downmarket, we adjust the prompt. It's tuning, not firefighting.&lt;/p&gt;

&lt;p&gt;If you're juggling more than one provider right now and every new one means another SDK, another auth flow, another set of error codes to learn — this pattern is worth the hour it takes to set up. It paid for itself for us within the first week.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Build Agent-to-Agent Communication With Events</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:22:10 +0000</pubDate>
      <link>https://dev.to/timalam01/build-agent-to-agent-communication-with-events-2fpc</link>
      <guid>https://dev.to/timalam01/build-agent-to-agent-communication-with-events-2fpc</guid>
      <description>&lt;p&gt;Your agents are stuck waiting on each other. One finishes a task, the next has no idea until you poll for it. That delay adds up fast, especially once you're running more than two or three agents at once. &lt;strong&gt;Agent-to-agent communication&lt;/strong&gt; fixes this. Events are the cleanest way to build it, and that's what this piece digs into: why direct calls fall apart, how event-driven messaging actually works, and what you need in place before things get messy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Agent-to-Agent Communication?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;Agent-to-agent communication&lt;/a&gt; is how one agent hands off information, requests, or results to another, no human required. Could be a research agent passing data to a summarizer. Could be a planner triggering three worker agents at once. Get it wrong and agents either sit around waiting or trip over each other's work. A clear agent-to-agent protocol keeps the handoff predictable, even when you can't predict much else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Direct Calls Don't Scale
&lt;/h2&gt;

&lt;p&gt;Most teams start with direct calls. Agent A calls Agent B, waits, moves on. Fine with two agents. Add a third, a fourth, a fifth, and the wiring turns into a mess fast. Every new agent means new connections you have to manage by hand.&lt;/p&gt;

&lt;p&gt;Distributed agent communication makes this worse. Agents often run on separate servers, with no shared memory to lean on when something goes wrong.&lt;/p&gt;

&lt;p&gt;Here's the comparison:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Direct calls&lt;/th&gt;
&lt;th&gt;Event-driven&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Adding a new agent&lt;/td&gt;
&lt;td&gt;Rewire existing agents&lt;/td&gt;
&lt;td&gt;Just subscribe to events&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure handling&lt;/td&gt;
&lt;td&gt;One slow agent blocks the rest&lt;/td&gt;
&lt;td&gt;Agents fail independently&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coordination&lt;/td&gt;
&lt;td&gt;Manual, agent by agent&lt;/td&gt;
&lt;td&gt;Centralized through events&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scaling&lt;/td&gt;
&lt;td&gt;Gets harder with each agent&lt;/td&gt;
&lt;td&gt;Stays flat as agents grow&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Direct calls create tight coupling. Every agent has to know who else exists and how to reach them. Brittle setup. Breaks the moment you add or remove an agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Events Change How Agents Talk
&lt;/h2&gt;

&lt;p&gt;An event is just a message saying something happened. "Task completed." "New data available." "Error detected." Agents publish events instead of calling each other directly, and other agents subscribe to whatever they care about, reacting when something fires.&lt;/p&gt;

&lt;p&gt;That's the core idea behind any solid agent communication protocol: nobody needs to know who's listening. The planner agent publishes a "plan ready" event. Three worker agents pick it up and start their tasks in parallel, no direct connection between any of them.&lt;/p&gt;

&lt;p&gt;It's what makes multi-agent coordination possible at scale. Add a tenth agent and it just subscribes to what it needs. No rewiring.&lt;/p&gt;

&lt;h2&gt;
  
  
  Core Pieces of an Event System
&lt;/h2&gt;

&lt;p&gt;A working AI agent communication protocol needs a few pieces in place, and skipping even one means agents either miss events entirely or drown in too many.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Event publisher&lt;/td&gt;
&lt;td&gt;Sends out events when something happens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Event subscriber&lt;/td&gt;
&lt;td&gt;Listens for specific event types&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Message broker&lt;/td&gt;
&lt;td&gt;Routes events to the right subscribers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Event schema&lt;/td&gt;
&lt;td&gt;Keeps event formats consistent across agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitoring layer&lt;/td&gt;
&lt;td&gt;Tracks what fired, when, and who responded&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Get the schema wrong and agents start misreading each other's messages. And without real monitoring, you won't know why an agent went quiet until something breaks in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where DNotifier Fits In
&lt;/h2&gt;

&lt;p&gt;Building this agent messaging architecture from scratch usually means stitching together a queue, a broker, and custom logging yourself. DNotifier skips that. Real-time pub/sub is built into one SDK, so agents publish and subscribe to events without extra infrastructure sitting on top.&lt;/p&gt;

&lt;p&gt;Every event gets traced automatically as it moves through the system. Say a worker agent never responds to a "plan ready" event. &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;DNotifier's monitoring and observability&lt;/a&gt; tools show you exactly where the message stopped, instead of leaving you to guess.&lt;/p&gt;

&lt;p&gt;For teams running multiple agents that need to coordinate in real time, this kind of agent collaboration infrastructure saves weeks of setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Mistakes to Avoid
&lt;/h2&gt;

&lt;p&gt;Publishing too many event types is a common trap. Fire an event for every tiny action and subscribers get flooded, then start tuning them out. Keep events meaningful: task done, error found, data ready.&lt;/p&gt;

&lt;p&gt;Skipping event versioning is another one. Change an event's structure and older subscribers break silently, no warning. Version from day one, even when it feels like overkill.&lt;/p&gt;

&lt;p&gt;And don't treat events as fire-and-forget. Track whether they were actually received and acted on. Skip that, and failures stay hidden until a customer finds them first.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What's the difference between an agent communication protocol and event-driven messaging?&lt;/strong&gt;&lt;br&gt;
A protocol defines the rules agents follow to exchange messages. Event-driven messaging is just one way to implement that protocol, using events instead of agents calling each other directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can agent-to-agent communication work without a message broker?&lt;/strong&gt;&lt;br&gt;
Technically, yes. But it gets messy fast. A broker routes events reliably and keeps agents decoupled, so skip one and you're back to managing direct connections by hand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How is multi-agent communication different from single-agent workflows?&lt;/strong&gt;&lt;br&gt;
Single-agent workflows run one process start to finish, nothing else involved. Multi-agent communication means several agents work in parallel and need a shared way to pass information back and forth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does event-driven architecture add latency?&lt;/strong&gt;&lt;br&gt;
A little, but not much. A well-built event system adds delay measured in milliseconds, not seconds. Worth the tradeoff once you're running more than a couple of agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;Agent-to-agent communication only gets harder to manage as you add more agents, unless it's built on events from day one. Start small. Keep your event types meaningful. Track what happens after each one fires.&lt;/p&gt;

&lt;p&gt;Want to see it in action? Explore the SDK at dnotifier.com.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Build a Production Agent Architecture With Node.js</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Thu, 10 Sep 2026 16:42:25 +0000</pubDate>
      <link>https://dev.to/timalam01/build-a-production-agent-architecture-with-nodejs-n8b</link>
      <guid>https://dev.to/timalam01/build-a-production-agent-architecture-with-nodejs-n8b</guid>
      <description>&lt;p&gt;Your AI agent works fine in the demo. Then real users show up, and it breaks in ways you never tested for. That gap, between a working prototype and a real production agent architecture, is where most Node.js teams get stuck.&lt;/p&gt;

&lt;p&gt;The problem usually isn't your prompt. It's your infrastructure. A single script calling an LLM API isn't an architecture, it's a proof of concept. That's the difference this guide covers: the production AI architecture patterns that separate a demo from something people can actually rely on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Production Agent Architecture Actually Is
&lt;/h2&gt;

&lt;p&gt;A &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;production agent architecture&lt;/a&gt; is the full system around an agent, not just the model call. It includes orchestration, memory, tool execution, monitoring, deployment. Everything the model call doesn't cover.&lt;/p&gt;

&lt;p&gt;Think of it like a car engine. The engine matters, but you still need wheels, brakes, and a fuel line to get anywhere. Skip those parts and it looks great in the demo, then stalls the second something unexpected happens.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Your Prototype Falls Apart in Production
&lt;/h2&gt;

&lt;p&gt;Most demos run one request, get one response, and stop there. Production traffic doesn't behave. Users send messages out of order. APIs time out. Models return broken JSON.&lt;/p&gt;

&lt;p&gt;A prototype has no memory between calls. No retry logic. No visibility into why something failed. These gaps stay hidden until real traffic hits them, and by then you're debugging live instead of building.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Core Layers Your Node.js Stack Needs
&lt;/h2&gt;

&lt;p&gt;A solid stack breaks into five layers, and each one has a different job. Input handles what comes in, whether that's a chat message or a webhook firing at 3am. Orchestration decides what happens next, which tool gets called, and in what order.&lt;/p&gt;

&lt;p&gt;Then there's execution, the layer that actually runs the functions, hits the APIs, queries the database. Memory keeps context around so the agent doesn't forget step two by the time it reaches step five. Semantic search helps here too, pulling in relevant history instead of dumping the whole conversation back into the prompt.&lt;/p&gt;

&lt;p&gt;Last is observability. Logs, traces, whatever tells you what the agent actually did instead of what you assumed it did.&lt;/p&gt;

&lt;p&gt;That's the idea behind decent AI agent system design: keep failures contained to one layer. If execution breaks, orchestration retries or falls back. The whole request doesn't have to die because one API call timed out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Single Agent or Multi-Agent? Pick Based on the Job
&lt;/h2&gt;

&lt;p&gt;Not every task needs multiple agents. A single agent handles narrow, linear work fine, like answering support tickets. Multi-agent setups earn their keep when the work splits into real roles: research, drafting, review.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;Single-Agent&lt;/th&gt;
&lt;th&gt;Multi-Agent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Best for&lt;/td&gt;
&lt;td&gt;Narrow, linear tasks&lt;/td&gt;
&lt;td&gt;Multi-step workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Complexity&lt;/td&gt;
&lt;td&gt;Low, easy to debug&lt;/td&gt;
&lt;td&gt;Higher, more moving parts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Faster, one call&lt;/td&gt;
&lt;td&gt;Slower, multiple handoffs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure points&lt;/td&gt;
&lt;td&gt;One&lt;/td&gt;
&lt;td&gt;Several&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coordination needed&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Orchestration layer required&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Multi-agent production architecture only pays off when a task actually needs separate roles. Bolt three agents onto a job one agent could handle, and you've just added three new ways for things to go wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Orchestration Is the Glue That Holds It Together
&lt;/h2&gt;

&lt;p&gt;Orchestration decides what happens next. It routes requests, manages handoffs between agents, and enforces the order tasks run in. Skip it and your agents just work in isolation, tripping over each other's outputs.&lt;/p&gt;

&lt;p&gt;In Node.js, this usually means an event-driven layer that tracks state and triggers the next action. That's what turns a loose set of scripts into a reliable agent architecture. DNotifier centralizes all of this in one SDK. You configure workflows and agent behavior from a single place instead of stitching five different services together.&lt;/p&gt;

&lt;h2&gt;
  
  
  You Can't Fix What You Can't See
&lt;/h2&gt;

&lt;p&gt;Monitoring and traceability aren't optional here. When an agent gives a wrong answer, you need to know which step caused it, not just that something went wrong.&lt;/p&gt;

&lt;p&gt;That's what enterprise AI agent architecture actually means in practice: traceability from day one, not something you bolt on after the incident already happened. Traceability logs every decision an agent makes, from the prompt it got to the tool it called. DNotifier's monitoring and traceability tools do this automatically, tracing a bad output back to its source in minutes instead of hours of log-diving.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployment Patterns That Actually Hold Up
&lt;/h2&gt;

&lt;p&gt;How you deploy an agent shapes how it fails. Pick based on your traffic pattern, not what's trendy this month.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;th&gt;Trade-off&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Serverless functions&lt;/td&gt;
&lt;td&gt;Spiky, unpredictable traffic&lt;/td&gt;
&lt;td&gt;Cold starts, time limits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Containers (Docker)&lt;/td&gt;
&lt;td&gt;Steady, predictable load&lt;/td&gt;
&lt;td&gt;More setup and upkeep&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-running process&lt;/td&gt;
&lt;td&gt;Real-time, stateful agents&lt;/td&gt;
&lt;td&gt;Needs manual scaling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Good AI agent infrastructure design rarely means picking just one option and calling it done. Plenty of teams run orchestration on a long-running process and push tool calls out to serverless functions instead. Whatever you pick, test prompt changes before they go live. DNotifier's prompt testing catches regressions before a user ever notices.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-Time Sync Keeps Agents From Going Stale
&lt;/h2&gt;

&lt;p&gt;Agents that talk to users, or to each other, need real-time updates. A pub/sub layer pushes events as they happen instead of forcing clients to poll for status.&lt;/p&gt;

&lt;p&gt;This matters most in chat-based agents, where users expect a response to stream in, not appear all at once. That's what AI application infrastructure actually looks like once you strip away the buzzwords. DNotifier's real-time pub/sub and chat system features cover this out of the box, so you're not building a WebSocket layer from scratch just to get agents talking.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What makes an architecture "production-ready" instead of just a working demo?&lt;/strong&gt;&lt;br&gt;
A production-ready setup handles failure, not just success. It includes retries, logging, monitoring, and recovery paths for when a model or tool call fails. A demo just has to work once. Production has to work every time, including the times it shouldn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need Kubernetes to run agents in production?&lt;/strong&gt;&lt;br&gt;
No, not really. Kubernetes helps once you're at scale, but plenty of teams run solid agents on a single container or one long-running Node process. Start simple. Add complexity only when traffic actually forces you to.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I stop one agent failure from breaking the whole system?&lt;/strong&gt;&lt;br&gt;
Isolate each layer so one failure doesn't cascade into five. If a tool call fails, orchestration should retry or fall back, not take the entire request down with it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can Node.js handle multi-agent systems at scale?&lt;/strong&gt;&lt;br&gt;
Yes, easily. Node's event loop is built for concurrent I/O, which is basically what multi-agent coordination needs. Most limits come from bad architecture, not the runtime itself.&lt;/p&gt;




&lt;p&gt;A production agent architecture was never about adding more agents. It's about building something that survives real traffic, bad inputs, and the failures you didn't see coming. Start with the layers. Worry about the feature list later.&lt;/p&gt;

&lt;p&gt;If you're building this in Node.js, DNotifier gives you orchestration, monitoring, and real-time infrastructure in one SDK. Explore it at &lt;a href="http://www.dnotifier.com" rel="noopener noreferrer"&gt;www.dnotifier.com&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Building a Simple Agent Runtime With Node.js</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Fri, 04 Sep 2026 08:40:51 +0000</pubDate>
      <link>https://dev.to/timalam01/building-a-simple-agent-runtime-with-nodejs-emk</link>
      <guid>https://dev.to/timalam01/building-a-simple-agent-runtime-with-nodejs-emk</guid>
      <description>&lt;p&gt;You built an agent that calls a model, picks a tool, and prints an answer. It works fine in a script. Then you try to run it for real, and it falls apart.&lt;/p&gt;

&lt;p&gt;No memory between steps. No retry when a call fails. No record of what actually happened. That's the moment every builder finds out they didn't build an agent. They built a function pretending to be one.&lt;/p&gt;

&lt;p&gt;What they actually needed was an &lt;strong&gt;agent runtime&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In this post, we'll build a small one in Node.js from scratch. No framework, no magic. Just the pieces that make an agent runtime work, so you understand what's happening under the hood before you reach for a bigger tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is an Agent Runtime, Really?
&lt;/h2&gt;

&lt;p&gt;An agent runtime is the system that keeps an agent alive between calls. It holds state, decides what step comes next, and routes work to models and tools.&lt;/p&gt;

&lt;p&gt;Without a runtime, your agent forgets everything the second the function returns. It can respond, but it can't act, retry, or remember. A runtime turns a single call into a process that runs, tracks, and finishes a task.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Loop Every Agent Execution Engine Runs
&lt;/h2&gt;

&lt;p&gt;Strip away the buzzwords and every agent execution engine does the same four things, over and over, until the task is done:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt; ┌─────────┐
 │ Perceive│  read the current state + input
 └────┬────┘
      ▼
 ┌─────────┐
 │  Decide │  ask the model what to do next
 └────┬────┘
      ▼
 ┌─────────┐
 │   Act   │  call a tool or return an answer
 └────┬────┘
      ▼
 ┌─────────┐
 │ Observe │  save the result, update state
 └────┬────┘
      │
      └──────► back to Perceive
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That loop is the entire job of an agent runtime. Everything else, memory, tools, logging, is built around keeping this loop honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting Up the Project
&lt;/h2&gt;

&lt;p&gt;Nothing fancy here. Just a plain Node project.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;mkdir &lt;/span&gt;agent-runtime &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;agent-runtime
npm init &lt;span class="nt"&gt;-y&lt;/span&gt;
npm &lt;span class="nb"&gt;install &lt;/span&gt;node-fetch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We'll keep everything in one file to start, then split it up once the pieces are clear.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the Agent Runtime Class
&lt;/h2&gt;

&lt;p&gt;This is the core of an LLM agent runtime: a class that owns the loop, the state, and a hard limit on how many steps it can take.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;AgentRuntime&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;constructor&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{},&lt;/span&gt; &lt;span class="nx"&gt;maxSteps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;maxSteps&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;maxSteps&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;history&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;maxSteps&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;step&lt;/span&gt;&lt;span class="o"&gt;++&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;final_answer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tool_call&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;callTool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
          &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tool&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="na"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;});&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Runtime stopped: max steps reached.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="nf"&gt;callTool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`No tool named &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;input&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`Tool &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; failed: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the &lt;code&gt;maxSteps&lt;/code&gt; guard. Without it, a confused model can loop forever. Every serious agent runtime architecture needs a hard stop like this.&lt;/p&gt;

&lt;h2&gt;
  
  
  Giving Your Agent Runtime Some Memory
&lt;/h2&gt;

&lt;p&gt;Right now, &lt;code&gt;state.history&lt;/code&gt; grows forever. That's fine for a demo, but a real agent state runtime needs to manage memory on purpose, not by accident.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;addMemory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="nf"&gt;trimHistory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;limit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;this&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;limit&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Call &lt;code&gt;trimHistory&lt;/code&gt; at the end of each loop. It keeps your context small and your costs predictable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring Up Tools
&lt;/h2&gt;

&lt;p&gt;Tools are just functions. Register them by name, and the runtime handles the rest.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;getWeather&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;city&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`Weather in &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;city&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;: 28°C, clear skies.`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;searchDocs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`Top result for "&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;query&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;": use the trimHistory method.`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;runtime&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;AgentRuntime&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;myModelFn&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;tools&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your model function just needs to return &lt;code&gt;{ type: "tool_call", tool, input }&lt;/code&gt; or &lt;code&gt;{ type: "final_answer", content }&lt;/code&gt;. That's the whole contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a Simple Runtime Breaks in Production
&lt;/h2&gt;

&lt;p&gt;The loop above works for a demo. It won't survive real traffic. Here's what changes once you move from a toy to a production agent runtime:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;Simple runtime&lt;/th&gt;
&lt;th&gt;Production runtime&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Failures&lt;/td&gt;
&lt;td&gt;Crashes or hangs&lt;/td&gt;
&lt;td&gt;Retries with backoff&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Visibility&lt;/td&gt;
&lt;td&gt;Console logs&lt;/td&gt;
&lt;td&gt;Full traceability of each step&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concurrency&lt;/td&gt;
&lt;td&gt;One task at a time&lt;/td&gt;
&lt;td&gt;Many agents running in parallel&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Communication&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Real-time updates via pub/sub&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging&lt;/td&gt;
&lt;td&gt;Guesswork&lt;/td&gt;
&lt;td&gt;Step-by-step monitoring&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is usually the point where teams stop hand-rolling everything. An autonomous agent runtime that handles many users needs monitoring, observability, and traceability baked in, not bolted on later.&lt;/p&gt;

&lt;p&gt;That's the gap &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;DNotifier&lt;/a&gt; is built to close. It gives you one SDK for orchestration, multi-agent coordination, and real-time pub/sub, so your agent runtime gets production features without you writing them from scratch. You still own the loop. You just stop reinventing the plumbing around it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing the Loop
&lt;/h2&gt;

&lt;p&gt;Before adding more features, write a fake model function that returns scripted decisions. Run it through your runtime and check the history at each step. If the loop behaves correctly with a fake model, it'll behave correctly with a real one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fakeModel&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;state&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;history&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;tool_call&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;tool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;getWeather&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Lahore&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;final_answer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Done checking the weather.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This kind of test catches loop bugs before they cost you an API bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What's the difference between an agent and an agent runtime?&lt;/strong&gt;&lt;br&gt;
An agent is a single decision-making call. An agent runtime is the system that runs that call repeatedly, tracks state, and manages tools and memory across steps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need a framework to build an agent runtime?&lt;/strong&gt;&lt;br&gt;
No. A small class with a loop, state, and a tool registry covers the basics. Frameworks help once you need retries, tracing, and multi-agent coordination at scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I add memory to an agent runtime?&lt;/strong&gt;&lt;br&gt;
Store history and key facts in a state object, then trim it on a schedule. Keep only what the next decision actually needs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can a simple &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;agent runtime&lt;/a&gt; handle production traffic?&lt;/strong&gt;&lt;br&gt;
Not on its own. You'll need retries, monitoring, and concurrency handling. That's usually when teams bring in a platform like DNotifier instead of building it by hand.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Migrating From LangChain to DNotifier: A Practical Guide</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Sat, 29 Aug 2026 18:56:23 +0000</pubDate>
      <link>https://dev.to/timalam01/migrating-from-langchain-to-dnotifier-a-practical-guide-3hie</link>
      <guid>https://dev.to/timalam01/migrating-from-langchain-to-dnotifier-a-practical-guide-3hie</guid>
      <description>&lt;p&gt;Anyone who's built something real with LangChain knows how this goes. The first prototype comes together fast, almost too fast. Then you try to add memory. Or hook up a second agent. Or figure out why a chain failed silently in production, and you're three layers deep in a stack trace that tells you nothing useful.&lt;/p&gt;

&lt;p&gt;That's usually when people start looking at a LangChain to &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;DNotifier&lt;/a&gt; migration. Not because LangChain is bad. It's just not built for what most teams actually need now: multiple agents working together, visibility into what they're doing, and something that doesn't quietly break at 2am.&lt;/p&gt;

&lt;p&gt;Here's what actually changes when you make the switch, step by step, no fluff.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Teams Are Leaving LangChain
&lt;/h2&gt;

&lt;p&gt;It's rarely one big problem. It's a dozen small ones that pile up.&lt;/p&gt;

&lt;p&gt;Chains get hard to follow past a certain point. Memory feels bolted on rather than built in. And the second you add another agent, you're writing glue code just to keep the two of them talking.&lt;/p&gt;

&lt;p&gt;DNotifier gets rid of that glue code. One SDK, one API, and orchestration, memory, and agent communication are already handled for you instead of something you duct-tape together yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Actually Changes With an AI Orchestration Platform
&lt;/h2&gt;

&lt;p&gt;The mental model shifts. In LangChain you're chaining function calls and hoping the flow holds together under load. With an AI orchestration platform like DNotifier, you define the workflow up front and let the platform handle execution, retries, and state.&lt;/p&gt;

&lt;p&gt;Sounds small. Isn't. When something breaks at 2am, you want a system that tells you exactly which step failed, not a stack trace and a guess.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Map Your Chains to Workflows&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Start dumb. List every chain you're running right now. Inputs, outputs, whatever tools it calls.&lt;/p&gt;

&lt;p&gt;Each one becomes a workflow in DNotifier. You're not rewriting your logic from scratch, you're translating a sequence of calls into something the platform can actually monitor.&lt;/p&gt;

&lt;p&gt;Migrate your busiest chain first. Test it end to end before touching anything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2: Migrate Memory and State&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is where most migrations get real, honestly. LangChain's memory objects work fine, but they're easy to lose track of once you're spanning sessions or multiple agents.&lt;/p&gt;

&lt;p&gt;DNotifier handles agent memory and state management natively. No passing memory objects between functions by hand, your workflow reads and writes state straight through the platform. Context sticks across sessions without you babysitting it.&lt;/p&gt;

&lt;p&gt;If anything you're building talks to customers directly, this step alone justifies the switch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3: Move Your RAG Pipeline&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Good news here: the core logic barely changes. Chunking, embedding, retrieval, all of it carries over.&lt;/p&gt;

&lt;p&gt;What changes is where the pipeline lives. DNotifier connects to your vector database and folds retrieval into the workflow itself, so a RAG pipeline doesn't need its own separate orchestration layer bolted on top. Still, check your retrieval quality after moving it. Settings usually carry over clean, but "usually" isn't "always."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4: Set Up Multi-Agent Systems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Honestly, this is the real reason most people migrate. Getting agents to coordinate in LangChain means building your own message-passing system and crossing your fingers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;DNotifier&lt;/a&gt; treats agent orchestration as a core piece, not an afterthought. You define each agent's role, and the platform handles the handoffs and keeps everyone in sync. Research agents, support agents, whatever the team looks like, the coordination layer already exists. You're configuring it, not building it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 5: Add Observability and Traceability&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once workflows are live, you need to see what's happening inside them. LangChain gives you logs, which is fine until it isn't.&lt;/p&gt;

&lt;p&gt;DNotifier gives you observability and traceability across every agent and every workflow. You can see which step failed, what data hit it, and why the decision went the way it did. In production, that's not a nice-to-have. It's the difference between a fragile demo and something a team can actually trust.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mistakes People Make Mid-Migration
&lt;/h2&gt;

&lt;p&gt;Don't move everything at once. One workflow, confirm it works, then the next one. Trying to flip the whole stack in a single sprint tends to create bugs nobody meant to write.&lt;/p&gt;

&lt;p&gt;And don't rush the memory and state piece. It's the part people skip, and it's the part users notice first when it's wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is DNotifier used for?&lt;/strong&gt;&lt;br&gt;
It's used to build, run, and monitor agents and workflows in one place. Basically the orchestration layer most teams end up hand-building on top of LangChain anyway.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is DNotifier good for production?&lt;/strong&gt;&lt;br&gt;
Yes. Built-in observability, traceability, and state handling are exactly what production needs and LangChain doesn't give you by default.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need to rewrite my RAG pipeline from scratch?&lt;/strong&gt;&lt;br&gt;
No. Your chunking and embedding logic usually stays as-is. What moves is the orchestration sitting on top of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long does a migration actually take?&lt;/strong&gt;&lt;br&gt;
Depends how many chains and agents you're running. Migrating one workflow at a time, figure a few days per workflow, not weeks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;Moving from LangChain to DNotifier isn't throwing out what you've built. It's giving it a foundation that can handle real traffic, more than one agent, and the debugging you're going to need eventually anyway.&lt;/p&gt;

&lt;p&gt;Start with one workflow. See how it feels. Go from there.&lt;/p&gt;

&lt;p&gt;Want to see what your migration looks like in practice? Explore the SDK at dnotifier.com.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Building an AI Data Analyst Agent With DNotifier</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Wed, 26 Aug 2026 07:34:18 +0000</pubDate>
      <link>https://dev.to/timalam01/building-an-ai-data-analyst-agent-with-dnotifier-ad2</link>
      <guid>https://dev.to/timalam01/building-an-ai-data-analyst-agent-with-dnotifier-ad2</guid>
      <description>&lt;p&gt;Analyzing raw enterprise data takes hours of manual SQL queries, dashboard creation, and endless script tweaks. Traditional setups break when schema definitions shift or data pipelines hit unforeseen bottlenecks. Engineering teams waste time stitching together custom scripts instead of focusing on core architecture.&lt;/p&gt;

&lt;p&gt;Building an &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;&lt;strong&gt;AI data analyst agent&lt;/strong&gt;&lt;/a&gt; with DNotifier changes how teams query internal databases and parse complex datasets. By leveraging the right AI agent infrastructure, developers can build production AI agents that handle intent routing, execute structured data tasks, and maintain persistent system memory without complex setups.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is an AI Data Analyst Agent?
&lt;/h2&gt;

&lt;p&gt;An AI data analyst agent is an autonomous digital worker that processes natural language queries, inspects raw database schemas, executes precise data retrieval commands, and generates structured analytical summaries.&lt;/p&gt;

&lt;p&gt;Unlike basic query generators, an autonomous AI data analyst agent works as a dedicated system. It evaluates edge cases, retries failed executions safely, and translates complex raw records into readable, executive-ready insights.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Use DNotifier for Agent Orchestration?
&lt;/h2&gt;

&lt;p&gt;Most open-source tools require complex glue code for routing, session logging, and state synchronization. DNotifier solves this by providing a unified AI agent platform with native agent runtime management, built-in vector databases, and real-time pub/sub features.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building the Data Analyst System Architecture
&lt;/h2&gt;

&lt;p&gt;An enterprise-ready AI data analyst agent relies on a multi-agent framework where specialized nodes work together. The system divides analytical workloads across four primary steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Intent Router Agent:&lt;/strong&gt; Receives the natural language request from the user, determines the core analytical goal, and routes the task to appropriate tools.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Schema Retrieval Pipeline:&lt;/strong&gt; Uses vector search to locate the exact database tables, column names, and metric definitions required for the query.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Execution Engine:&lt;/strong&gt; Safely constructs validated SQL statements or data retrieval scripts and executes them against your database.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Insight Synthesizer:&lt;/strong&gt; Interprets the raw dataset returned from the database and packages it into structured business reports.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Step-by-Step Implementation Guide
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1.Initialize the DNotifier SDK:&lt;/strong&gt;&lt;br&gt;
Set up environment credentials and import the core libraries.Import the SDK into your project environment and configure your secret key to authorize connection with the managed production runtime.&lt;br&gt;
&lt;strong&gt;2.Configure Schema Retrieval (RAG Pipeline):&lt;/strong&gt;&lt;br&gt;
Index enterprise database definitions for semantic lookup.Upload your data warehouse schema definitions, metric rules, and table structures into the &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;DNotifier&lt;/a&gt; document store to enable context-aware query building.&lt;br&gt;
&lt;strong&gt;3.Define Specialized Agents:&lt;/strong&gt;&lt;br&gt;
Set up router and analytics roles within the workflow.Establish dedicated agent personas within your workflow—assigning specific responsibility roles for parsing user intent and translating context into execution statements.&lt;br&gt;
&lt;strong&gt;4.Build and Execute the Orchestrated Workflow:&lt;/strong&gt;&lt;br&gt;
Chain agent execution and monitor output in real time.Connect your agents into a unified sequence. DNotifier automatically handles session context, data handoffs between steps, and real-time observability logging.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is DNotifier an AI agent framework?&lt;/strong&gt;&lt;br&gt;
Yes, DNotifier is an enterprise-grade AI agent framework that combines multi-agent orchestration, managed RAG pipelines, and real-time observability in a single platform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How does DNotifier handle state management across agent?&lt;/strong&gt;&lt;br&gt;
DNotifier provides managed sessions and workflow context that automatically persist state, variable handoffs, and session history across multi-agent pipelines.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I monitor agent actions and LLM calls in real time?&lt;/strong&gt;&lt;br&gt;
Yes, DNotifier includes real-time tracing, workflow execution graphs, and dashboard analytics to monitor latency, tool execution, and prompt logs.Explore the platform at dnotifier.com to start building production-ready data agents today.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>What is DNotifier and how does it work in production?</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Thu, 20 Aug 2026 16:47:38 +0000</pubDate>
      <link>https://dev.to/timalam01/what-is-dnotifier-and-how-does-it-work-in-production-2ak7</link>
      <guid>https://dev.to/timalam01/what-is-dnotifier-and-how-does-it-work-in-production-2ak7</guid>
      <description>&lt;p&gt;Building an autonomous AI agent in a Jupyter notebook feels amazing. Getting Deploying &lt;strong&gt;DNotifier Agents to Production&lt;/strong&gt; right for active users is a completely different challenge. Local prototypes rarely handle network spikes, broken state, or rogue API calls.&lt;/p&gt;

&lt;p&gt;Moving from a local demo to a enterprise-ready application requires reliable AI agent infrastructure. You need clear trace logs, instant state management, and tight LLM orchestration.&lt;/p&gt;

&lt;p&gt;Here is how you can deploy your &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;DNotifier &lt;/a&gt;AI agent workforce to live infrastructure safely and cleanly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy9h8stb0zlpk0t4y2eiz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy9h8stb0zlpk0t4y2eiz.png" alt=" " width="799" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What is DNotifier and how does it work in production?
&lt;/h2&gt;

&lt;p&gt;DNotifier is a lightweight AI agent framework designed to simplify Deploying DNotifier Agents to Production. It combines model routing, state management, and real-time pub/sub messaging into a unified SDK and API layer.&lt;/p&gt;

&lt;p&gt;Instead of chaining together multiple libraries, DNotifier acts as all-in-one AI middleware. It handles background job queues, multi-agent communication, and LLM observability through a simple single-entry setup.&lt;/p&gt;

&lt;p&gt;When users interact with your system, the DNotifier runtime orchestrates execution across your agents automatically.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Initialize your production AI agent with the DNotifier SDK
from dnotifier import DNotifierClient, Agent

client = DNotifierClient(api_key="dn_live_key")

agent = Agent(
    name="CustomerSupportAgent",
    model="gpt-4o",
    tools=["knowledge_base_search", "refund_calculator"],
    persistence=True
)

# Publish your agent workflow to production
client.deploy(agent, env="production")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;1. Set up persistent state and agent memory&lt;/strong&gt;&lt;br&gt;
Production agents fail when they forget context mid-conversation. Basic prototypes keep short-term memory in RAM, which drops every time your cloud container restarts.&lt;/p&gt;

&lt;p&gt;Production AI agents require durable memory backends. DNotifier manages state persistence across user sessions out of the box using built-in database hooks.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enable state storage in your config file before deployment.&lt;/li&gt;
&lt;li&gt;Assign unique session IDs to every user request.&lt;/li&gt;
&lt;li&gt;Store system prompts in prompt management layers instead of hardcoding text.&lt;/li&gt;
&lt;li&gt;Use vector database integrations for long-term semantic context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;2. Configure real-time pub/sub for multi-agent workflows&lt;/strong&gt;&lt;br&gt;
When multiple autonomous AI agents work together, direct API calls quickly become a web of unmaintainable code. Production multi-agent platform architecture relies on event-driven communication.&lt;/p&gt;

&lt;p&gt;Using real-time pub/sub, your researchers, writers, and code agents talk through dedicated message buses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Implement human-in-the-loop safeguards&lt;/strong&gt;&lt;br&gt;
Autonomous execution can cause unwanted behavior if left completely unsupervised. High-risk actions—like issuing refunds or sending live emails—demand human review before execution.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;DNotifier includes &lt;strong&gt;&lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;human-in-the-loop&lt;/a&gt;&lt;/strong&gt; workflows at the runtime level.&lt;/li&gt;
&lt;li&gt;Configure action policies inside your agent definition.&lt;/li&gt;
&lt;li&gt;Flag high-impact tools as requires_approval=True.&lt;/li&gt;
&lt;li&gt;Route pending actions to a review dashboard using real-time alerts.&lt;/li&gt;
&lt;li&gt;Resume execution once an admin approves the action.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;4. Set up AI observability and prompt testing&lt;/strong&gt;&lt;br&gt;
You cannot optimize what you do not trace. Debugging non-deterministic LLM chains requires granular log tracking for every agent step.&lt;/p&gt;

&lt;p&gt;DNotifier logs token usage, step latency, tool inputs, and raw model outputs automatically.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[INFO] Agent Run ID: run_98234x
[TRACE] Step 1: Retrieval augmented generation pipeline queried successfully. (42ms)
[TRACE] Step 2: Tool `refund_calculator` called with args: {"user_id": 402}. (112ms)
[SUCCESS] Execution finished. Total Tokens: 412
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use prompt testing environments to evaluate new system prompts against real production traces before pushing updates live.&lt;/p&gt;

&lt;h2&gt;
  
  
  DNotifier vs LangChain for production workloads
&lt;/h2&gt;

&lt;p&gt;Teams often compare DNotifier vs LangChain when choosing an AI orchestration platform. While open-source chaining tools work well for early experiments, DNotifier is built specifically for production stability and easy maintenance.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;LangChain:&lt;/strong&gt; Flexible, huge library ecosystem, but complex to maintain in large apps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DNotifier:&lt;/strong&gt; Unified API, native pub/sub, built-in monitoring, and lower deployment overhead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want a single SDK that replaces fragmented tracing tools, state databases, and task runners, DNotifier gives you a cleaner path to launch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is DNotifier an AI agent framework?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Yes, DNotifier is a complete AI agent framework that provides tools, memory management, and orchestration for production workloads.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I build a RAG application with DNotifier?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Connect your vector database to the DNotifier document loader to create an event-driven retrieval pipeline in a few lines of code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is DNotifier good for production?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Yes, DNotifier is engineered specifically for production environments with built-in tracing, failovers, and multi-agent pub/sub messaging.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Build DNotifier Human in the Loop Workflows for Production AI</title>
      <dc:creator>T. Alam</dc:creator>
      <pubDate>Wed, 19 Aug 2026 08:40:29 +0000</pubDate>
      <link>https://dev.to/timalam01/how-to-build-dnotifier-human-in-the-loop-workflows-for-production-ai-277n</link>
      <guid>https://dev.to/timalam01/how-to-build-dnotifier-human-in-the-loop-workflows-for-production-ai-277n</guid>
      <description>&lt;p&gt;Fully autonomous agents fail in production when edge cases break business rules. You need humans to review risky decisions without slowing down your &lt;strong&gt;AI agent workflow&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Implementing &lt;strong&gt;DNotifier human in the loop&lt;/strong&gt; patterns gives you safety and control. This guide shows you how to pause &lt;strong&gt;autonomous AI agents&lt;/strong&gt;, request approval, and resume execution cleanly.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Human-in-the-Loop in AI Workflows?
&lt;/h2&gt;

&lt;p&gt;Human-in-the-loop (HITL) pauses an AI agent execution path until a real person approves, rejects, or edits the state. It prevents hallucinated actions from reaching production environments.&lt;/p&gt;

&lt;p&gt;Instead of letting an AI writer agent publish content automatically, HITL routes the draft to a manager. The workflow resumes only after explicit authorization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Use DNotifier for HITL Orchestration?
&lt;/h2&gt;

&lt;p&gt;Traditional &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;&lt;strong&gt;AI agent frameworks&lt;/strong&gt;&lt;/a&gt; force you to write custom polling loops or manage external databases for paused states. This adds fragile boilerplate code to your &lt;strong&gt;AI infrastructure&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;dnotifier framework&lt;/strong&gt; simplifies state suspension using an &lt;strong&gt;event-driven agent&lt;/strong&gt; runtime. &lt;br&gt;
&lt;strong&gt;Native State Suspension:&lt;/strong&gt; Freeze the execution state without losing event context.&lt;br&gt;
&lt;strong&gt;Real-Time Pub/Sub:&lt;/strong&gt; Stream review requests directly to your human UI.&lt;br&gt;
&lt;strong&gt;Unified Observability:&lt;/strong&gt; Trace every prompt, model response, and human intervention in one audit log.&lt;br&gt;
Comparing &lt;strong&gt;LangChain&lt;/strong&gt; vs &lt;strong&gt;DNotifier&lt;/strong&gt;, dnotifier ai handles messaging and state natively in one AI SDK.&lt;/p&gt;
&lt;h2&gt;
  
  
  Step-by-Step: Implementing Human Approval with DNotifier
&lt;/h2&gt;

&lt;p&gt;Here is how to set up human verification for an &lt;strong&gt;AI automation agents&lt;/strong&gt; system using the &lt;strong&gt;DNotifier SDK&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 1: Initialize the DNotifier Client&lt;/strong&gt;&lt;br&gt;
Set up your connection using the &lt;strong&gt;DNotifier agent framework&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;DNotifier&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@dnotifier/sdk&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;dnotifier&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;DNotifier&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;appId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_APP_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DNOTIFIER_APP_SECRET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2: Define the Suspended Workflow State&lt;/strong&gt;&lt;br&gt;
When your &lt;strong&gt;AI agent workflow framework&lt;/strong&gt; hits a sensitive step, pause execution and emit a review event.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;processRefund&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;workflow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;dnotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;workflows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Refund Processing&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;amount&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// Pause workflow and request human approval&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;dnotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;publish&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;channel&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;human-approvals&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;approval_required&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;workflowId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PENDING_HUMAN_REVIEW&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PAUSED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;workflowId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;executeRefund&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;userId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 3: Handle the Human Decision Signal&lt;/strong&gt;&lt;br&gt;
When the human manager approves the action in your dashboard, send a resume signal back to the &lt;strong&gt;AI orchestrator&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;handleHumanDecision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;workflowId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;approved&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;boolean&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;approved&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;dnotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;workflows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;resume&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;workflowId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;APPROVED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Workflow &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;workflowId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt; resumed by operator.`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;dnotifier&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;workflows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;cancel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;workflowId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Rejected by human reviewer&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This pattern ensures safe execution without manual database state stitching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architectural Patterns for Human-Agent Collaboration
&lt;/h2&gt;

&lt;p&gt;Different business problems need different &lt;strong&gt;AI agent architecture&lt;/strong&gt;&lt;br&gt;
patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The Gatekeeper PatternThe&lt;/strong&gt;&lt;br&gt;
 AI agent processes tasks autonomously until a risk threshold is met. High-value transfers or public communications pause for human sign-off. &lt;br&gt;
&lt;strong&gt;2. The Interactive Copilot Pattern&lt;/strong&gt;&lt;br&gt;
The human and AI customer support agents work together in real-time. The agent drafts responses while the human edits before sending.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is DNotifier used for?&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;DNotifier&lt;/strong&gt; is a unified &lt;a href="https://www.dnotifier.com/" rel="noopener noreferrer"&gt;&lt;strong&gt;AI infrastructure&lt;/strong&gt;&lt;/a&gt; platform providing orchestration, real-time messaging, and multi-agent coordination. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is DNotifier good for production?&lt;/strong&gt;&lt;br&gt;
Yes, &lt;strong&gt;dnotifier production&lt;/strong&gt; deployments scale reliably using event-driven real-time infrastructure and multi-model support.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do I build an AI agent with DNotifier?&lt;/strong&gt;&lt;br&gt;
Initialize the SDK, define model roles, attach enterprise data sources, and trigger execution using the &lt;strong&gt;dnotifier tutorial&lt;/strong&gt; docs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I build RAG workflows with DNotifier?&lt;/strong&gt;&lt;br&gt;
Yes, you can build a &lt;strong&gt;full RAG pipeline&lt;/strong&gt; using the built-in &lt;strong&gt;DNotifier vector database&lt;/strong&gt; capabilities.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
