<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Seenivasa Ramadurai</title>
    <description>The latest articles on DEV Community by Seenivasa Ramadurai (@sreeni5018).</description>
    <link>https://dev.to/sreeni5018</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1829954%2Fe57edf87-9dae-48c9-a528-0f57f54aac70.png</url>
      <title>DEV Community: Seenivasa Ramadurai</title>
      <link>https://dev.to/sreeni5018</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sreeni5018"/>
    <language>en</language>
    <item>
      <title>Jev Decides, OpenAI Reasons: A Practical Hybrid Agent Architecture</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Tue, 22 Sep 2026 02:14:40 +0000</pubDate>
      <link>https://dev.to/sreeni5018/jev-decides-openai-reasons-a-practical-hybrid-agent-architecture-4ol3</link>
      <guid>https://dev.to/sreeni5018/jev-decides-openai-reasons-a-practical-hybrid-agent-architecture-4ol3</guid>
      <description>&lt;p&gt;&lt;strong&gt;Code Controls, Jev Decides, OpenAI Reasons: A Hybrid Agent Architecture That Doesn't Waste LLM Calls&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4ba8zco4bysyavrdi8n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa4ba8zco4bysyavrdi8n.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When we build an AI agent today, the default architecture puts an LLM in the middle of everything. The user sends a request, the LLM interprets it, reasons about what to do, picks a tool, generates arguments, reads the result, reasons again, and finally responds.&lt;/p&gt;

&lt;p&gt;That approach works. But while experimenting with agents, I started asking a different question: &lt;strong&gt;does every decision inside an agent really need a general purpose LLM?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Consider a simple request:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I forgot my password."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The application already knows the possible workflows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;reset_password&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;unlock_account&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;check_order&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;cancel_order&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;openai_reasoning&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;human_support&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We're not asking AI to invent a new solution. We just need it to answer: &lt;strong&gt;which of these known capabilities best matches the request?&lt;/strong&gt; That's a very different problem from asking an LLM to investigate an outage, explain an architecture, or compare options.&lt;/p&gt;

&lt;p&gt;This distinction led me to a hybrid agent architecture using &lt;strong&gt;Jev&lt;/strong&gt; and &lt;strong&gt;OpenAI&lt;/strong&gt;. The basic idea:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CODE    = CONTROL
JEV     = DECIDE
OPENAI  = REASON
TOOLS   = ACT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of asking one model to do everything, the agent uses the right kind of intelligence for each step.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Jev?
&lt;/h2&gt;

&lt;p&gt;Jev is TypeSafe AI's first public &lt;strong&gt;System One Model&lt;/strong&gt; a model built specifically for decisions inside software, not for open ended generation.&lt;/p&gt;

&lt;p&gt;A general-purpose LLM works like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Context → LLM → Generate tokens → Explanation / Code / JSON / Tool Call / Answer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev works differently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application State → Jev → Typed Probabilistic Decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;TypeSafe describes this as &lt;em&gt;unstructured state in, typed probabilistic decisions out&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Suppose our application sends the user request plus the available capabilities. Jev might return something conceptually like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;reset_password&lt;/span&gt;       &lt;span class="err"&gt;1.00&lt;/span&gt;
&lt;span class="err"&gt;unlock_account&lt;/span&gt;       &lt;span class="err"&gt;0.00&lt;/span&gt;
&lt;span class="err"&gt;check_order&lt;/span&gt;          &lt;span class="err"&gt;0.00&lt;/span&gt;
&lt;span class="err"&gt;cancel_order&lt;/span&gt;         &lt;span class="err"&gt;0.00&lt;/span&gt;
&lt;span class="err"&gt;openai_reasoning&lt;/span&gt;     &lt;span class="err"&gt;0.00&lt;/span&gt;
&lt;span class="err"&gt;human_support&lt;/span&gt;        &lt;span class="err"&gt;0.00&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application executes &lt;code&gt;reset_password()&lt;/code&gt;. Jev doesn't write a paragraph explaining what to do — its job is the &lt;strong&gt;bounded semantic judgment&lt;/strong&gt;: which known capability best fits this context?&lt;/p&gt;

&lt;h2&gt;
  
  
  Jev Is Not Just a Smaller LLM
&lt;/h2&gt;

&lt;p&gt;It's easy to assume Jev is a cheaper model used for classification, but that misses the point.&lt;/p&gt;

&lt;p&gt;A general-purpose LLM is optimized for flexible language generation it can explain, reason, summarize, write, generate code, compare alternatives, plan, and troubleshoot. Jev focuses on a narrower interface: decisions whose structure the application already defines. TypeSafe's workflow examples use three decision primitives &lt;strong&gt;Choice&lt;/strong&gt;, &lt;strong&gt;Score&lt;/strong&gt;, and &lt;strong&gt;Noul&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Choice&lt;/strong&gt; - "Which team? Billing / Shipping / Technical / Account"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Score&lt;/strong&gt; - "How severe is this issue? Low / Medium / High / Critical"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Probabilistic yes/no&lt;/strong&gt; -"Should this transaction be escalated? Yes → 0.91, No → 0.09"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The application defines the decision space. Jev is most useful when the system already knows the possible actions but still needs semantic understanding to pick the right one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Jev vs. an LLM
&lt;/h2&gt;

&lt;p&gt;Two requests make the difference clear.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Request 1:&lt;/strong&gt; "I forgot my password." The possible workflows are already known this is a bounded decision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Request 2:&lt;/strong&gt; "Our application authenticates successfully through SSO, but users start getting errors when their OAuth access token expires. Analyze what could be happening and explain what we should investigate."&lt;/p&gt;

&lt;p&gt;There's no predefined answer here. The system might need to think about refresh tokens, token rotation, scopes, audience, session expiration, IdP configuration, or token caching. That's open ended reasoning.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;th&gt;General-Purpose LLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary role&lt;/td&gt;
&lt;td&gt;Decision&lt;/td&gt;
&lt;td&gt;Reasoning + generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answer space&lt;/td&gt;
&lt;td&gt;Defined beforehand&lt;/td&gt;
&lt;td&gt;Potentially open-ended&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Typed decisions&lt;/td&gt;
&lt;td&gt;Generated content&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Probabilities&lt;/td&gt;
&lt;td&gt;Core part of interface&lt;/td&gt;
&lt;td&gt;Possible, not primary&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long explanation&lt;/td&gt;
&lt;td&gt;Not the purpose&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semantic routing&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;td&gt;Possible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool selection from known set&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;td&gt;Possible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Troubleshooting&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content creation&lt;/td&gt;
&lt;td&gt;Not the purpose&lt;/td&gt;
&lt;td&gt;Strong fit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Planning&lt;/td&gt;
&lt;td&gt;Bounded&lt;/td&gt;
&lt;td&gt;Open-ended&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The right question isn't &lt;em&gt;Jev versus OpenAI, which is better&lt;/em&gt; it's &lt;em&gt;which one fits the decision I'm making right now?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Build a Hybrid Agent?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv4oy63dcyzwhtapf8mpi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv4oy63dcyzwhtapf8mpi.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent doesn't contain only one kind of problem. Inside the same interaction we usually have deterministic rules, semantic decisions, open ended reasoning, and real world actions. Forcing all four through the same model is wasteful. Instead, separate the responsibilities:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    USER
                      │
                      ▼
              CONVERSATION STATE
                      │
                      ▼
                     JEV
                Decide / Route
                      │
        ┌─────────────┼──────────────┐
        │             │              │
        ▼             ▼              ▼
 Deterministic     OpenAI          Human
   Workflow        Reasoning       Support
        │             │
        └──────┬──────┘
               │
               ▼
         TOOLS &amp;amp; SYSTEMS
               │
               ▼
           UPDATE STATE
               │
               ▼
         RESPONSE TO USER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Figure 1&lt;/strong&gt; Jev decides, OpenAI reasons, code orchestrates, and tools execute.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let's Build It
&lt;/h2&gt;

&lt;p&gt;The implementation is surprisingly small. Start with two API keys.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Environment Configuration
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="py"&gt;TYPESAFE_API_KEY&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;your_typesafe_api_key&lt;/span&gt;
&lt;span class="py"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;your_openai_api_key&lt;/span&gt;
&lt;span class="py"&gt;OPENAI_MODEL&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;gpt-6-astra&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;dotenv&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_dotenv&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="nf"&gt;load_dotenv&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;TYPESAFE_API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TYPESAFE_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;OPENAI_API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;OPENAI_MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENAI_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-astra&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;openai_client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenAI's current Responses API uses &lt;code&gt;client.responses.create(...)&lt;/code&gt; and exposes the generated text through &lt;code&gt;response.output_text&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Create a Capability Registry
&lt;/h3&gt;

&lt;p&gt;Instead of a giant &lt;code&gt;if/elif&lt;/code&gt; routing block, register capabilities so they can describe themselves:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;TOOLS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;register_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decorator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;func&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;TOOLS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;function&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;func&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;func&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;decorator&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@register_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reset_password&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use when the user forgot their password or wants to reset their password.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reset_password&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Password reset workflow started.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="nd"&gt;@register_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;unlock_account&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use when the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s account is locked or needs to be unlocked.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;unlock_account&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Account unlock workflow started.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We can also register the LLM itself as a capability:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@register_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai_reasoning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use when the request requires explanation, analysis, comparison, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;troubleshooting, summarization, writing, or open-ended reasoning.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;openai_reasoning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the part I find most interesting: OpenAI becomes &lt;strong&gt;one capability&lt;/strong&gt; available to the agent, rather than automatically being the whole agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3:  Dynamically Build Jev's Choices
&lt;/h3&gt;

&lt;p&gt;The tool descriptions &lt;em&gt;become&lt;/em&gt; Jev's decision space:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;build_jev_choices&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;TOOLS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;items&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add a new capability — say &lt;code&gt;refund_order&lt;/code&gt; — and it automatically becomes another candidate. No router rewrite needed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Let Jev Route the Request
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typesafe_sdk&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Choice&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TypeSafeClient&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;route_with_jev&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;choices&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;build_jev_choices&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user_request&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;available_capabilities&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;list&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;keys&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nc"&gt;TypeSafeClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;system_one&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;questions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;route&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Choice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                    &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Choose the single best capability for handling the &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s request. Use openai_reasoning when the task &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;requires explanation, analysis, troubleshooting, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;comparison, writing, or open-ended reasoning.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="p"&gt;),&lt;/span&gt;
                    &lt;span class="n"&gt;criteria&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;
                &lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;route&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev's job is now narrow: given the request and the available capabilities, return the best one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example — password reset:&lt;/strong&gt; the user says "I forgot my password," Jev sees the six capabilities, and selects &lt;code&gt;reset_password&lt;/code&gt; with a confidence of 1.0. Job done.&lt;/p&gt;

&lt;h2&gt;
  
  
  But Routing Alone Doesn't Make an Agent
&lt;/h2&gt;

&lt;p&gt;My first implementation just returned "Password reset workflow started." Technically correct, not very useful. A real agent needs to continue the conversation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You:   I forgot my password.
Agent: I can help with that. What email address is associated with your account?
You:   sreeni@example.com
Agent: I've initiated the password-reset workflow. Please check your registered email.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That means the agent needs &lt;strong&gt;state&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Add Conversation State
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SESSION&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending_tool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending_step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{},&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;history&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After Jev chooses &lt;code&gt;reset_password&lt;/code&gt;, the app stores &lt;code&gt;pending_tool = "reset_password"&lt;/code&gt; and &lt;code&gt;pending_step = "awaiting_email"&lt;/code&gt;. When the user replies with an email address, we don't send that back to Jev the application already knows what it means, because it's answering the question the active workflow just asked.&lt;/p&gt;

&lt;p&gt;This led me to a principle I keep coming back to: &lt;strong&gt;don't ask AI to decide something your application already knows.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example — order cancellation&lt;/strong&gt; follows the same shape: Jev selects &lt;code&gt;cancel_order&lt;/code&gt; once, up front, and everything after that (asking for the order number, confirming, executing) is normal software. We don't need an LLM to understand "Yes."&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6: Add OpenAI as the Reasoning Capability
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@register_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai_reasoning&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use when the user&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s request requires open-ended reasoning, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;explanation, analysis, comparison, summarization, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;troubleshooting, or generation.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;openai_reasoning&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;openai_client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;OPENAI_MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;instructions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are the reasoning component inside a hybrid AI agent. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Provide a clear and concise answer.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output_text&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenAI is no longer called for every request. Jev decides, up front, whether this matches a known workflow or whether it actually needs open ended reasoning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example a request that actually needs an LLM:&lt;/strong&gt; "Our application works with SSO initially, but starts failing after the access token expires. Explain what could be wrong and what we should investigate." Jev evaluates the six capabilities and selects &lt;code&gt;openai_reasoning&lt;/code&gt;. Only now does OpenAI get involved, and it can discuss refresh-token expiration, token rotation, scope problems, audience, session expiration, IdP configuration, and caching behavior exactly the kind of work where an LLM earns its cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  Two Architectures, Side by Side
&lt;/h3&gt;

&lt;p&gt;A conventional LLM-centric agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;USER → LLM understands → LLM reasons → LLM chooses tool → TOOL
     → LLM reads result → LLM reasons again → ANSWER
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The hybrid architecture:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;USER → JEV → Which capability?
              ├── Deterministic → CODE
              ├── Open-ended    → OPENAI
              └── Uncertain     → HUMAN
            → TOOLS → RESULT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither is universally better. They optimize for different workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 7: The Core Agent Loop
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;

    &lt;span class="c1"&gt;# Continue an existing workflow
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;SESSION&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending_tool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;execute_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;SESSION&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending_tool&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
            &lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;SESSION&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending_step&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Otherwise ask Jev to route
&lt;/span&gt;    &lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;route_with_jev&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;selected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choice&lt;/span&gt;
    &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;confidence&lt;/span&gt;

    &lt;span class="c1"&gt;# Low confidence: ask for clarification
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;m not completely sure what you&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;d like me to do. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Could you provide a little more detail?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Execute selected capability
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;execute_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;selected&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;user_message&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole architecture in one function: is a workflow already active? If so, continue it no LLM call needed. If not, ask Jev which capability fits, then route to code, OpenAI, or a human.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Confidence Fits
&lt;/h2&gt;

&lt;p&gt;One thing I like about Jev's interface is that decisions come with &lt;strong&gt;probabilities and confidence&lt;/strong&gt; as first-class outputs, not an afterthought.&lt;/p&gt;

&lt;p&gt;If Jev returns &lt;code&gt;**reset_password**: 0.97&lt;/code&gt;, the app can safely continue. But if it returns something closer to &lt;code&gt;**reset_password**: 0.43, **unlock_account**: 0.39, **human_support**: 0.18&lt;/code&gt;, the harness shouldn't just pick the highest score blindly it can ask a clarifying question instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.85&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.50&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;ask_for_clarification&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;escalate_to_human&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact thresholds need to be tuned per use case rather than treated as universal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Typed Doesn't Mean Semantically Correct
&lt;/h2&gt;

&lt;p&gt;There's an important limitation here. If the allowed outputs are &lt;code&gt;**reset_password**&lt;/code&gt;, &lt;code&gt;**unlock_account**&lt;/code&gt;, and &lt;code&gt;**human_support**&lt;/code&gt;, Jev will always return something structurally valid but a structurally valid decision can still be the &lt;em&gt;wrong&lt;/em&gt; one. The model might select &lt;code&gt;**reset_password**&lt;/code&gt; when &lt;code&gt;**unlock_account**&lt;/code&gt; was correct. Both are valid outputs; only one matches reality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Typed output does not mean guaranteed correctness.&lt;/strong&gt; The application still needs evaluation, confidence policies, validation, guardrails, human escalation, observability, and testing. That doesn't go away just because the decision is typed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four Responsibilities, Not One "AI Problem"
&lt;/h2&gt;

&lt;p&gt;The experiment got clearer once I stopped treating everything as an AI problem. There are really four distinct responsibilities:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CODE = CONTROL&lt;/strong&gt; — use normal code when the rule is already known (&lt;code&gt;if order.status == "SHIPPED": cancellation_allowed = False&lt;/code&gt;). There's no reason to ask a model to rediscover a deterministic business rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;JEV = DECIDE&lt;/strong&gt; — use Jev when the possible outcomes are known but semantic understanding is required: which workflow, which tool, which specialist agent, relevant or not, retry or stop, escalate or continue. This is &lt;strong&gt;bounded semantic judgment&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OPENAI = REASON&lt;/strong&gt; — use a general-purpose LLM when the problem requires analysis, explanation, planning, generation, comparison, or synthesis for example, "analyze the last five production incidents and identify recurring failure patterns." That's not choosing from a menu; it requires reasoning across information.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TOOLS = ACT&lt;/strong&gt; — tools perform the real-world action: Okta, Entra ID, ServiceNow, Jira, Salesforce, OpenSearch, Qdrant, databases, REST APIs, MCP servers, A2A agents. The model shouldn't just say "the password has been reset" unless the identity provider actually did it.&lt;/p&gt;

&lt;p&gt;Above all four sits the &lt;strong&gt;agent harness&lt;/strong&gt;, which orchestrates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Gets More Interesting
&lt;/h2&gt;

&lt;p&gt;Password reset is intentionally a simple example. The architecture gets more interesting in real agent systems, where an agent is repeatedly asking bounded questions: which tool should I use, which MCP server should receive this request, is this retrieved document relevant, should I retry this failed operation, which specialist agent should handle this task, should this action require human approval, has enough evidence been gathered? Many of these don't need a full general-purpose reasoning model every time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev + MCP.&lt;/strong&gt; One obvious next step is dynamic MCP tool discovery — instead of manually registering capabilities, the system pulls them from an MCP server's &lt;code&gt;list_tools()&lt;/code&gt; call, and Jev's decision space becomes dynamic:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MCP SERVER → list_tools() → Tool names + descriptions → JEV → Select best tool → call_tool()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Jev + multi-agent systems.&lt;/strong&gt; The same idea works one level up — instead of selecting a tool, Jev selects a specialist agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ORCHESTRATOR
                      │
                      ▼
                     JEV
                      │
             Which specialist?
                      │
        ┌─────────────┼─────────────┐
        ▼             ▼             ▼
     Finance       Supply        Engineering
      Agent        Chain Agent      Agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The choices are known. The intelligence is in picking the one that fits the current context.&lt;/p&gt;

&lt;h2&gt;
  
  
  So When Should I Use Jev?
&lt;/h2&gt;

&lt;p&gt;Ask: &lt;strong&gt;can I clearly define the possible answers before making the model call?&lt;/strong&gt; If yes, this is probably a good candidate for bounded semantic judgment — approve/review/reject, billing/shipping/technical, retry/stop/escalate, low/medium/high risk.&lt;/p&gt;

&lt;p&gt;If you can't define the answer space because the problem requires exploration, explanation, synthesis, or creation, a general-purpose LLM is the better fit something like "analyze everything we know about this outage and identify the three most plausible root causes" is genuinely open ended.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Lesson
&lt;/h2&gt;

&lt;p&gt;For years, AI architecture discussions started with &lt;em&gt;which LLM should we use?&lt;/em&gt; Agentic systems push us toward a better question: &lt;strong&gt;what kind of intelligence does this particular step require?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sometimes the answer is code. Sometimes Jev. Sometimes OpenAI. Sometimes no model at all just call the tool. That produces a more modular architecture:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gvtvc0j5ggaush0mivi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5gvtvc0j5ggaush0mivi.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     AGENT HARNESS

          ┌──────────────┼──────────────┐
          │              │              │
          ▼              ▼              ▼
        CODE            JEV           OPENAI
       Control         Decide          Reason
          │              │              │
          └──────────────┼──────────────┘
                         │
                         ▼
                       TOOLS
                         │
                         ▼
                        ACT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I don't think the future of AI agents means finding &lt;strong&gt;one giant model and letting it control every decision. A more interesting&lt;/strong&gt; architecture uses the right intelligence for the right decision: code when the rule is deterministic, Jev when the choices are known but semantic judgment is required, OpenAI (or another LLM) when genuine reasoning or generation is needed, and tools when something actually has to happen in the real world.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code controls. Jev decides. OpenAI reasons. Tools act. The agent harness orchestrates everything.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the hybrid agent architecture I wanted to explore.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent Output
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fooedyymqzo9tie6qudbi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fooedyymqzo9tie6qudbi.png" alt=" " width="800" height="604"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
    <item>
      <title>Stop Sending Every Decision to an LLM: Code vs. Jev vs. Claude</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Mon, 21 Sep 2026 06:03:19 +0000</pubDate>
      <link>https://dev.to/sreeni5018/stop-sending-every-decision-to-an-llm-code-vs-jev-vs-claude-32e4</link>
      <guid>https://dev.to/sreeni5018/stop-sending-every-decision-to-an-llm-code-vs-jev-vs-claude-32e4</guid>
      <description>&lt;p&gt;&lt;strong&gt;Why bounded semantic decisions may become an important layer in agentic AI architecture.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  𝗝𝗲𝘃, 𝗔𝗽𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗦𝘁𝗮𝘁𝗲, and Why I Think of It as 𝗛𝗼𝗺𝗲 𝗗𝗲𝗹𝗶𝘃𝗲𝗿𝘆 vs. a 𝗖𝗹𝗮𝘂𝗱𝗲 𝗕𝘂𝗳𝗳𝗲𝘁
&lt;/h2&gt;

&lt;p&gt;For the last few years, our default AI architecture has been surprisingly simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Have a problem? Send it to an LLM.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Need to classify something? LLM.&lt;/p&gt;

&lt;p&gt;Need to choose a tool? LLM.&lt;/p&gt;

&lt;p&gt;Need to decide whether an agent should retry? LLM.&lt;/p&gt;

&lt;p&gt;Need to determine whether a document contains enough evidence? LLM.&lt;/p&gt;

&lt;p&gt;Need to choose one option from three possibilities? Again, LLM.&lt;/p&gt;

&lt;p&gt;General-purpose models such as Claude are extremely capable, but while experimenting with &lt;strong&gt;TypeSafe AI's 𝗝𝗲𝘃&lt;/strong&gt;, I started asking a different architectural question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why hire a chef when all I need is someone to pick the right item from an already-defined menu?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is where my &lt;strong&gt;𝗛𝗼𝗺𝗲 𝗗𝗲𝗹𝗶𝘃𝗲𝗿𝘆 vs. 𝗕𝘂𝗳𝗳𝗲𝘁&lt;/strong&gt; analogy came from.&lt;/p&gt;

&lt;h2&gt;
  
  
  But before getting to the analogy, it is important to understand what Jev actually is.
&lt;/h2&gt;

&lt;h1&gt;
  
  
  What Is 𝗝𝗲𝘃?
&lt;/h1&gt;

&lt;p&gt;Jev is TypeSafe AI's first public &lt;strong&gt;𝗦𝘆𝘀𝘁𝗲𝗺 𝗢𝗻𝗲 𝗠𝗼𝗱𝗲𝗹&lt;/strong&gt;. Unlike a general-purpose LLM whose primary interface is generated text, Jev is designed for &lt;strong&gt;decisions inside software&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9piqfwayetbib6kdmypg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9piqfwayetbib6kdmypg.png" alt=" " width="800" height="1000"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The basic idea is remarkably simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗔𝗽𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗦𝘁𝗮𝘁𝗲 → 𝗧𝘆𝗽𝗲𝗱 𝗤𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀 → 𝗣𝗿𝗼𝗯𝗮𝗯𝗶𝗹𝗶𝘀𝘁𝗶𝗰 𝗗𝗲𝗰𝗶𝘀𝗶𝗼𝗻𝘀&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;TypeSafe describes the model as taking program state and returning typed decisions that software can consume directly. The possible outputs are defined ahead of time instead of allowing the model to generate an arbitrary string.&lt;/p&gt;

&lt;p&gt;That difference is important.&lt;/p&gt;

&lt;p&gt;With Claude, I might ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Analyze this candidate's experience and explain whether this person would be appropriate for an enterprise AI architecture role.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The output could be several paragraphs.&lt;/p&gt;

&lt;p&gt;With Jev, I can instead define a bounded question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which role best matches this state?

A. Traditional Backend Developer
B. Generative and Agentic AI Architect
C. Database Administrator
D. Manual Test Engineer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now I am not asking the model to generate language.&lt;/p&gt;

&lt;p&gt;I am asking it to &lt;strong&gt;make a decision within a vocabulary that my application already controls&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is why, as my own mental model, I think of Jev as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“𝗝𝘂𝘀𝘁 𝗘𝗻𝗿𝗶𝗰𝗵𝗲𝗱 𝗩𝗼𝗰𝗮𝗯𝘂𝗹𝗮𝗿𝘆” — not the official expansion of Jev, but a useful way for me to remember what the architecture is doing.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The vocabulary belongs to my application.&lt;/p&gt;

&lt;p&gt;Jev adds &lt;strong&gt;semantic judgment&lt;/strong&gt; around it.&lt;/p&gt;




&lt;h1&gt;
  
  
  𝗝𝗲𝘃 Starts With the 𝗦𝘁𝗮𝘁𝗲 of the Application
&lt;/h1&gt;

&lt;p&gt;This is the part I find most interesting.&lt;/p&gt;

&lt;p&gt;Instead of beginning with a chat prompt such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“You are an expert AI architect. Analyze the following information…”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I start by defining the &lt;strong&gt;𝗦𝘁𝗮𝘁𝗲 that exists right now&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Sreeni Ramadurai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"myboss"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Krishna K"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"skills"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GenAI, AgenticAI, AWS, Azure"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not really a prompt in the traditional chatbot sense.&lt;/p&gt;

&lt;p&gt;It is a &lt;strong&gt;snapshot of what the application currently knows&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The state might contain a customer message, an account status, a transaction, retrieved documents, an agent trace, tool results, user permissions, workflow history, or any other information required to make the next decision.&lt;/p&gt;

&lt;p&gt;TypeSafe's own description emphasizes this distinction: Jev inputs are oriented around &lt;strong&gt;structured program state&lt;/strong&gt;, while conventional LLM interfaces tend to be organized around sequential conversational messages.&lt;/p&gt;

&lt;p&gt;That gives me a useful mental model:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;State is not what I want the model to say. State is what the system currently knows.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Then I separately define what I want the system to decide.
&lt;/h2&gt;

&lt;h1&gt;
  
  
  𝗦𝘁𝗮𝘁𝗲 First, 𝗤𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀 Second
&lt;/h1&gt;

&lt;p&gt;Using the small state above, I experimented with questions such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Is there sufficient evidence of architecture readiness?

Which enterprise AI role best matches the available evidence?

How strong is the evidence?

Does this state support a production deployment claim?

Does this state establish that Krishna K is Sreeni's manager?

Which conclusion would require information that is not present?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the architecture.&lt;/p&gt;

&lt;p&gt;I am not putting everything into one large prompt and asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Think about all of this and tell me what you think.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead, the application provides &lt;strong&gt;one state and several independent decision questions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is very close to normal software engineering.&lt;/p&gt;

&lt;p&gt;We define the data.&lt;/p&gt;

&lt;p&gt;We define the contract.&lt;/p&gt;

&lt;p&gt;We define the permitted outputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then intelligence evaluates the state against those contracts.
&lt;/h2&gt;

&lt;h1&gt;
  
  
  Three Kinds of Decisions
&lt;/h1&gt;

&lt;p&gt;In the current Jev interface, these bounded judgments are expressed through typed question forms such as &lt;strong&gt;𝗡𝗼𝘂𝗹, 𝗖𝗵𝗼𝗶𝗰𝗲, and 𝗦𝗰𝗼𝗿𝗲&lt;/strong&gt;. TypeSafe's public materials describe Jev broadly as returning typed decisions with probabilities and confidence values rather than open-ended generated text.&lt;/p&gt;

&lt;p&gt;For example, a binary judgment might ask whether the state contains enough evidence for a particular claim.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;𝗖𝗵𝗼𝗶𝗰𝗲&lt;/strong&gt; might ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Traditional Backend Developer
Generative and Agentic AI Architect
Database Administrator
Manual Test Engineer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;strong&gt;𝗦𝗰𝗼𝗿𝗲&lt;/strong&gt; might ask how strongly the available evidence supports enterprise AI capability on a predefined scale.&lt;/p&gt;

&lt;p&gt;The important point is not the names of these primitives.&lt;/p&gt;

&lt;p&gt;The important point is that &lt;strong&gt;the shape of the answer exists before inference begins&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The application owns the decision boundary.&lt;/p&gt;

&lt;p&gt;The model judges what belongs inside it.&lt;/p&gt;

&lt;h1&gt;
  
  
  Jev Playground with my own State
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fns8q9j46l2aifx77gzyz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fns8q9j46l2aifx77gzyz.png" alt=" " width="800" height="432"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  Now the Restaurant Analogy Becomes Clear
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjnej7hjdrm9ck5knhktn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjnej7hjdrm9ck5knhktn.png" alt=" " width="800" height="1000"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Imagine opening a food-delivery application.&lt;/p&gt;

&lt;p&gt;The restaurant has only three dishes available:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pizza&lt;br&gt;
Burger&lt;br&gt;
Biryani&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Then I provide some state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Vegetarian
Likes spicy food
Wants something filling
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The menu has already been established.&lt;/p&gt;

&lt;p&gt;Jev does not need to invent another meal.&lt;/p&gt;

&lt;p&gt;It does not need to write a recipe.&lt;/p&gt;

&lt;p&gt;It does not need to explain the history of biryani.&lt;/p&gt;

&lt;p&gt;Its job is simply to evaluate the state against the known choices.&lt;/p&gt;

&lt;p&gt;The result might look conceptually like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Biryani    82%
Pizza      14%
Burger      4%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is why I call Jev &lt;strong&gt;𝗶𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝘁 𝗵𝗼𝗺𝗲 𝗱𝗲𝗹𝗶𝘃𝗲𝗿𝘆&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The intelligence is real, because choosing correctly may require semantic understanding.&lt;/p&gt;

&lt;p&gt;But the destination is bounded.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;𝗛𝗼𝗺𝗲 𝗗𝗲𝗹𝗶𝘃𝗲𝗿𝘆 = 𝗸𝗻𝗼𝘄𝗻 𝗺𝗲𝗻𝘂 + 𝗰𝗼𝗻𝘁𝗲𝘅𝘁 + 𝗷𝘂𝗱𝗴𝗺𝗲𝗻𝘁.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h1&gt;
  
  
  𝗖𝗹𝗮𝘂𝗱𝗲 Is the 𝗕𝘂𝗳𝗳𝗲𝘁
&lt;/h1&gt;

&lt;p&gt;Now imagine walking into a buffet.&lt;/p&gt;

&lt;p&gt;I can ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What should I eat?&lt;/p&gt;

&lt;p&gt;Build me a vegetarian dinner.&lt;/p&gt;

&lt;p&gt;Compare these dishes.&lt;/p&gt;

&lt;p&gt;Explain why one is healthier.&lt;/p&gt;

&lt;p&gt;Suggest something I haven't considered.&lt;/p&gt;

&lt;p&gt;Create a completely different meal.&lt;/p&gt;

&lt;p&gt;Plan my meals for the entire week.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Now I am no longer choosing from one narrowly defined application vocabulary.&lt;/p&gt;

&lt;p&gt;I want exploration.&lt;/p&gt;

&lt;p&gt;I want synthesis.&lt;/p&gt;

&lt;p&gt;I want explanation.&lt;/p&gt;

&lt;p&gt;I may not even know the answer space before I start asking questions.&lt;/p&gt;

&lt;p&gt;That is where a general-purpose model such as Claude becomes valuable.&lt;/p&gt;

&lt;p&gt;So in my architecture analogy:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;𝗝𝗲𝘃 is 𝗛𝗼𝗺𝗲 𝗗𝗲𝗹𝗶𝘃𝗲𝗿𝘆. 𝗖𝗹𝗮𝘂𝗱𝗲 is the 𝗕𝘂𝗳𝗳𝗲𝘁.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Home delivery does not mean unintelligent.&lt;/p&gt;

&lt;p&gt;And buffet does not mean better.&lt;/p&gt;

&lt;p&gt;They solve different problems.&lt;/p&gt;




&lt;h1&gt;
  
  
  But Before Jev, There Is Still 𝗖𝗼𝗱𝗲
&lt;/h1&gt;

&lt;p&gt;There is an even simpler layer.&lt;/p&gt;

&lt;p&gt;Suppose my application already has this business rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IF vegetarian
    remove all meat dishes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I do not need Jev.&lt;/p&gt;

&lt;p&gt;And I certainly do not need Claude.&lt;/p&gt;

&lt;p&gt;The business has already determined exactly what should happen.&lt;/p&gt;

&lt;p&gt;Use code.&lt;/p&gt;

&lt;p&gt;That gives me three architectural layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CODE
Explicit rule is already known
        ↓
JEV
Options are known,
but choosing requires semantic judgment
        ↓
CLAUDE
Answer space is open,
requiring reasoning, synthesis or generation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or, using my restaurant analogy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CODE
"The restaurant rule already decides."

        ↓

JEV — HOME DELIVERY
"The menu is fixed.
Choose the best item for this context."

        ↓

CLAUDE — BUFFET
"Explore, compare, combine,
explain and create."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That distinction is much more useful to me than asking which model is “smarter.”&lt;/p&gt;




&lt;h1&gt;
  
  
  Where Jev Gets Really Interesting: 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗔𝗜
&lt;/h1&gt;

&lt;p&gt;Now move away from restaurants and consider an enterprise agent.&lt;/p&gt;

&lt;p&gt;During one workflow, an agent may need to decide:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which tool should I call?

Is this retrieved document relevant?

Did the tool actually satisfy the task?

Should I retry?

Should I continue?

Should I escalate to a human?

Is this transaction low, medium or high risk?

Does the evidence support this claim?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Many of these are &lt;strong&gt;semantic decisions&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;They cannot always be expressed as simple &lt;code&gt;if/else&lt;/code&gt; statements.&lt;/p&gt;

&lt;p&gt;But they also do not necessarily need paragraphs of generated reasoning.&lt;/p&gt;

&lt;p&gt;Their answer spaces often look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Tool
Salesforce | ServiceNow | SharePoint | None

Action
Continue | Retry | Escalate

Risk
Low | Medium | High

Evidence
Supported | Partial | Unsupported
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where I see Jev fitting naturally inside an &lt;strong&gt;𝗔𝗴𝗲𝗻𝘁 𝗛𝗮𝗿𝗻𝗲𝘀𝘀&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     AGENT HARNESS
                          │
           ┌──────────────┼──────────────┐
           │              │              │
         CODE            JEV          CLAUDE
           │              │              │
         Rules         Semantic       Reasoning
                       Judgment       Generation
           │              │              │
           └──────────────┼──────────────┘
                          │
                        TOOLS
                          │
                        ACTION
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Code controls deterministic behavior.&lt;/p&gt;

&lt;p&gt;Jev makes &lt;strong&gt;bounded semantic judgments&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Claude handles the parts where broader reasoning or generation is actually necessary.&lt;/p&gt;

&lt;p&gt;That architecture interests me much more than simply routing everything through one giant model.&lt;/p&gt;




&lt;h1&gt;
  
  
  Is Jev's “𝗔𝗽𝗽𝗹𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗦𝘁𝗮𝘁𝗲” Related to 𝗛𝗔𝗧𝗘𝗢𝗔𝗦?
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F76hxy06wyfe04ekxjxyv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F76hxy06wyfe04ekxjxyv.png" alt=" " width="800" height="1000"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When I started thinking about Jev in terms of &lt;strong&gt;application state&lt;/strong&gt;, another architecture immediately came to mind:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗛𝗔𝗧𝗘𝗢𝗔𝗦 — Hypermedia as the Engine of Application State.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is definitely a conceptual connection, but they are &lt;strong&gt;not the same thing&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In REST, HATEOAS means the server representation provides hypermedia controls that tell the client what resources or transitions are available next. Rather than hard-coding every possible navigation path, the client can discover permitted next actions from the representation it receives. This is part of the REST architectural style described by Roy Fielding.&lt;/p&gt;

&lt;p&gt;For example, an order might be represented as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"orderId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1001&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pending"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"links"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"rel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"cancel"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"href"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/orders/1001/cancel"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"rel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pay"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"href"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/orders/1001/payment"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The current state determines which transitions are available.&lt;/p&gt;

&lt;p&gt;That sounds somewhat similar to what we are doing with Jev—but there is an important difference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗛𝗔𝗧𝗘𝗢𝗔𝗦 exposes the allowed next actions.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗝𝗲𝘃 can semantically judge which allowed action best fits the current context.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That leads to an architecture I find fascinating.&lt;/p&gt;

&lt;p&gt;Imagine the application state says:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"orderStatus"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pending"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"customerMessage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"I accidentally placed this order twice."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"paymentCaptured"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"availableActions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"cancel"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"continue"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"escalate"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;HATEOAS could tell the application:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;These are the actions currently available.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev could then answer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Given the complete state,
which available action best fits the situation?

cancel      94%
escalate     5%
continue     1%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then deterministic code executes the selected transition subject to whatever confidence threshold, permissions, policy checks, and guardrails the application requires.&lt;/p&gt;

&lt;p&gt;So I would summarize the relationship this way:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;𝗛𝗔𝗧𝗘𝗢𝗔𝗦 tells the client what it may do next.&lt;br&gt;
𝗝𝗲𝘃 can help the application judge which permitted action makes sense next.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is not part of the formal definition of HATEOAS, nor is Jev an implementation of HATEOAS.&lt;/p&gt;

&lt;p&gt;It is simply a useful architectural connection.&lt;/p&gt;

&lt;p&gt;HATEOAS gives us &lt;strong&gt;state-driven discoverability&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Jev gives us &lt;strong&gt;state-driven semantic judgment&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Put together, they suggest an interesting pattern for intelligent APIs and agents.&lt;/p&gt;




&lt;h1&gt;
  
  
  𝗦𝘁𝗮𝘁𝗲 → 𝗣𝗼𝘀𝘀𝗶𝗯𝗶𝗹𝗶𝘁𝗶𝗲𝘀 → 𝗝𝘂𝗱𝗴𝗺𝗲𝗻𝘁 → 𝗔𝗰𝘁𝗶𝗼𝗻
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8bfjc3dtdologo2w3yqm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8bfjc3dtdologo2w3yqm.png" alt=" " width="800" height="1000"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This may actually be the most useful way for me to think about Jev.&lt;/p&gt;

&lt;p&gt;Not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prompt → LLM → Text
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CURRENT STATE
      ↓
AVAILABLE DECISIONS
      ↓
SEMANTIC JUDGMENT
      ↓
PROBABILITY / CONFIDENCE
      ↓
BUSINESS RULES + GUARDRAILS
      ↓
ACTION
      ↓
NEW STATE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And then the cycle repeats.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;State₀
   ↓
Decision
   ↓
Action
   ↓
State₁
   ↓
Decision
   ↓
Action
   ↓
State₂
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now Jev starts looking less like a chatbot and more like an &lt;strong&gt;intelligence primitive inside a stateful software system&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is the part I find most compelling.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why Not Just Give Claude Structured Output?
&lt;/h1&gt;

&lt;p&gt;This is the obvious question.&lt;/p&gt;

&lt;p&gt;Claude can absolutely be given:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Choose exactly one:

cancel
continue
escalate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It can return JSON.&lt;/p&gt;

&lt;p&gt;It can use tool calling.&lt;/p&gt;

&lt;p&gt;It can conform to a schema.&lt;/p&gt;

&lt;p&gt;So the question is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Can Claude make bounded decisions?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Of course it can.&lt;/p&gt;

&lt;p&gt;The architecture question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If my application needs thousands of small semantic judgments whose output spaces are already known, should every one of them go through a general-purpose text-generation architecture?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Jev was explicitly designed around this narrower machine-facing task. TypeSafe describes it as optimizing for &lt;strong&gt;structured, calibrated decisions&lt;/strong&gt; rather than free-form strings.&lt;/p&gt;

&lt;p&gt;This becomes especially interesting inside agentic systems where a single user request may cause many internal decisions.&lt;/p&gt;

&lt;p&gt;One model call may be insignificant.&lt;/p&gt;

&lt;p&gt;Thousands of decisions across many agents, tools, users, and workflow steps are a different architectural problem.&lt;/p&gt;




&lt;h1&gt;
  
  
  𝗕𝗼𝘂𝗻𝗱𝗲𝗱 Does Not Mean 𝗖𝗼𝗿𝗿𝗲𝗰𝘁
&lt;/h1&gt;

&lt;p&gt;There is one distinction I think is essential.&lt;/p&gt;

&lt;p&gt;Suppose Jev produces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Biryani 82%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That does not make Biryani objectively correct.&lt;/p&gt;

&lt;p&gt;Likewise:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"risk"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;being perfectly valid structured output does not prove that the risk assessment itself is correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗧𝘆𝗽𝗲 𝘀𝗮𝗳𝗲𝘁𝘆 ≠ 𝗦𝗲𝗺𝗮𝗻𝘁𝗶𝗰 𝗰𝗼𝗿𝗿𝗲𝗰𝘁𝗻𝗲𝘀𝘀.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So we still need evaluations, thresholds, observability, guardrails, and human review where the consequence demands it.&lt;/p&gt;

&lt;p&gt;The difference is that the &lt;strong&gt;decision surface is constrained and machine-consumable&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That can make the surrounding system easier to reason about.&lt;/p&gt;




&lt;h1&gt;
  
  
  How I Would Use Jev in an Enterprise Application
&lt;/h1&gt;

&lt;p&gt;My pattern would be simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Build the application state
        ↓
2. Define the decisions the application actually needs
        ↓
3. Express those decisions as bounded typed questions
        ↓
4. Let Jev evaluate the state
        ↓
5. Read probabilities/confidence
        ↓
6. Apply deterministic thresholds and policies in code
        ↓
7. Execute a tool, ask Claude for deeper reasoning,
   or escalate to a human
        ↓
8. Update the application state
        ↓
9. Repeat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model should not own the entire workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗧𝗵𝗲 𝗵𝗮𝗿𝗻𝗲𝘀𝘀 𝗼𝘄𝗻𝘀 𝘁𝗵𝗲 𝘄𝗼𝗿𝗸𝗳𝗹𝗼𝘄.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model contributes intelligence at specific decision points.&lt;/p&gt;

&lt;p&gt;That separation matters.&lt;/p&gt;




&lt;h1&gt;
  
  
  My Mental Model After Using Jev
&lt;/h1&gt;

&lt;p&gt;I now think about AI architecture using three increasingly flexible layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗖𝗢𝗗𝗘 = 𝗥𝗨𝗟𝗘𝗦&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The correct behavior is already explicitly known.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If payment has already been captured, do not cancel automatically.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;𝗝𝗘𝗩 = 𝗝𝗨𝗗𝗚𝗠𝗘𝗡𝗧 / 𝗛𝗢𝗠𝗘 𝗗𝗘𝗟𝗜𝗩𝗘𝗥𝗬&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The options are known, but the correct choice depends on interpreting context.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Cancel, Continue, or Escalate?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;𝗖𝗟𝗔𝗨𝗗𝗘 = 𝗥𝗘𝗔𝗦𝗢𝗡𝗜𝗡𝗚 + 𝗚𝗘𝗡𝗘𝗥𝗔𝗧𝗜𝗢𝗡 / 𝗕𝗨𝗙𝗙𝗘𝗧&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The problem requires exploration, explanation, planning, synthesis, or creation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Analyze the entire situation, explain the trade-offs, and propose a resolution strategy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is not a competition between models.&lt;/p&gt;

&lt;p&gt;It is &lt;strong&gt;separation of concerns&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  The Bigger Question: Are We Overusing Generation?
&lt;/h1&gt;

&lt;p&gt;Software architecture has always been about choosing the appropriate abstraction.&lt;/p&gt;

&lt;p&gt;We do not use a database where a queue belongs.&lt;/p&gt;

&lt;p&gt;We do not use Kubernetes to execute one shell script.&lt;/p&gt;

&lt;p&gt;And perhaps we should not use &lt;strong&gt;open-ended generation&lt;/strong&gt; where all we really need is a &lt;strong&gt;bounded semantic decision&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is the larger lesson I took from experimenting with Jev.&lt;/p&gt;

&lt;p&gt;The future of AI applications may not be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application → One giant LLM → Everything
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It may look more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application State
        │
        ▼
      Harness
        │
   ┌────┼────────────┐
   │    │            │
 Code  Jev         Claude
   │    │            │
Rules Judgment    Reasoning
                  Generation
   └────┼────────────┘
        │
       Tools
        │
      Actions
        │
     New State
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And once I saw it this way, my restaurant analogy finally clicked.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗖𝗼𝗱𝗲 already knows the restaurant rules.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗝𝗲𝘃 is intelligent 𝗛𝗼𝗺𝗲 𝗗𝗲𝗹𝗶𝘃𝗲𝗿𝘆 from a defined menu.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;𝗖𝗹𝗮𝘂𝗱𝗲 gives me the 𝗕𝘂𝗳𝗳𝗲𝘁 when I genuinely need exploration and creation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So before making the next expensive generative call, perhaps the architecture question should not be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Which model is smartest?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Does this decision need a 𝗿𝘂𝗹𝗲, intelligent 𝗵𝗼𝗺𝗲 𝗱𝗲𝗹𝗶𝘃𝗲𝗿𝘆, or the entire 𝗯𝘂𝗳𝗳𝗲𝘁?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sometimes we need the chef.&lt;/p&gt;

&lt;p&gt;But sometimes all we needed was someone intelligent enough to pick the right item from the menu.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sreeni Ramadorai&lt;/strong&gt;&lt;br&gt;
&lt;em&gt;AI Architect&lt;/em&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
    </item>
    <item>
      <title>Why Hire a Chef Just to Pick From the Menu?</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Sun, 20 Sep 2026 18:29:31 +0000</pubDate>
      <link>https://dev.to/sreeni5018/why-hire-a-chef-just-to-pick-from-the-menu-4a0n</link>
      <guid>https://dev.to/sreeni5018/why-hire-a-chef-just-to-pick-from-the-menu-4a0n</guid>
      <description>&lt;h2&gt;
  
  
  Jev, Claude, and a Different Way to Think About AI in a Token-Based Economy
&lt;/h2&gt;

&lt;p&gt;For the last few years, whenever we had a problem that required some form of intelligence, our instinct was almost automatic: &lt;strong&gt;send the context to a large language model and ask it what to do.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That approach made sense. Models such as &lt;strong&gt;Claude&lt;/strong&gt; are remarkably capable because they can understand a &lt;strong&gt;situation&lt;/strong&gt;, &lt;strong&gt;reason about it&lt;/strong&gt;, &lt;strong&gt;explain their thinking&lt;/strong&gt;, &lt;strong&gt;generate content&lt;/strong&gt;, &lt;strong&gt;write code&lt;/strong&gt;, &lt;strong&gt;use tools&lt;/strong&gt;, and work through multi step problems. If the problem is open ended, that flexibility is exactly what we want.&lt;/p&gt;

&lt;p&gt;But recently I started thinking about a different class of problems inside enterprise software.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A customer request needs to be routed to Billing, Shipping, or Technical Support.&lt;/strong&gt; A transaction needs to be labeled Low, Medium, or High Risk. A workflow needs to Continue, Retry, Escalate, or Stop. A request needs to be Approved, Reviewed, or Rejected.&lt;/p&gt;

&lt;p&gt;Notice what is different about these problems.&lt;/p&gt;

&lt;p&gt;The choices already exist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The system is not asking AI to invent an answer.&lt;/strong&gt; It is asking AI to understand some messy, &lt;strong&gt;unstructured context and choose the most appropriate answer from a known set of possibilities.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is what made &lt;strong&gt;TypeSafe AI's Jev&lt;/strong&gt; interesting to me.&lt;/p&gt;

&lt;p&gt;It also gave me a simple analogy that helped me understand the difference between a decision-oriented model such as Jev and a general-purpose model such as Claude:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Are we sometimes hiring a chef when all we really need is someone intelligent enough to pick the right item from the menu?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjd8ri6yoe4crr817axmw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjd8ri6yoe4crr817axmw.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Jev Starts With the Menu
&lt;/h2&gt;

&lt;p&gt;TypeSafe describes Jev as its first &lt;strong&gt;System One Model&lt;/strong&gt;, designed around a simple idea: take unstructured state as input and produce typed, probabilistic decisions as output.&lt;/p&gt;

&lt;p&gt;That is a very different starting point from a chatbot.&lt;/p&gt;

&lt;p&gt;Imagine a support application has a customer message, account history, product information, previous interactions, and perhaps some transaction context. The application already knows which questions it needs answered: which department should handle the request, how urgent it is, and whether it should be escalated.&lt;/p&gt;

&lt;p&gt;The possible answers are already defined.&lt;/p&gt;

&lt;p&gt;Jev's job is to make the semantic judgment.&lt;/p&gt;

&lt;p&gt;Conceptually, I think of it like this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Application context → predefined decision space → semantic judgment → typed probabilities → software action&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the important part. Jev is not primarily designed to compose a beautiful sentence explaining that a ticket probably belongs to Technical Support. If the software ultimately needs &lt;code&gt;Technical Support&lt;/code&gt; with a confidence score, that structured decision is the product.&lt;/p&gt;

&lt;p&gt;TypeSafe says Jev is built around typed outputs, a parallel sampling approach rather than ordinary autoregressive text generation, and a training approach called Reinforcement Learning for Calibrated Decisions. The goal is not simply to pick an answer, but to attach useful uncertainty to the decision.&lt;/p&gt;

&lt;p&gt;That last part matters a lot in enterprise systems.&lt;/p&gt;

&lt;p&gt;If a decision is made with high confidence, perhaps the application can continue automatically. If confidence is lower, another system can verify it. If confidence is very low, the request can be routed to a human.&lt;/p&gt;

&lt;p&gt;I would still make one important distinction. A typed output does not mean the semantic judgment is guaranteed to be correct. Jev can still choose the wrong category. What the architecture gives us is a more constrained output space and a decision that is naturally easier for software to consume.&lt;/p&gt;

&lt;p&gt;In other words, &lt;strong&gt;type safety is not the same thing as semantic correctness&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Starts With the Problem
&lt;/h2&gt;

&lt;p&gt;Claude begins from a much broader place.&lt;/p&gt;

&lt;p&gt;Instead of saying, “Here are the choices; tell me which one fits,” we can give Claude a situation and ask it to figure out what should happen.&lt;/p&gt;

&lt;p&gt;Using the same support example, I could simply ask Claude which department should receive the ticket, and it could answer “Technical Support.”&lt;/p&gt;

&lt;p&gt;But I could also ask it to explain why, summarize the customer problem, propose troubleshooting steps, draft the customer response, determine whether the issue needs a Jira ticket, call an appropriate tool, inspect the result, and decide what to do next.&lt;/p&gt;

&lt;p&gt;Now we are no longer talking about classification.&lt;/p&gt;

&lt;p&gt;We are talking about reasoning, planning, generation, and possibly tool use.&lt;/p&gt;

&lt;p&gt;That is Claude's strength.&lt;/p&gt;

&lt;p&gt;So I do not see this as Jev being “better” than Claude or Claude being “too expensive.” They are designed around different kinds of work.&lt;/p&gt;

&lt;p&gt;The distinction I find useful is this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;th&gt;Claude&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Question&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Which predefined option fits?&lt;/td&gt;
&lt;td&gt;Given this situation, what should happen?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Answer space&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Bounded&lt;/td&gt;
&lt;td&gt;Open&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Primary role&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Semantic decision&lt;/td&gt;
&lt;td&gt;Reasoning and generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Output&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Typed decisions and probabilities&lt;/td&gt;
&lt;td&gt;Language, code, structured output, tool calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Best fit&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Repeated bounded judgments&lt;/td&gt;
&lt;td&gt;Complex, exploratory or generative work&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And this is where the restaurant analogy becomes useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Home Delivery vs. Buffet
&lt;/h2&gt;

&lt;p&gt;Imagine I open a food-delivery application and the restaurant has three choices: Pizza, Burger, and Biryani.&lt;/p&gt;

&lt;p&gt;I tell the system, “I want something spicy, vegetarian, and filling.”&lt;/p&gt;

&lt;p&gt;The menu has already been created. The system does not need to invent another dish. It does not need to explain where biryani came from or write a recipe. It only needs enough intelligence to understand my preference and decide which existing choice is the best match.&lt;/p&gt;

&lt;p&gt;Maybe the result is something like:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Biryani 82% — Pizza 14% — Burger 4%&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is how I think about Jev.&lt;/p&gt;

&lt;p&gt;The intelligence is in selecting correctly from a bounded menu.&lt;/p&gt;

&lt;p&gt;Now imagine that instead of ordering delivery, I walk into a buffet.&lt;/p&gt;

&lt;p&gt;I might ask, “What should I eat?” But I could just as easily ask, “Build me a healthy plate,” “Compare these dishes,” “Suggest something I have not considered,” or “Take these ingredients and create a new meal.”&lt;/p&gt;

&lt;p&gt;Now the answer space is much larger.&lt;/p&gt;

&lt;p&gt;That is how I think about Claude.&lt;/p&gt;

&lt;p&gt;Claude can certainly choose biryani. But it is capable of doing much more than choosing biryani. It can reason about the whole dining experience.&lt;/p&gt;

&lt;p&gt;So for me:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev is intelligent selection from the menu. Claude can reason beyond the menu.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  But Couldn't Claude Do the Same Thing?
&lt;/h2&gt;

&lt;p&gt;This was my next question.&lt;/p&gt;

&lt;p&gt;If Claude is already capable of understanding the request, why not simply give it a strict prompt, a schema, some few-shot examples, structured output constraints, or an Agent Skill and force it to return only Pizza, Burger, or Biryani?&lt;/p&gt;

&lt;p&gt;Of course we can.&lt;/p&gt;

&lt;p&gt;Claude can be constrained very effectively. Modern LLM applications already do this with system instructions, tool schemas, JSON outputs, guardrails, validators, and application logic.&lt;/p&gt;

&lt;p&gt;But that changes the question from “Can Claude do it?” to “Is a general-purpose reasoning and generation model the right architecture for every bounded decision?”&lt;/p&gt;

&lt;p&gt;We can hire a chef and tell the chef, “Do not cook anything. Just look at the three menu items and tell me which one I should order.”&lt;/p&gt;

&lt;p&gt;The chef can absolutely do that.&lt;/p&gt;

&lt;p&gt;The more interesting question is whether we needed the chef in the first place.&lt;/p&gt;

&lt;p&gt;This becomes especially relevant when the same decision is being made millions of times inside an enterprise application.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrt3oiqt8qa7fvuhgyb4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrt3oiqt8qa7fvuhgyb4.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Token Economy Changes the Architecture Conversation
&lt;/h2&gt;

&lt;p&gt;Generative AI has made us think in tokens. We send prompts, documents, conversation history, tool definitions, instructions, and other context into a model. The model processes those tokens and may produce reasoning, generated text, structured outputs, or tool calls.&lt;/p&gt;

&lt;p&gt;For many tasks, that cost is worthwhile because we genuinely need the reasoning and generation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcs02qii4hwntar1iua78.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcs02qii4hwntar1iua78.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But suppose a workflow is repeatedly asking only:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Approve, Review, or Reject?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Or:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Continue, Retry, Escalate, or Stop?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Or:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Billing, Shipping, or Technical?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the final answer is always one choice from a small, predefined set, it is worth asking whether every one of those decisions needs to travel through a fully general generative architecture.&lt;/p&gt;

&lt;p&gt;That is the idea I find more interesting than any individual benchmark or vendor claim.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why pay for generation when the application does not need generation?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This does not mean “replace LLMs.” It means we should become more deliberate about where we use them.&lt;/p&gt;

&lt;h2&gt;
  
  
  And Sometimes We Don't Need AI at All
&lt;/h2&gt;

&lt;p&gt;There is another boundary that is just as important.&lt;/p&gt;

&lt;p&gt;If the rule is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If the user is not authorized, deny access.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We do not need Jev.&lt;/p&gt;

&lt;p&gt;We do not need Claude.&lt;/p&gt;

&lt;p&gt;We need code.&lt;/p&gt;

&lt;p&gt;If the rule is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If retry count is greater than three, stop.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Again, use code.&lt;/p&gt;

&lt;p&gt;AI becomes interesting when the outcome is bounded but the decision cannot easily be expressed as a clean deterministic rule.&lt;/p&gt;

&lt;p&gt;For example, “Determine whether this supplier request represents Low, Medium, or High operational risk based on the description, previous activity, contract information, and current business context.”&lt;/p&gt;

&lt;p&gt;The three possible outputs are known.&lt;/p&gt;

&lt;p&gt;But choosing among them requires understanding meaning.&lt;/p&gt;

&lt;p&gt;That is where I now find this simple architecture model useful:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code = Rules&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev = Judgment&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude = Reasoning + Generation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I do not mean this as a rigid rule. I use it as a mental model.&lt;/p&gt;

&lt;p&gt;If the decision is deterministic, code should handle it. If the possible answers are known but choosing among them requires semantic judgment, a decision-oriented model becomes interesting. If the problem requires exploration, explanation, planning, synthesis, creation, or dynamic tool use, a general-purpose reasoning model belongs there.&lt;/p&gt;

&lt;p&gt;That is more useful to me than thinking only in terms of “small model versus big model.”&lt;/p&gt;

&lt;h2&gt;
  
  
  This Could Become Even More Important in Agentic AI
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1a509ft06xw6c2d9jwr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc1a509ft06xw6c2d9jwr.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Agentic AI makes this distinction even more interesting.&lt;/p&gt;

&lt;p&gt;Think about the questions an agent asks itself during execution:&lt;/p&gt;

&lt;p&gt;Which tool should I use? Should I retry? Is this result relevant? Did the previous action succeed? Should I escalate? Is the request risky? Should I continue or stop?&lt;/p&gt;

&lt;p&gt;Many of those questions have a bounded answer space.&lt;/p&gt;

&lt;p&gt;They are runtime judgments, not necessarily open-ended generation problems.&lt;/p&gt;

&lt;p&gt;Today, we often send every one of these decisions back through the same large reasoning model because that is convenient. But I am not convinced that mature agent architectures will continue doing that forever.&lt;/p&gt;

&lt;p&gt;I can imagine a harness where different components do different jobs:&lt;/p&gt;

&lt;p&gt;Code handles predictable control flow. A decision-oriented model handles repeated semantic judgments. Claude handles the difficult reasoning, planning, synthesis, and generation. Tools perform the actual actions. The harness manages how these pieces work together.&lt;/p&gt;

&lt;p&gt;That looks less like “an LLM application” and more like mature software architecture.&lt;/p&gt;

&lt;p&gt;Different components are optimized for different responsibilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Question Is Not Jev vs. Claude
&lt;/h2&gt;

&lt;p&gt;I would not design a system by asking, “Should I use Jev or Claude?”&lt;/p&gt;

&lt;p&gt;I would start one level higher.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What kind of problem is this particular model call solving?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can normal code solve it reliably? Then use code.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Are the possible outcomes already known, but selecting the right one requires understanding unstructured context? That is where a decision-oriented model such as Jev becomes interesting.&lt;/p&gt;

&lt;p&gt;Does the problem require exploring possibilities, explaining something, writing, planning, synthesizing information, using tools dynamically, or generating something new? That is where a model such as Claude earns its place.&lt;/p&gt;

&lt;p&gt;The first phase of Generative AI was about discovering everything large language models could do.&lt;/p&gt;

&lt;p&gt;I think the next phase will be about learning when we &lt;strong&gt;do not need them to do everything&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;As AI moves deeper into enterprise software, model intelligence will still matter. But architecture will matter just as much. We will care about latency, cost, uncertainty, reliability, observability, and how easily AI decisions compose with ordinary software.&lt;/p&gt;

&lt;p&gt;The future may not be &lt;strong&gt;LLM everywhere&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It may be a combination of &lt;strong&gt;rules where rules are enough, judgment where choices are bounded, and reasoning where the problem is open&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Sometimes we really need the chef.&lt;/p&gt;

&lt;p&gt;But sometimes the menu already exists.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;We just need enough intelligence to choose the right meal.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
    </item>
    <item>
      <title>When AI Writes the Code, What Makes an Engineer Valuable?</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Fri, 11 Sep 2026 05:31:30 +0000</pubDate>
      <link>https://dev.to/sreeni5018/when-ai-writes-the-code-what-makes-an-engineer-valuable-1bp1</link>
      <guid>https://dev.to/sreeni5018/when-ai-writes-the-code-what-makes-an-engineer-valuable-1bp1</guid>
      <description>&lt;p&gt;&lt;strong&gt;Claude&lt;/strong&gt; and &lt;strong&gt;OpenAI&lt;/strong&gt; are bringing what feels like a real &lt;strong&gt;storm to IT&lt;/strong&gt;. I don’t expect that storm to &lt;strong&gt;hit every kind of work equally&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My view is&lt;/strong&gt; that work reducible to &lt;strong&gt;“given clear intent, produce correct code”&lt;/strong&gt; will face increasing automation. That includes tasks across experience levels. The question for us is how we grow beyond producing code to taking responsibility for the systems that code creates.&lt;/p&gt;

&lt;p&gt;Writing an implementation is one part of engineering. Understanding the problem, deciding what should happen, anticipating what could go wrong, and verifying the outcome are equally essential.&lt;/p&gt;

&lt;p&gt;With AI agents, those responsibilities become especially visible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A capable model still needs a dependable system around it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5avatmmznekuko4f3xur.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5avatmmznekuko4f3xur.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;An agent may understand a request, select a tool, and execute an action. But can we trust it to choose the right action, use the right information, respect the right boundaries, and recognize when it should stop?&lt;/p&gt;

&lt;p&gt;That is where &lt;strong&gt;harness&lt;/strong&gt; and &lt;strong&gt;loop&lt;/strong&gt; &lt;strong&gt;engineering&lt;/strong&gt; matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Harness engineering&lt;/strong&gt; shapes the environment the agent operates within: its instructions, tools, context, memory, permissions, approval gates, execution limits, and observability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Loop engineering&lt;/strong&gt; shapes how execution progresses: the model chooses an action, the system executes it, results return, and the agent decides whether to continue, recover, ask for help, or finish.&lt;/p&gt;

&lt;p&gt;The harness provides the controls and resources. The loop turns the goal into successive actions within those controls.&lt;/p&gt;

&lt;p&gt;Neither becomes dependable simply because the underlying model becomes more capable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A successful tool call is only part of the story&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Consider an agent asked to update an enterprise record through a tool exposed by an MCP server.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The tool executes and reports success. That tells us something about the operation—but it doesn’t establish that the agent selected the correct record, made the intended change, or acted with the required business approval.&lt;/p&gt;

&lt;p&gt;The tool could work exactly as designed while the overall task still fails.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“Tool-call success is not task success in the agentic AI world.”&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We need to verify the outcome against the original intent. Was the correct record updated? Were the values correct? Were permissions and approval requirements enforced? Can we trace the decision and recover if something went wrong?&lt;/p&gt;

&lt;p&gt;Authorization must be enforced by the system. It cannot depend on the model remembering to behave.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Engineering includes recognizing quiet failures&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some failures are obvious: a tool times out, a connection fails, or an exception appears.&lt;/p&gt;

&lt;p&gt;Others are harder to detect. An agent uses outdated context, selects a similarly named record, treats incomplete evidence as sufficient, or declares completion before checking the result.&lt;/p&gt;

&lt;p&gt;The workflow finishes. The answer sounds confident. The business outcome is still wrong.&lt;/p&gt;

&lt;p&gt;This is why evaluations, monitoring, and human escalation belong in the design from the beginning. We need evidence that the system performs acceptably, including when information is missing, requests are ambiguous, or tools fail.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Confidence is not correctness. Completion is not proof.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The frameworks will change. Judgment must keep growing.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Learning &lt;strong&gt;MCP&lt;/strong&gt;, &lt;strong&gt;A2A&lt;/strong&gt;, &lt;strong&gt;Agent Skills&lt;/strong&gt;, and &lt;strong&gt;evaluation&lt;/strong&gt; frameworks is useful. &lt;strong&gt;But no particular protocol or framework is a permanent career advantage&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Implementation patterns will evolve. Some work we perform manually today may become a standard platform feature tomorrow.&lt;/p&gt;

&lt;p&gt;The more durable skill is engineering judgment: &lt;strong&gt;understanding&lt;/strong&gt; the domain, &lt;strong&gt;identifying failure modes&lt;/strong&gt;, &lt;strong&gt;choosing appropriate controls&lt;/strong&gt;, and knowing what evidence is sufficient to trust an outcome.&lt;/p&gt;

&lt;p&gt;That judgment includes deciding when &lt;strong&gt;an agent is appropriate at all&lt;/strong&gt;. Some tasks are better served by a &lt;strong&gt;deterministic workflow&lt;/strong&gt;. Others need an agent with constrained autonomy. Some decisions require human review.&lt;/p&gt;

&lt;p&gt;Good engineering means making that choice deliberately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Business value remains the goal&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An agent taking more steps or using more tools is not automatically delivering more value.&lt;/p&gt;

&lt;p&gt;We should ask whether it improves quality, reduces turnaround time, controls cost, and earns users’ trust. We should also know who owns the outcome when it makes a mistake.&lt;/p&gt;

&lt;p&gt;Greater autonomy should come with clear accountability and verification proportionate to the consequences.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For me, this is the opportunity ahead&lt;/strong&gt;: combining &lt;strong&gt;software&lt;/strong&gt; fundamentals, &lt;strong&gt;domain knowledge&lt;/strong&gt;, and a reliability mindset with the ability to build and evaluate AI systems.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No skill guarantees that we will remain unaffected by change.&lt;/strong&gt; But learning to take responsibility for the whole system gives us a stronger foundation for adapting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don’t just learn to code with an agent. Learn to engineer the agent and keep re-engineering your understanding of what that requires.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;As agents take on more responsibility in our work, where should we strengthen our engineering practices first?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>NUMBERS: The Secret Language of AI Intelligence</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Thu, 16 Jul 2026 07:53:16 +0000</pubDate>
      <link>https://dev.to/sreeni5018/numbers-the-secret-language-of-ai-intelligence-5c76</link>
      <guid>https://dev.to/sreeni5018/numbers-the-secret-language-of-ai-intelligence-5c76</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Every time someone asks me to explain how a large language model actually works, I end up back at the same starting point: underneath the chat interface, everything a model does comes down to numbers moving through matrices. There's no understanding of language happening in the way people usually picture it. There's arithmetic, at a scale most people never see.&lt;/p&gt;

&lt;p&gt;To make that concrete, I picked seven core ideas that, together, explain most of what's going on inside a neural network, from a single neuron firing to a model turning its final calculation into an actual word. I arranged them so the first letters spell NUMBERS:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;N&lt;/strong&gt;eurons, the basic computing unit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;U&lt;/strong&gt;pdate, how the network learns from its mistakes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;M&lt;/strong&gt;emory, where what it learned gets stored&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;B&lt;/strong&gt;ias, a small but essential parameter (and a separate, more familiar meaning worth untangling)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;E&lt;/strong&gt;mbedding, how raw input becomes numbers in the first place&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;R&lt;/strong&gt;elationships, how the model figures out what's relevant to what&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;S&lt;/strong&gt;oftmax, the last step that turns a calculation into a choice&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Below, I'll take each letter in order, the same order data actually flows through a model: input becomes numbers, numbers get related to each other, and the result gets turned into a decision.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu00u4uiy0j6fm3c3q36m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu00u4uiy0j6fm3c3q36m.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  N — Neurons
&lt;/h2&gt;

&lt;p&gt;A neuron in a neural network is not the biological thing it's named after. It's a small unit that does one job: take in some numbers, multiply each by a &lt;strong&gt;weight&lt;/strong&gt;, add them up, add a &lt;strong&gt;bias&lt;/strong&gt;, and pass the result through an &lt;strong&gt;activation function&lt;/strong&gt;. That's it. No cell body, no dendrites doing anything mysterious.&lt;/p&gt;

&lt;p&gt;What makes neurons powerful isn't any single one of them. It's the fact that you stack thousands or millions of them into &lt;strong&gt;layers&lt;/strong&gt;, and each layer transforms the numbers it receives into a slightly more useful representation for the next layer. A neuron on its own is a simple function. A network of them is a very high-dimensional &lt;strong&gt;function approximator&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  U — Update
&lt;/h2&gt;

&lt;p&gt;Training a model is really just a loop: make a prediction, measure how wrong it was, and nudge the weights slightly in the direction that would have made the prediction less wrong. That nudge is the &lt;strong&gt;update&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The mechanism behind it is &lt;strong&gt;gradient descent&lt;/strong&gt;. You compute the &lt;strong&gt;gradient&lt;/strong&gt; of the &lt;strong&gt;loss function&lt;/strong&gt; with respect to every weight in the network, which tells you which direction increases the error and which direction decreases it. Then you take a small step in the direction that decreases it. Do this millions of times, in small &lt;strong&gt;batches&lt;/strong&gt;, and the weights slowly settle into values that make good predictions. Nothing about training a model is a single moment of learning. It's an enormous number of tiny corrections.&lt;/p&gt;

&lt;h2&gt;
  
  
  M — Memory
&lt;/h2&gt;

&lt;p&gt;People sometimes assume a model remembers things the way a person does, by storing episodes and recalling them later. That's not what's happening in most of the architecture. The &lt;strong&gt;weights&lt;/strong&gt; themselves are the memory. Every fact, pattern, and association the model has picked up during training is compressed into the numeric values of its &lt;strong&gt;parameters&lt;/strong&gt;. There's no filing cabinet of facts anywhere. It's all baked into billions of floating point numbers.&lt;/p&gt;

&lt;p&gt;There's a second, more literal kind of memory worth knowing about if you work with transformers: the &lt;strong&gt;KV cache&lt;/strong&gt;. During generation, the model stores the &lt;strong&gt;key&lt;/strong&gt; and &lt;strong&gt;value vectors&lt;/strong&gt; from earlier tokens so it doesn't have to recompute them for every new token. That's short-term, working memory, and it's a real engineering concern when you're thinking about inference cost and &lt;strong&gt;context length&lt;/strong&gt;. Long-term memory lives in the weights. Short-term memory lives in the cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  B — Bias
&lt;/h2&gt;

&lt;p&gt;This is the one word in the acronym that means two different things depending on who's using it, so it's worth being precise.&lt;/p&gt;

&lt;p&gt;In the strict architectural sense, the &lt;strong&gt;bias term&lt;/strong&gt; is a learnable number added to a neuron's weighted sum before the activation function is applied. Its job is to let the neuron shift its output independent of the input. Without a bias term, every neuron's output would be forced to pass through the origin, which limits what the network can represent. It's a small parameter, but every neuron has one, and they get updated during training just like the weights do.&lt;/p&gt;

&lt;p&gt;In the everyday sense that shows up in conversations about fairness, &lt;strong&gt;algorithmic bias&lt;/strong&gt; means something else: a systematic skew in a model's outputs that traces back to patterns in the training data. If the data overrepresents certain viewpoints, demographics, or outcomes, the model will reproduce that imbalance in its predictions. This kind of bias isn't a single parameter you can point to. It's an emergent property of what the model was trained on.&lt;/p&gt;

&lt;p&gt;Both are real, both matter, and it's worth knowing which one you're talking about when the word comes up.&lt;/p&gt;

&lt;h2&gt;
  
  
  E — Embedding
&lt;/h2&gt;

&lt;p&gt;Neural networks only operate on numbers, so before a model can do anything with a word, an image, or a chunk of data, that input has to be converted into a &lt;strong&gt;vector&lt;/strong&gt; of numbers. That vector is an &lt;strong&gt;embedding&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;What makes embeddings genuinely useful isn't just that they're numeric. It's that the training process arranges them so that things which are semantically related end up close together in that &lt;strong&gt;vector space&lt;/strong&gt;. Words with similar meanings cluster near each other. Concepts that show up in similar contexts get pulled toward each other during training. This is the property that makes &lt;strong&gt;retrieval&lt;/strong&gt;, &lt;strong&gt;similarity search&lt;/strong&gt;, and &lt;strong&gt;RAG&lt;/strong&gt; pipelines work at all. You're not matching strings, you're measuring distance in a high-dimensional space.&lt;/p&gt;

&lt;h2&gt;
  
  
  R — Relationships
&lt;/h2&gt;

&lt;p&gt;Once you have everything represented as vectors, the next question is how the model figures out which pieces of information matter to each other. In transformer architectures, this is the job of the &lt;strong&gt;attention mechanism&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Attention computes a relevance score between every &lt;strong&gt;token&lt;/strong&gt; and every other token in the context, using &lt;strong&gt;query&lt;/strong&gt;, &lt;strong&gt;key&lt;/strong&gt;, and &lt;strong&gt;value&lt;/strong&gt; projections. The scores determine how much weight each token gets when the model builds its next representation. This is how a model can connect a pronoun to the noun it refers to several sentences earlier, or weigh a modifier against the word it's modifying. It's a numeric answer to the question "what here is relevant to what," recomputed at every layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  S — Softmax
&lt;/h2&gt;

&lt;p&gt;At the very end of a prediction, a model has a list of raw scores, one for every possible next token, called &lt;strong&gt;logits&lt;/strong&gt;. Logits can be any real number, positive or negative, and they don't sum to anything meaningful on their own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Softmax&lt;/strong&gt; converts that list of raw scores into a &lt;strong&gt;probability distribution&lt;/strong&gt;. It exponentiates each score and divides by the sum of all the exponentiated scores, so the results are all positive and add up to exactly one. The token with the highest score gets the highest probability, but every other token still has some nonzero chance of being picked, which is what lets sampling strategies like &lt;strong&gt;temperature&lt;/strong&gt; and &lt;strong&gt;top-p&lt;/strong&gt; do their work. Softmax is the last numeric step before the model's internal computation turns into an actual choice.&lt;/p&gt;




&lt;p&gt;None of these seven ideas is hard to understand on its own. A &lt;strong&gt;neuron&lt;/strong&gt; is just simple math. An update is just a small correction. &lt;strong&gt;Softmax&lt;/strong&gt; is just a way of turning scores into probabilities. By themselves, none of these look like "intelligence."&lt;/p&gt;

&lt;p&gt;So where does the intelligence actually come from? The answer is &lt;strong&gt;scale&lt;/strong&gt;. A model like &lt;strong&gt;GPT&lt;/strong&gt; or &lt;strong&gt;Claude&lt;/strong&gt; isn't made of one &lt;strong&gt;neuron&lt;/strong&gt;, it's made of &lt;strong&gt;billions of them&lt;/strong&gt;. It isn't trained with one update, it's trained with trillions of tiny corrections. It doesn't compare two pieces of information once, it compares millions of them at the same time, across every sentence you type. No single piece is doing anything clever. What looks like intelligence shows up only when you take all seven of these simple ideas and run them together, over and over, at a scale no person could ever do by hand.&lt;/p&gt;

&lt;p&gt;That should make the whole thing more impressive, not less. Knowing that a model is "&lt;strong&gt;just&lt;/strong&gt;" &lt;strong&gt;numbers&lt;/strong&gt; and &lt;strong&gt;simple math doesn't take away the magic&lt;/strong&gt;, it explains where the magic comes from. The same way a single ant isn't smart but an ant colony can build complex structures, or a single water droplet doesn't have weather but billions of them together create a storm, a single neuron isn't smart, but billions of them working together can write, reason, and hold a conversation with you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Your Brain Beat Silicon to Every Idea in Modern AI - Nature Shipped It First</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Mon, 13 Jul 2026 05:36:51 +0000</pubDate>
      <link>https://dev.to/sreeni5018/your-brain-beat-silicon-to-every-idea-in-modern-ai-nature-shipped-it-first-47kl</link>
      <guid>https://dev.to/sreeni5018/your-brain-beat-silicon-to-every-idea-in-modern-ai-nature-shipped-it-first-47kl</guid>
      <description>&lt;p&gt;I put together an infographic recently that maps eleven core ideas in machine learning against eleven stages of how a human mind actually develops, from a small child looking at the world for the first time to an adult who plans, acts, and works with other people to get things done. The same child appears in every panel, growing up one stage at a time: reading under a tree while learning to observe, running across a field while learning sequence and order, studying at a desk while learning to remember. The visual repetition was deliberate. I wanted it to read less like a glossary of ML architectures and more like a single life, watched end to end, with the machine learning term sitting next to whatever human milestone it happens to rhyme with.&lt;/p&gt;

&lt;p&gt;That's the idea I want to unpack here. Eleven ideas from machine learning, laid out in the order a mind actually develops them, not the order textbooks introduce them or papers got published. &lt;strong&gt;CNNs before transformers. Memory before efficiency. Compression before creation.&lt;/strong&gt; Once you see it as a developmental ladder instead of a list of unrelated tricks, ML architecture starts looking less like a pile of clever engineering and more like a single story about how any intelligence, biological or artificial, comes to know a world.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3vbo2hawaifb286w0qj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi3vbo2hawaifb286w0qj.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Observe — CNNs and the Grammar of Seeing
&lt;/h2&gt;

&lt;p&gt;Before a child can do anything else, she has to look. Her senses take in the world, and somewhere in the visual cortex, raw light gets turned into edges, edges get combined into shapes, and shapes get combined into objects. Nobody teaches a toddler this hierarchy explicitly. It builds itself, layer by layer, from exposure.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;convolutional neural network (CNN)&lt;/strong&gt; does the same thing on purpose. Early layers learn to detect edges. Middle layers combine those edges into textures and parts. Later layers assemble parts into whole objects. The architecture isn't imitating vision as a clever engineering hack, it is rediscovering a principle that biological perception already runs on: complexity is built by composing simple, local patterns into progressively more abstract ones.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Recognize Order and Context — RNNs and the Arrow of Time
&lt;/h2&gt;

&lt;p&gt;Perception alone gives you snapshots. Understanding requires sequence. A child running through a park has to process motion, not stills, she needs to know what just happened to predict what happens next, and language works the same way, one word making sense only in light of the words before it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recurrent Neural Networks (RNNs)&lt;/strong&gt; were the first serious attempt to give machines that sense of "before." A &lt;strong&gt;hidden state&lt;/strong&gt; carries information forward from one step to the next, so the network's read of the current input is always colored by what it has already seen. It is a primitive kind of memory, just enough to string moments into a narrative instead of a slideshow.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Remember — LSTMs and the Discipline of Forgetting
&lt;/h2&gt;

&lt;p&gt;Pure sequence processing has a flaw: everything fades. Try to carry a thought across fifty steps with a vanilla RNN and it dissolves before it arrives. A child who studies for an exam doesn't just accumulate information, she has to decide what to keep and what to let go, and so does any system that wants to reason over a long span of time.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;LSTM (Long Short-Term Memory)&lt;/strong&gt; solves this with three gates working in concert: a &lt;strong&gt;forget gate&lt;/strong&gt; that decides what to discard, an &lt;strong&gt;input gate&lt;/strong&gt; that decides what new information deserves storage, and an &lt;strong&gt;output gate&lt;/strong&gt; that decides what to actually use right now. Underneath sits a &lt;strong&gt;cell state&lt;/strong&gt;, a kind of long-term memory highway that information can travel along largely undisturbed. This is not just an engineering patch. It is a formal answer to a question every learner faces: memory isn't about holding onto everything, it's about curating what matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Learn Faster with Less — GRUs and the Value of Simplicity
&lt;/h2&gt;

&lt;p&gt;Somewhere between childhood and competence, a person needs fewer examples to get better. A kid learning to run doesn't need to consciously relearn balance every time, the skill compresses and generalizes.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Gated Recurrent Unit (GRU)&lt;/strong&gt; is the LSTM's leaner sibling, merging some of those gates, cutting parameters, training faster with less data while holding onto most of the benefit. It's the architectural version of an athlete who's stopped thinking about every individual muscle movement and just runs. Efficiency, in both biological and artificial learning, tends to arrive after the underlying pattern has already been earned the hard way.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Focus on the Essence — Autoencoders and the Art of Compression
&lt;/h2&gt;

&lt;p&gt;Somewhere in adolescence or earlier, we all learn to filter. Not every detail in a scene matters, not every sentence in a conversation is load-bearing, and part of maturing is learning what the essence of a thing actually is.&lt;/p&gt;

&lt;p&gt;An &lt;strong&gt;autoencoder&lt;/strong&gt; is built entirely around that skill. An &lt;strong&gt;encoder&lt;/strong&gt; squeezes the input down into a compact &lt;strong&gt;latent representation&lt;/strong&gt;, throwing away noise and keeping structure, and a &lt;strong&gt;decoder&lt;/strong&gt; tries to reconstruct the original from that compressed form. What makes this interesting isn't the reconstruction, it's the bottleneck in the middle. Forcing information through a narrow space is exactly what forces a system, biological or artificial, to learn what's essential rather than what's merely present.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Imagine New Things — GANs and the Tension That Creates Novelty
&lt;/h2&gt;

&lt;p&gt;A child inventing a new idea out of familiar pieces, a new story built from characters she already knows, a new drawing that recombines shapes she's seen, is doing something qualitatively different from memorizing. She's generating.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generative Adversarial Networks (GANs)&lt;/strong&gt; formalize creation as a contest. A &lt;strong&gt;generator&lt;/strong&gt; tries to produce convincing fake data, a &lt;strong&gt;discriminator&lt;/strong&gt; tries to catch the fakes, and the two improve in lockstep, each one's progress forcing the other to get sharper. Real creativity, it turns out, often looks less like a single mind conjuring novelty from nothing and more like this: two internal pressures, one generative and one critical, pushing against each other until something genuinely new survives the argument.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Handle Uncertainty — VAEs and Living with Probability
&lt;/h2&gt;

&lt;p&gt;At some point, certainty stops being available and a person has to get comfortable with distributions instead of facts. Not "this is definitely what will happen," but "here's a reasonable range of what might."&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Variational Autoencoder (VAE)&lt;/strong&gt; takes the compression idea from step five and adds probability to it. Instead of encoding an input to a single fixed point, it encodes it to a distribution, then samples from that distribution to generate. This is a small mathematical move with a large philosophical consequence: it means the system's internal representation of the world isn't a single rigid answer, it's a shape of possibilities, which is a much more honest way to represent anything genuinely uncertain, including most of real life.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Attend to What Matters — Transformers and the End of Sequential Reading
&lt;/h2&gt;

&lt;p&gt;Sequential processing, even with memory gates, has a bottleneck: everything has to pass through one step at a time. But a mature mind doesn't read a room, or a paragraph, or a problem one token at a time and forget the beginning by the time it reaches the end. It holds the whole thing in view and decides, dynamically, what deserves the most attention right now.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;transformer's self-attention mechanism&lt;/strong&gt; does exactly that. Every token can attend directly to every other token, weighing relevance in parallel instead of waiting in a queue. It captures global context in one pass, and it is arguably the single architectural idea that unlocked everything that came after it, because a system that can attend to everything at once can also reason about everything at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Turn Thought into Words — LLMs and the Discipline of Expression
&lt;/h2&gt;

&lt;p&gt;Understanding is not the same as being able to say what you understand. A child who grasps an idea still has to learn to put it into language, one word chosen after another, each one constrained by everything said before it.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;Large Language Model (LLM)&lt;/strong&gt; does this at scale: given context, predict the next token, repeat, and coherent language emerges from that repetition. It's easy to describe this reductively, "it's just next-token prediction," but the same is technically true of human speech production and nobody finds that description satisfying there either. Expression, biological or artificial, is compression and prediction working together in real time, under the constraint of having to commit to one word before you're allowed to choose the next.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Plan, Use Tools, Retry — Agents and the Loop of Doing
&lt;/h2&gt;

&lt;p&gt;Knowing things and saying things is still short of acting in the world. Acting requires a loop: perceive the situation, form a plan, act, observe what happened, evaluate whether it worked, and if it didn't, retry with what you've learned.&lt;/p&gt;

&lt;p&gt;This is the architecture of an &lt;strong&gt;AI agent&lt;/strong&gt;, and it is also, not coincidentally, the architecture of anyone who has ever learned a skill by doing it badly first. The agent's advantage over a static model isn't intelligence, it's the loop itself, the willingness to be wrong, observe the consequences, and adjust. That loop is arguably where most real competence, human or artificial, actually comes from.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Use Tools and Other People — MCP and the Return to the World
&lt;/h2&gt;

&lt;p&gt;The last stage isn't really about the machine at all. It's about connection. A person's competence means very little in isolation, it becomes valuable when it can plug into other people, other systems, other sources of truth.&lt;/p&gt;

&lt;p&gt;The &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt; is the standardized version of that instinct: a protocol that lets an AI securely reach out to external tools, data sources, and people rather than reasoning in a sealed room. It's a fitting place to end the ladder, because it points outward. Every earlier stage was about building an internal world model, perception, memory, compression, expression. This last one is about admitting that no internal model, however good, is complete without a real connection to what's outside it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Part the Diagram Doesn't Explain
&lt;/h2&gt;

&lt;p&gt;Here's what stays with me after mapping all eleven stages side by side: the machine got remarkably good at replicating the mechanics. Perception, memory, generation, attention, action, connection, all of it, in some form, now runs in silicon. But a line I keep coming back to is that the machine became the clearest mirror humanity ever built, and mirrors don't originate anything, they reflect. The human carried something the mirror could not verify on its own: love, meaning, presence, whatever you want to call the thing that isn't a computation over tokens but shows up anyway, uninvited, whenever two people actually pay attention to each other.&lt;/p&gt;

&lt;p&gt;I don't think that's a knock on the architecture. If anything, building a system that mimics cognition this closely is what makes the gap so visible. You can formalize the forget gate. You cannot formalize why someone stays up late worrying about a person they love. The ladder above gets you from raw perception all the way to an agent that plans, acts, and connects to the world through protocols like MCP. It does not get you the thing beyond the latent space. That part, for now, is still ours to carry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Activation Functions, Explained With a Drone Show</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Thu, 09 Jul 2026 21:16:54 +0000</pubDate>
      <link>https://dev.to/sreeni5018/activation-functions-explained-with-a-drone-show-4lo7</link>
      <guid>https://dev.to/sreeni5018/activation-functions-explained-with-a-drone-show-4lo7</guid>
      <description>&lt;p&gt;I was watching a drone light show over a city skyline last week, a thousand drones forming shapes against the night sky, and it struck me that this is one of the cleanest ways to explain what an &lt;strong&gt;activation function&lt;/strong&gt; actually does inside a neural network. No math notation required to get the intuition. &lt;strong&gt;Just drones, brightness, and a rule.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the &lt;strong&gt;analogy&lt;/strong&gt;, the actual definition behind it, and where this concept sits inside a real deep learning pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is an Activation Function?
&lt;/h2&gt;

&lt;p&gt;An &lt;strong&gt;activation function is a mathematical rule&lt;/strong&gt; applied to the output of a neuron after it computes a &lt;strong&gt;weighted sum of its inputs&lt;/strong&gt;. It decides whether that neuron &lt;strong&gt;fires&lt;/strong&gt;, how strongly it fires, and what shape the signal takes before it gets passed to the next layer.&lt;/p&gt;

&lt;p&gt;Formally, a single neuron computes &lt;strong&gt;z = (w1*x1 + w2*x2 + ... + wn*xn) + b,&lt;/strong&gt; a linear combination of inputs, weights, and a bias term. The activation function is the step applied right after: &lt;strong&gt;a = f(z).&lt;/strong&gt; That f is the activation function, and the choice of f determines the behavior of the entire network.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The core reason it exists:&lt;/strong&gt; &lt;strong&gt;non linearity&lt;/strong&gt;. &lt;strong&gt;Without&lt;/strong&gt; an &lt;strong&gt;activation&lt;/strong&gt; function, or with a &lt;strong&gt;purely linear one&lt;/strong&gt;, stacking any number of &lt;strong&gt;layers collapses mathematically into a single linear transformation&lt;/strong&gt;. A hundred layer network with no activation function is functionally identical to one layer. You cannot model curves, boundaries, or complex patterns with that. &lt;strong&gt;Activation functions are what let a network bend, separate, and represent non linear relationships in data&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where It's Used in Deep Learning
&lt;/h2&gt;

&lt;p&gt;Activation functions show up at two distinct points in almost every deep learning architecture:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hidden layers&lt;/strong&gt;: every neuron in every hidden layer of a &lt;strong&gt;CNN&lt;/strong&gt;, &lt;strong&gt;RNN&lt;/strong&gt;, &lt;strong&gt;transformer&lt;/strong&gt;, or plain feedforward network applies an activation function to its weighted sum before passing output forward. This is where &lt;strong&gt;ReLU&lt;/strong&gt; and its variants dominate today.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Output layers:&lt;/strong&gt; the final layer's activation function is chosen based on the task. &lt;strong&gt;Sigmoid&lt;/strong&gt; for binary classification, &lt;strong&gt;softmax&lt;/strong&gt; for multi class classification, and often no activation, or a linear one, for regression tasks like price prediction.&lt;/p&gt;

&lt;p&gt;This is not an optional add-on. Every trained deep network you have used, image classifiers, language models, recommendation systems, has an activation function baked into every single neuron. It is one of the most fundamental architectural choices in the entire design, right alongside layer depth and width.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup for the Analogy
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Picture 1,000 drones in the sky&lt;/strong&gt;. Each drone is a neuron. A control signal tells each drone how bright to shine. That signal is the weighted sum described above, and on its own it carries no shape, no decision, no meaning yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgjeraxi22lhkwrcv3s7v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgjeraxi22lhkwrcv3s7v.png" alt=" " width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens Without an Activation Function
&lt;/h2&gt;

&lt;p&gt;If every drone just displays its raw signal value directly, you get a mess. Negative brightness values do not make physical sense, so you would need to clip or scale them arbitrarily, and even then there is no rule enforcing which drones should stand out and which should stay quiet. The result is a formless blob of light. &lt;strong&gt;No shape, no picture, no decision boundary.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the exact real world consequence of skipping the activation function: linear layers stacked without non linearity learn nothing more than a straight line through your data, regardless of depth.&lt;/p&gt;

&lt;h2&gt;
  
  
  ReLU: The Decisive Drone
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ReLU&lt;/strong&gt;, short for &lt;strong&gt;Rectified Linear Unit&lt;/strong&gt;, follows a blunt rule:&lt;/p&gt;

&lt;p&gt;If the &lt;strong&gt;incoming signal is negative&lt;/strong&gt;, the drone turns off completely. &lt;strong&gt;No light.&lt;/strong&gt;&lt;br&gt;
If the &lt;strong&gt;signal is positive&lt;/strong&gt;, the drone shines at exactly that &lt;strong&gt;brightness&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Mathematically, &lt;strong&gt;f(x) = max(0, x).&lt;/strong&gt; A signal of -5 means the drone goes dark. A signal of 8 means the drone shines at brightness 8.&lt;/p&gt;

&lt;p&gt;The effect at the network level is sharp, clean shapes. Only the neurons that are confident, meaning they &lt;strong&gt;received a positive signal, contribute anything.&lt;/strong&gt; &lt;strong&gt;Everyone else stays silent.&lt;/strong&gt; That is exactly why ReLU became the default activation function for hidden layers across most modern deep learning architectures: it produces sparse, decisive activations, it is cheap to compute since it is just a threshold check, and it largely avoids the gradient problems that plagued earlier activation functions.&lt;/p&gt;

&lt;p&gt;The tradeoff worth knowing: if a neuron's signal stays negative across every input it ever sees, it goes permanently dark and stops learning entirely. This is called the dead ReLU problem, and it is why variants like &lt;strong&gt;Leaky ReLU and GELU exist&lt;/strong&gt;, giving negative signals a small nonzero slope instead of a hard zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sigmoid: The Gradient Drone
&lt;/h2&gt;

&lt;p&gt;Sigmoid takes a completely different approach. Instead of an on/off switch, every signal gets smoothly converted into a brightness somewhere between 0% and 100%.&lt;/p&gt;

&lt;p&gt;A very negative signal fades toward 0%, but never fully reaches zero.&lt;br&gt;
A very positive signal climbs toward 100%, but never fully reaches max.&lt;/p&gt;

&lt;p&gt;The formula is &lt;strong&gt;f(x) = 1 / (1 + e^-x).&lt;/strong&gt; Every drone always emits some light, even if barely visible. The result across the network is soft, glowing gradients instead of hard shapes, which is why sigmoid is still the standard choice for output layers doing binary classification, where you want a smooth probability between 0 and 1 rather than a hard cutoff.&lt;/p&gt;

&lt;p&gt;The tradeoff: at the extremes, the curve flattens out almost completely. A signal of -10 and a signal of -20 produce nearly identical output. That flat region means the gradient shrinks toward zero during &lt;strong&gt;back-propagation&lt;/strong&gt;, which is the vanishing gradient problem. Stack enough sigmoid layers and the earliest layers in the network barely learn anything at all. That limitation is a big part of why ReLU displaced sigmoid as the default for hidden layers once networks started getting deep.&lt;/p&gt;

&lt;h2&gt;
  
  
  Other Activation Functions Worth Knowing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Tanh:&lt;/strong&gt; similar shape to sigmoid but outputs between -1 and 1 instead of 0 and 1. Centered at zero, which helps gradients flow slightly better than sigmoid, though it still suffers from the same vanishing gradient issue at the extremes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Softmax:&lt;/strong&gt; used almost exclusively in the output layer for multi-class classification. Converts a vector of raw scores into a probability distribution that sums to 1 across all classes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GELU:&lt;/strong&gt; a smoother variant of ReLU used in most modern transformer architectures, including the ones behind large language models. It weights inputs by their magnitude rather than applying a hard cutoff at zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping the Analogy Back to the Architecture
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuv4aspa2t5686hq658o8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuv4aspa2t5686hq658o8.png" alt=" " width="800" height="309"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The One Line Takeaway
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;An activation function is the rule each neuron follows to decide how strongly to fire.&lt;/strong&gt; Without that rule, you get a formless blob of numbers instead of a network capable of recognizing a pattern. With the right rule, sharp decisive shapes with &lt;strong&gt;ReLU&lt;/strong&gt;, smooth probabilistic gradients with sigmoid, or the smoother curves of &lt;strong&gt;GELU&lt;/strong&gt; &lt;strong&gt;inside a transformer&lt;/strong&gt;, you get a network that can actually learn something.&lt;/p&gt;

&lt;p&gt;Next time someone asks why activation functions matter, skip the equations first. Start with the drones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The Architecture of Time</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Wed, 08 Jul 2026 21:56:28 +0000</pubDate>
      <link>https://dev.to/sreeni5018/the-architecture-of-time-52j2</link>
      <guid>https://dev.to/sreeni5018/the-architecture-of-time-52j2</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;em&gt;What AI Taught Me About Living in the Present&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every AI model I build &lt;strong&gt;depends entirely on the past&lt;/strong&gt;. Every &lt;strong&gt;mindfulness teacher&lt;/strong&gt; I've listened to tells me to &lt;strong&gt;let the past go&lt;/strong&gt;. At first those felt like &lt;strong&gt;contradictions&lt;/strong&gt;, one treating history as everything, the other treating it as something to release. The longer &lt;strong&gt;I spent building AI systems&lt;/strong&gt;, the more I understood they weren't opposing ideas at all. They were describing two different kinds of intelligence, and somewhere between them &lt;strong&gt;I found something worth writing down.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxh5zcwcth588zwar60u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxxh5zcwcth588zwar60u.png" alt=" " width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Rearview Mirror of the Machine&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;One of the &lt;strong&gt;biggest misconceptions about AI is that it looks into the future.&lt;/strong&gt; It doesn't. &lt;strong&gt;Every prediction a machine makes is built entirely from yesterday&lt;/strong&gt;, whether it's forecasting demand, recommending your next movie, catching fraud before it happens, or &lt;strong&gt;generating the next word in a sentence&lt;/strong&gt;. There is exactly one source of knowledge behind all of it &lt;strong&gt;historical data.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Take that history away and the model doesn't grow cautious. It grows incapable. &lt;strong&gt;No patterns, no inference, no prediction, just silence. Machines don't see the future.&lt;/strong&gt; They learn the past well enough to respond to a moment they've never seen before, and that isn't a limitation we're still engineering our way out of. That's the entire mechanism.&lt;/p&gt;

&lt;p&gt;The more I sat with that, the more I suspected we aren't so different. Our experiences are our training data. Every success teaches us something, every mistake leaves a pattern behind, every failure quietly updates our model of the world. &lt;strong&gt;The past isn't baggage. It's data.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Human Edge&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;This is where we stop resembling the machine.&lt;/strong&gt; A model cannot choose to ignore its training. We can. Spend too long living in yesterday and memory starts masquerading as identity. Old disappointments harden into expectations, past failures get filed as permanent truths, and yesterday's fear starts making today's decisions without asking permission. We call it experience. Sometimes it's just unprocessed history, still running the meeting.&lt;/p&gt;

&lt;p&gt;The present hands us something no machine has ever touched choice. You can interrupt a habit halfway through the reflex. You can forgive someone after years of meaning not to. You can change your mind for no reason except that you've changed, and become someone slightly different between one breath and the next. &lt;strong&gt;That capacity isn't a gap in human intelligence. It's the whole point of it.&lt;/strong&gt; Our history shapes us, but it was never supposed to finish the sentence for us.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Real Algorithm&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;So who's right, the engineer who trusts the data or the teacher who says to release it?&lt;/strong&gt; Both, describing the same truth from opposite ends. Ignore the past completely and nothing accumulates, you just repeat what you haven't yet learned from. Live entirely inside it and you calcify, the human version of an overfit model, tuned to perfection for a world that has already moved on without you. &lt;strong&gt;Wisdom sits in the narrow space between those two failures. Not the absence of history. Not its rule. Its use.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Building Your Own Architecture&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Next time an old mistake finds its way back into your thoughts, resist calling it a verdict. Call it what it is training data&lt;/strong&gt;. Take the lesson, and leave the weight where it fell. Your history can brief the decision, but it doesn't get to make it. That belongs to whoever you are in this exact moment, choosing.&lt;/p&gt;

&lt;p&gt;Maybe that's the only architecture worth building, one where history informs the moment without ever being allowed to replace it. The future was never something we predict from a distance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;em&gt;It's something assembled, one present moment at a time, by someone who remembered enough to know better and stayed awake enough to choose anyway.&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The Architecture of Becoming: Why Your Life is a Transformer Network</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Thu, 02 Jul 2026 03:46:57 +0000</pubDate>
      <link>https://dev.to/sreeni5018/the-architecture-of-becoming-why-your-life-is-a-transformer-network-4aco</link>
      <guid>https://dev.to/sreeni5018/the-architecture-of-becoming-why-your-life-is-a-transformer-network-4aco</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;For most of my life, I thought my journey followed a predictable, linear script &lt;strong&gt;Go to school. Get a degree. Get a job. Gain experience. Get promoted.&lt;/strong&gt; That was the whole story or so I believed.&lt;/p&gt;

&lt;p&gt;Then I spent a &lt;strong&gt;few years designing Generative AI systems for a living, and that neat little script stopped making sense.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most of my days are spent building on top of &lt;strong&gt;Large Language Models&lt;/strong&gt; (LLM)  navigating Transformer architectures, &lt;strong&gt;RAG pipelines&lt;/strong&gt;, multi &lt;strong&gt;agent workflows&lt;/strong&gt;, &lt;strong&gt;MCP servers&lt;/strong&gt;, and the &lt;strong&gt;guardrail harnesses&lt;/strong&gt; required to keep the whole thing from going off the rails. One evening, while working through the architecture for an Agentic AI solution to solve a complex enterprise use case, something stopped me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It wasn't just a system diagram anymore.&lt;/strong&gt; It looked like a map of my own life. Not in a poetic, greeting card way but in a literal, structural way. The stages matched up almost embarrassingly well.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Here is what I saw, phase by phase.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffs1n1ecngzr922741zev.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffs1n1ecngzr922741zev.png" alt=" " width="800" height="794"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The Vector Space of Childhood
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Embeddings&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkfrild33xczbtkspnbx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkfrild33xczbtkspnbx.png" alt=" " width="326" height="489"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Before an LLM can process a single sentence, it has to turn words into &lt;strong&gt;embeddings&lt;/strong&gt;. Raw text means nothing to a model; it’s just empty symbols. An embedding model places those symbols into a &lt;strong&gt;high dimensional vector space where distance and direction carry meaning.&lt;/strong&gt; Words end up near each other because they share context. Nothing useful happens downstream until this space exists.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Life Parallel:&lt;/strong&gt; This is the hidden work of childhood. I wasn’t merely collecting facts; I was building the high dimensional coordinate system those facts would later live in.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;My early teachers weren't just handing me information they were shaping my cognitive geometry.&lt;/strong&gt; A lesson would place a new idea near something I already half understood. A correction would nudge two concepts a little further apart. None of it meant anything in isolation. Looking back, they weren't filling a blank disk; they were establishing the baseline vectors I would use to make sense of the entire universe.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The Layers of Understanding
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: The Encoder Stack&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fixsqd59y2qbqaiprqx39.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fixsqd59y2qbqaiprqx39.png" alt=" " width="311" height="483"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once you have embeddings, the encoder takes over. An encoder doesn’t generate new text; its sole purpose is to build a deeper, more abstract representation of the input. Each layer takes the output of the previous layer and distills it further.&lt;/p&gt;

&lt;p&gt;Formal education did exactly this to me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Elementary School:&lt;/strong&gt; Decoded raw symbols on a page.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Middle School:&lt;/strong&gt; Learned to reason with core concepts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High School:&lt;/strong&gt; Introduced systemic abstraction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;University:&lt;/strong&gt; Grounded those abstractions into engineering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Graduate Work:&lt;/strong&gt; Shifted the focus entirely to systems thinking.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each stage wasn't just a new textbook it was a &lt;strong&gt;new layer stacked on top of the last&lt;/strong&gt;. The outside world hadn't changed at all. What changed was the depth of the representation I could build. When I solve a complex problem today, it’s not because the problem got easier. It’s because I have more layers processing the input.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. The Reality Check
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Validation Sets &amp;amp; Loss Signals&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwnq28y0tcwy8nt0s4fjw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwnq28y0tcwy8nt0s4fjw.png" alt=" " width="331" height="488"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Every model eventually has to leave the clean, synthetic world of the training loop.&lt;/strong&gt; For me, that happened on day one of my first real job. School had been a highly curated training set labeled, clean, and forgiving. The workplace was noisy, unlabeled, and entirely unimpressed by my resume.&lt;/p&gt;

&lt;p&gt;My first &lt;strong&gt;internship&lt;/strong&gt; was a brutal &lt;strong&gt;validation set&lt;/strong&gt;. Reality doesn't grade on a curve; it checks your output against hard constraints.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Every bug I shipped was a &lt;strong&gt;loss signal&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Every rough code review was a &lt;strong&gt;weight adjustment&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Every design that collapsed in production closed the gap between what I thought would work and what actually did.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I used to think failure meant the learning process had stalled. Eventually, I understood that &lt;strong&gt;failure &lt;em&gt;was&lt;/em&gt; the mechanism&lt;/strong&gt;. It wasn't an error in the system; it was the gradient descent optimization of my career.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. The Act of Generation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Auto Regressive Decoders&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5w8mnninbm3qmwynv9l0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5w8mnninbm3qmwynv9l0.png" alt=" " width="283" height="487"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At some point, the balance flipped. I transitioned from absorbing context to producing it.&lt;/p&gt;

&lt;p&gt;This is the part of a Transformer that gets less attention than it deserves &lt;strong&gt;the decoder cannot see what is coming next&lt;/strong&gt;. It only knows what it has already produced, and every new token depends entirely on the sequence that came before it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is exactly what a career feels like after a decade.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The architectural decision I made last year still shapes what I can build today.&lt;/li&gt;
&lt;li&gt;The professional reputation I built five years ago quietly decides which opportunities are visible to me now.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I don't get to go back and retroactively edit the tokens I’ve already put out into the world. I can only take the sequence as given and generate the next token as intelligently as possible.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The Shift to External Context
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Retrieval Augmented Generation (RAG)&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpz1kw0tnxaz6njkx52rk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpz1kw0tnxaz6njkx52rk.png" alt=" " width="308" height="482"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is a persistent myth that being an expert means having all the answers memorized.&lt;/strong&gt; &lt;strong&gt;Modern AI abandoned that ideology a long time ago.&lt;/strong&gt; A model that only knows what is baked into its static weights is severely limited. Real production systems reach outside themselves triggering a &lt;strong&gt;RAG lookup against a technical document or querying a database.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The moment I stopped trying to hold everything in my head and started building better external retrieval systems, my engineering velocity exploded. The strongest leads I know aren't the ones with the most trivia crammed into short-term memory. They are the ones with the most &lt;strong&gt;efficient retrieval instincts&lt;/strong&gt; they know exactly which document, API, or expert to query at the exact moment they hit a constraint.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Protocols of Connection
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Model Context Protocol (MCP)&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Folcxtzafoi03u1ip2r9p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Folcxtzafoi03u1ip2r9p.png" alt=" " width="322" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;As enterprise AI matured, engineers ran into a scaling wall &lt;strong&gt;you can't hardcode a model to every single custom tool it might need.&lt;/strong&gt; It doesn't scale. &lt;strong&gt;Tools like MCP (Model Context Protocol)&lt;/strong&gt; solve this by creating an open, standard interface so models can safely read data and touch tools without tight coupling.&lt;/p&gt;

&lt;p&gt;I hit that exact same scaling wall when my career grew past what I could personally manage. I couldn’t sit in every meeting or audit every codebase.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Documentation became my API.&lt;/strong&gt; Writing clear wikis, run books, and design docs allowed other teams to query my context without needing to interrupt my runtime. I stopped equating value with personal execution and started thinking of myself as an interface that others could cleanly build upon.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Dynamic Adaptability
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Agent Skills &amp;amp; Lightweight Adapters&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fly4631uxeskum3tiryhf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fly4631uxeskum3tiryhf.png" alt=" " width="310" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A production agent doesn't retrain its entire multi billion parameter foundation model just to learn how to use a new piece of software. Instead, we register &lt;strong&gt;Agent Skills&lt;/strong&gt; scoped, &lt;strong&gt;modular capabilities&lt;/strong&gt; &lt;strong&gt;loaded dynamically at the application layer when the environment calls for them&lt;/strong&gt;, leaving the underlying foundational weights completely untouched.&lt;/p&gt;

&lt;p&gt;This is exactly how acquiring a new skill works. Learning Kubernetes didn't rewrite how I think about distributed systems; it just loaded a new "&lt;strong&gt;skill&lt;/strong&gt;" on top of the engineering fundamentals I already possessed. New tech stacks don't erase your foundational experience they are just specialized functional blocks loaded into your prompt context when a specific task demands them.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. The Enterprise Hivemind
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Multi-Agent Workflows &amp;amp; A2A Communication&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqko6evg1swncws2lz01n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqko6evg1swncws2lz01n.png" alt=" " width="318" height="443"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nothing serious runs on a single, isolated model anymore. Robust architectures rely on &lt;strong&gt;A2A (Agent-to-Agent) communication&lt;/strong&gt; within multi agent workflows. An &lt;strong&gt;engineering agent&lt;/strong&gt; writes code, a &lt;strong&gt;security agent&lt;/strong&gt; audits it, a &lt;strong&gt;compliance agent&lt;/strong&gt; checks it against regulations, and a financial agent estimates the API cost. None of these agents can see inside each other’s prompt history. They don't need to. They talk asynchronously, exchanging structured messages back and forth. Yet, through this protocol, the hivemind converges on a coherent solution.&lt;/p&gt;

&lt;p&gt;That is the exact definition of a cross functional corporate team. Every alignment meeting, sprint handoff, and architecture review is an exercise in A2A communication an exchange of structured text payloads between humans who cannot see each other's internal reasoning. We were running multi agent workflows long before we wrote code for them.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. The Governance Layer
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Guardrails &amp;amp; Evals&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmsn06tbxng7o2dnyg9rf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmsn06tbxng7o2dnyg9rf.png" alt=" " width="295" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you build production AI systems long enough, you learn a humbling lesson: &lt;strong&gt;the most capable model is often the most dangerous one if left unguided.&lt;/strong&gt; Without evaluators, systemic guardrails, and a strict system prompt, a brilliant model will eventually drift into confident, destructive nonsense. Raw compute requires a harness to be useful.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The Life Parallel:&lt;/strong&gt; This is what character actually is.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Intelligence tells you what you &lt;em&gt;can&lt;/em&gt; do; character determines what you &lt;em&gt;should&lt;/em&gt; do. Discipline, integrity, and humility are not products of raw cognitive horsepower. They are the governance layer built around your mind. Over a long enough time horizon, the strength of your guardrails matters infinitely more than the speed of your processing core.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. The Anchor of the Past
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;em&gt;Metaphor: Causal Self Attention&lt;/em&gt;
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe98hk1lozg0cpapkml3m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe98hk1lozg0cpapkml3m.png" alt=" " width="310" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is a common assumption that education is a phase that ends when work begins a clean line between learning and doing. The architecture of a Transformer proves otherwise.&lt;/p&gt;

&lt;p&gt;Most modern models don’t have a separate, isolated encoder pass. Every single token is generated via &lt;strong&gt;causal self-attention&lt;/strong&gt;, looking back across the entire historical sequence generated so far. Nothing gets thrown away once it is in the context window. A token generated at step 5 still directly alters the probability distribution of step 50,000.&lt;/p&gt;

&lt;p&gt;That is how my past actually functions.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An architectural decision I make today is still weighted by a principle an old mentor shared twenty years ago.&lt;/li&gt;
&lt;li&gt;A leadership choice I make this morning carries the heavy context of my worst early management failures.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The past isn't a dusty archive I occasionally visit; it lives inside my active context window, one attention pass away at all times.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Break in the Metaphor: Continuous Optimization
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Here is where the comparison fundamentally breaks, and it’s the most beautiful part of the realization.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Production machine learning models eventually &lt;strong&gt;freeze their weights&lt;/strong&gt;. &lt;strong&gt;The training loop closes&lt;/strong&gt;, the checkpoint ships, and from that second onward, the model runs static inference. It cannot learn from the user interactions it handles today unless an engineering team aggregates the logs and kicks off an expensive retraining run later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Humans don't have that limitation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is no boundary between our &lt;strong&gt;"training phase"&lt;/strong&gt; and our &lt;strong&gt;"production phase."&lt;/strong&gt; Every hard conversation, every failed launch, and every sudden market shift changes our internal weights in real time while we are actively doing the job. We don't wait for a maintenance window to optimize. We evolve mid flight.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is no final deployment checkpoint for a person.&lt;/strong&gt; There is only the next token you choose to generate, and the system you build while generating it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrh09fkbrdkh4s7ihmu1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvrh09fkbrdkh4s7ihmu1.png" alt=" " width="800" height="43"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>career</category>
      <category>llm</category>
    </item>
    <item>
      <title>Stop Chasing Smarter Models. Start Engineering Better Context.</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Sun, 21 Jun 2026 04:37:35 +0000</pubDate>
      <link>https://dev.to/sreeni5018/stop-chasing-smarter-models-start-engineering-better-context-5014</link>
      <guid>https://dev.to/sreeni5018/stop-chasing-smarter-models-start-engineering-better-context-5014</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;We spend our entire lives &lt;strong&gt;chasing happiness in the next job&lt;/strong&gt;, the &lt;strong&gt;next city&lt;/strong&gt;, the &lt;strong&gt;next relationship&lt;/strong&gt; only to realize somewhere along the way that happiness was never out there. It was a state of mind. Something you &lt;strong&gt;cultivate from within&lt;/strong&gt;, not something you find.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enterprise AI is making the exact same mistake.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every week, teams pour budget into the next frontier model &lt;strong&gt;bigger parameters, wider context windows&lt;/strong&gt;, stronger benchmarks convinced that intelligence is the thing they're missing. But the agents still &lt;strong&gt;hallucinate&lt;/strong&gt;. The pipelines still break. The outputs still disappoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Because the problem was never the model.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Just like happiness, the answer isn't out there in a smarter LLM.&lt;/strong&gt; It's in the environment you build around the model you already have. It's in the &lt;strong&gt;quality of the context the data it receives&lt;/strong&gt;, the memory it carries, the instructions it operates on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is Context Engineering.&lt;/strong&gt; And it's the discipline most teams are ignoring while they wait for the next model release to save them.&lt;/p&gt;

&lt;p&gt;You can &lt;strong&gt;hand a genius a disorganized pile of corrupted documents, conflicting instructions, and broken APIs&lt;/strong&gt; and they will still fail. &lt;strong&gt;Conversely&lt;/strong&gt;, &lt;strong&gt;hand an average worker a clean playbook, sharp guardrails, and exactly the data they need and they will execute flawlessly every single time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That single observation should reshape how every enterprise AI team allocates its engineering budget.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feehdwtbwtwz1htd5jcv7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feehdwtbwtwz1htd5jcv7.png" alt=" " width="800" height="377"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Illusion of the "Smart" Agent
&lt;/h2&gt;

&lt;p&gt;Most &lt;strong&gt;enterprise teams treat LLMs like databases stuffed with world knowledge.&lt;/strong&gt; They are not. An &lt;strong&gt;LLM is a reasoning engine.&lt;/strong&gt; It &lt;strong&gt;doesn't retrieve answers it reasons toward them&lt;/strong&gt;. The quality of that reasoning is almost entirely determined by the quality of what you put in front of it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When an agent fails in production, the gut reaction is to upgrade to a larger, more expensive model.&lt;/strong&gt; But when you actually dig into the execution logs, the root cause is rarely a lack of raw intelligence. It is almost always a data failure:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The system fetched the wrong vector chunk.&lt;/li&gt;
&lt;li&gt;The API schema passed to the agent was ambiguous or incomplete.&lt;/li&gt;
&lt;li&gt;The conversation history was an unpruned wall of noise.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of those failures get fixed by a bigger model. They get fixed by better context.&lt;/p&gt;

&lt;p&gt;A well optimized, smaller model operating on clean, deterministic context will consistently outperform a massive frontier model reasoning in the dark. I have watched this play out on real enterprise workloads. The model upgrade didn't help. The context pipeline redesign did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Token Economy Problem
&lt;/h2&gt;

&lt;p&gt;Chasing raw model intelligence isn't just an engineering trap it's an economic liability.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvifk04a43tfxqz0srkcn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvifk04a43tfxqz0srkcn.png" alt=" " width="800" height="166"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In enterprise AI, every token is a financial transaction.&lt;/strong&gt; When you dump unrefined, unstructured data into a massive context window and rely on the model's "intelligence" to sort through the noise, you are paying a premium for a problem that should have been solved upstream. Runtime costs spike. Latency climbs. And the outputs are still inconsistent because the ambiguity was never resolved it was just delegated to the model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is lazy architecture.&lt;/strong&gt; It feels like a shortcut, but it compounds into a long-term cost problem.&lt;/p&gt;

&lt;p&gt;The alternative is Context Engineering: building deterministic pipelines that deliver exactly the right information, in the right structure, at the right moment. When you do this well, you stop needing frontier-scale models for routine tasks. You route cheaper, faster, specialized agents to handle the work — and you reserve the heavier models for the exceptions that actually need them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context Engineering in Practice
&lt;/h2&gt;

&lt;p&gt;Context Engineering is not a single technique. It is a discipline that spans three layers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wvlvy5z5tr17kjaqcgi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2wvlvy5z5tr17kjaqcgi.png" alt=" " width="800" height="302"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Retrieval Quality&lt;/strong&gt; — Moving Beyond Simple Vector Search&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Traditional RAG systems retrieve based on semantic similarity.&lt;/strong&gt; That works for surface level lookups, but enterprise data is relational. A customer record connects to contracts, which connect to support history, which connect to renewal status. &lt;strong&gt;GraphRAG architectures&lt;/strong&gt; capture those relationships explicitly so the agent receives not just a matching chunk, but the business context that surrounds it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Simple vector distance gets you close. Knowledge graphs get you accurate.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Dynamic Context Pruning&lt;/strong&gt; — Protecting the Token Budget&lt;/p&gt;

&lt;p&gt;Not everything in your context window deserves to be there. Long conversation histories, redundant tool outputs, and boilerplate instructions accumulate fast. Dynamic pruning is the practice of continuously evaluating what stays, what gets compressed, and what gets dropped — ensuring the agent only operates on high-signal data.&lt;/p&gt;

&lt;p&gt;Every token you cut from noise is a token you can spend on signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. State and Memory Management&lt;/strong&gt; — Teaching Agents What to Remember&lt;/p&gt;

&lt;p&gt;An agent with poor memory management treats every turn like it's the first. It re-fetches data it already has, loses track of prior decisions, and fails to carry forward the context that matters. Proper short-term and long-term memory architecture ensures agents accumulate useful state across a workflow and discard the rest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Raw Intelligence Still Matters
&lt;/h2&gt;

&lt;p&gt;This is not an argument for dumb models. Intelligence and context are not competitors they are complements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context does the heavy lifting for the predictable 80%.&lt;/strong&gt; It ensures the agent knows where it is, what it has to work with, and what it's supposed to do. For that 80%, &lt;strong&gt;a well contextualized smaller model is not just good enough it's better, faster, cheaper, and more consistent.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But the other 20% the undocumented API errors, the unexpected user edge cases, the multi-step recoveries that is where raw reasoning earns its place. You still need intelligence for exception handling. The difference is that you should be deploying it surgically, not as a substitute for good architecture.&lt;/p&gt;

&lt;p&gt;Think of model intelligence as the safety net, not the foundation. It catches what falls through. Context engineering is the foundation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Shift That Matters
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwee1zss0sbjd6xn1f81n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwee1zss0sbjd6xn1f81n.png" alt=" " width="799" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The teams building production grade AI systems in 2025 are not the ones waiting for the next model release. They are the ones mastering the &lt;strong&gt;scaffolding that surrounds the model the Agent Harness&lt;/strong&gt; the &lt;strong&gt;context pipelines&lt;/strong&gt;, &lt;strong&gt;the memory layers&lt;/strong&gt;, the &lt;strong&gt;routing logic&lt;/strong&gt;, the &lt;strong&gt;governance controls.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The competitive advantage in enterprise AI is not access to a bigger model. Everyone has access to the same frontier models. The advantage is in how well you engineer the environment those models operate in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stop waiting for a smarter brain. Start building a better world for the brain you already have.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Most Enterprise AI Agents Fail in Production for the Same Reason And It's Not the Model</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Fri, 19 Jun 2026 01:49:33 +0000</pubDate>
      <link>https://dev.to/sreeni5018/most-enterprise-ai-agents-fail-in-production-for-the-same-reason-and-its-not-the-model-4ad7</link>
      <guid>https://dev.to/sreeni5018/most-enterprise-ai-agents-fail-in-production-for-the-same-reason-and-its-not-the-model-4ad7</guid>
      <description>&lt;h2&gt;
  
  
  Because intelligence alone is never enough.
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F542evrgz6cgnyyuq53tc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F542evrgz6cgnyyuq53tc.png" alt=" " width="799" height="168"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There's a question I keep hearing from enterprise teams who are just starting to productionize  AI agents:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;"We've got great prompts. The model performs well in testing. Why does it still fail in production?"&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The reason is almost always the same &lt;strong&gt;they built the intelligence.&lt;/strong&gt; &lt;strong&gt;They didn't build the system around it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's it. That's the whole failure pattern. The model is fine. The engineering discipline surrounding it wasn't applied.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Here's the analogy I use to explain the difference.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  An AI Agent Is a Self-Driving Car
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Not metaphorically. Structurally.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both operate in dynamic, unpredictable environments.&lt;/strong&gt; &lt;strong&gt;Both make&lt;/strong&gt; &lt;strong&gt;real time decisions with incomplete information&lt;/strong&gt;. Both can fail not because they're dumb, but because the environment surprises them in ways nobody anticipated. And in both cases, the intelligence of the system (the model, the sensors, the neural net) is only one layer of what makes it trustworthy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fod6bmgridtsr4snvi2ct.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fod6bmgridtsr4snvi2ct.png" alt=" " width="800" height="60"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When you break it down, &lt;strong&gt;three distinct engineering disciplines make a self-driving car work.&lt;/strong&gt; The same three disciplines make an &lt;strong&gt;AI agent work.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 1: Prompt Engineering = Destination and Driving Instructions
&lt;/h2&gt;

&lt;p&gt;Before you put a self driving car on the road, you configure it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Where are we going?&lt;/li&gt;
&lt;li&gt;Which route is preferred?&lt;/li&gt;
&lt;li&gt;What's the speed limit?&lt;/li&gt;
&lt;li&gt;Are there constraints? (No highways. No toll roads. Arrive by 3 PM.)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2tp79vfgtrekbubs818.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo2tp79vfgtrekbubs818.png" alt=" " width="766" height="1716"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The car doesn't invent the mission. You give it one precisely, explicitly, in a format it can act on.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt Engineering does exactly the same thing for an AI agent. It defines:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The goal and scope of the task&lt;br&gt;
The rules and constraints it must follow&lt;br&gt;
The persona and tone it should operate with&lt;br&gt;
The guardrails that bound its behavior&lt;br&gt;
The expected format and outcome of its output&lt;/p&gt;

&lt;p&gt;Without clear prompts, the agent does what a car does without a destination it moves, but not toward anything useful. It might wander into edge cases, confabulate, or execute the wrong task with full confidence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real example:&lt;/strong&gt; An &lt;strong&gt;Ecommerce support agent&lt;/strong&gt; told only to "help customers" will happily process a refund, cancel an active shipment, and escalate to a manager all for the same complaint because nobody told it which action to take first, or when escalation is appropriate. The model is working fine. The briefing failed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt Engineering is the briefing. It's not optional, and it's not a one-time job. As your tasks evolve, so should the prompts.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 2: Context Engineering = Situational Awareness
&lt;/h2&gt;

&lt;p&gt;A self-driving car with perfect instructions will still crash if it can't see what's around it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's why autonomous vehicles carry:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;GPS and real-time maps&lt;/li&gt;
&lt;li&gt;Lidar and radar sensors&lt;/li&gt;
&lt;li&gt;Camera feeds processing the road ahead&lt;/li&gt;
&lt;li&gt;Weather and road condition data&lt;/li&gt;
&lt;li&gt;Traffic pattern feeds&lt;/li&gt;
&lt;li&gt;Pedestrian detection systems&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc5ep2bv8u10s06no4yy0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc5ep2bv8u10s06no4yy0.png" alt=" " width="764" height="1704"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All of this is context live, environmental, dynamic information that allows the vehicle to make intelligent decisions in the moment, not just based on pre-loaded instructions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An AI agent has the same problem. The base LLM is trained on historical data.&lt;/strong&gt; It doesn't know about your &lt;strong&gt;enterprise data&lt;/strong&gt;, &lt;strong&gt;your customer's current account status&lt;/strong&gt;, the document that was updated yesterday, or the conversation that happened last week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Real example:&lt;/strong&gt; A banking support agent is asked "what's the status of my loan application?" The model knows everything about loans in general. It knows nothing about this customer's application filed three days ago. Without retrieval &lt;strong&gt;RAG pulling the customer's record in real time the agent either hallucinates a status or says it doesn't have access&lt;/strong&gt;. Both outcomes destroy trust. The model is fine. The context layer wasn't built.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Context Engineering fills that gap. It's how you inject:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;RAG and GraphRAG&lt;/strong&gt; — retrieval of relevant documents and structured knowledge&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory systems&lt;/strong&gt; — both short-term (within session) and long-term (across sessions)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP Servers&lt;/strong&gt; — access to external tools, APIs, and services&lt;/li&gt;
&lt;li&gt;Enterprise knowledge bases — internal policies, product documentation, historical data&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User history and preferences&lt;/strong&gt; — the personalization layer&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time data feeds&lt;/strong&gt; — current state of the world the agent is operating in&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Context is not a prompt engineering problem. It's an infrastructure problem.&lt;/strong&gt; Getting the right information to the agent at the right moment, in the right format, with the right freshness that's an entirely different discipline with its own architecture, its own tooling, and its own failure modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A well prompted agent with poor context is like a skilled driver in a blindfolded car.&lt;/strong&gt; The instructions are clear. The execution is impossible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Layer 3: Harness Engineering = Safety, Recovery, and Accountability
&lt;/h2&gt;

&lt;p&gt;Here's where most teams underinvest.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Even the most advanced autonomous vehicle isn't deployed without a full safety stack.&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Collision detection and emergency braking&lt;/li&gt;
&lt;li&gt;Lane departure warnings&lt;/li&gt;
&lt;li&gt;Route recalculation when roads are blocked&lt;/li&gt;
&lt;li&gt;Telemetry for monitoring vehicle state&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Black-box logging for post-incident investigation&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Human override capability&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Regulatory compliance systems&lt;/li&gt;
&lt;li&gt;Redundant sensor fusion&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzvv8abe219kzb0uopz9p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzvv8abe219kzb0uopz9p.png" alt=" " width="800" height="1667"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the harness — the layer that doesn't make the car smarter, but makes it safer. It's the layer that catches failures before they become disasters, and that proves what happened when they do.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent Harness Engineering is the same idea applied to AI systems&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;State Management&lt;/strong&gt; — knowing where the agent is in a multi-step workflow&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Checkpointing&lt;/strong&gt; — saving progress so failures don't require starting over&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-in-the-Loop (HITL)&lt;/strong&gt; — escalation paths when confidence is low or stakes are high&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observability&lt;/strong&gt; — traces, logs, and dashboards that show you what the agent did and why&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails and Content Controls&lt;/strong&gt; — preventing harmful or out-of-scope outputs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool Access Control&lt;/strong&gt; — scoping what the agent can call and with what permissions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluation Pipelines&lt;/strong&gt; — continuous testing against ground truth to catch regression&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recovery Logic&lt;/strong&gt; — graceful degradation when tools fail or context is unavailable&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security and Governance&lt;/strong&gt; — audit trails, access controls, compliance hooks&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Real example:&lt;/strong&gt; An HR onboarding agent is mid-workflow — it has created a user account, sent a welcome email, and is about to provision software licenses when the identity service times out. Without checkpointing, the entire workflow restarts from scratch: duplicate account, duplicate email, confused new hire. Without observability, the engineering team doesn't even know it happened until someone complains. The model executed perfectly. The harness wasn't there to catch the infrastructure failure.&lt;/p&gt;

&lt;p&gt;The harness doesn't change what the agent can do. It changes what the agent will do under pressure  which is when it matters most.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Failures Still Happen Even When You've Done Everything Right
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgeygfoyqjoc8an0ngg2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgeygfoyqjoc8an0ngg2.png" alt=" " width="799" height="165"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here's the truth every production AI team eventually confronts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Even with all three layers in place&lt;/strong&gt; solid &lt;strong&gt;prompts&lt;/strong&gt;, rich &lt;strong&gt;context&lt;/strong&gt;, a well engineered &lt;strong&gt;harness&lt;/strong&gt; your &lt;strong&gt;agent will still make mistakes. Not occasionally.&lt;/strong&gt; Regularly enough that you need a plan for it.&lt;/p&gt;

&lt;p&gt;This is not a model quality problem. &lt;strong&gt;It is a fundamental property of the environment these systems operate in.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both autonomous vehicles and AI agents face the same four realities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic environments&lt;/strong&gt; — the world changes faster than any training set or prompt update cycle can track&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Incomplete information&lt;/strong&gt; — no matter how good your retrieval is, the context is always partial&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unseen edge cases&lt;/strong&gt; — production traffic will surface combinations that no benchmark, red team, or test suite anticipated&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cascading conditions&lt;/strong&gt; — two situations your agent handles perfectly in isolation can combine into something it has never encountered&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;No amount of engineering eliminates these realities. What engineering does is change how you respond to them.&lt;/p&gt;

&lt;p&gt;You can have:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Clear, tested prompts&lt;/li&gt;
&lt;li&gt;Rich, well-curated context&lt;/li&gt;
&lt;li&gt;A well-designed harness with observability and recovery&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;And the agent will still make mistakes. The difference is whether those mistakes are visible, recoverable, and traceable — or silent, destructive, and impossible to debug.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The goal is never zero failures. The goal is:&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Detect failures earlier. Recover faster. Prove what happened. Continuously improve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's what the harness is for. That's what observability is for. That's what HITL is for.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;If someone asks you to explain all three disciplines in a single breath&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftzgwpgh8qftlru2m8rm4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftzgwpgh8qftlru2m8rm4.png" alt=" " width="706" height="666"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt Engineering&lt;/strong&gt; tells the agent where to go. &lt;strong&gt;Context Engineering&lt;/strong&gt; helps it understand where it is. &lt;strong&gt;Harness Engineering&lt;/strong&gt; helps it arrive safely, recover when things go wrong, and prove what happened along the way.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means for Enterprise AI Teams
&lt;/h2&gt;

&lt;p&gt;Most teams are over invested in &lt;strong&gt;Layer 1&lt;/strong&gt; and under invested in &lt;strong&gt;Layers 2 and 3&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt Engineering&lt;/strong&gt; gets the most attention because it's visible, iterable, and produces immediate results. It's also the layer that impresses in demos. &lt;strong&gt;Context Engineering&lt;/strong&gt; is harder because it requires data infrastructure, retrieval pipelines, and integration work. &lt;strong&gt;Harness Engineering&lt;/strong&gt; is hardest because it requires thinking about failure modes before they happen.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiu3u2o231y1w8ubfkp5z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiu3u2o231y1w8ubfkp5z.png" alt=" " width="800" height="653"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But here's the practical reality: in production, the agents that stay in production are the ones with solid harnesses. Not the ones with the most creative prompts.&lt;/p&gt;

&lt;p&gt;The teams that deploy reliably aren't just asking "did the agent get the right answer?" They're asking "when it gets the wrong answer, how fast do we know? How do we recover? What's the audit trail? Who can intervene?"&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's the shift from building demos to building systems.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;The autonomous vehicle analogy works because it shifts the conversation from capability to reliability. Nobody debates whether self-driving cars are technically impressive. The debate is always about whether they're trustworthy enough to operate at scale without human supervision.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7y5sa9zcwzqaelascd7z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7y5sa9zcwzqaelascd7z.png" alt=" " width="786" height="634"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's exactly where enterprise AI is right now.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The LLMs are impressive. The question is whether the systems around them are engineering grade.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prompt&lt;/strong&gt;, &lt;strong&gt;Context&lt;/strong&gt;, and &lt;strong&gt;Harness&lt;/strong&gt; Engineering are how you close that gap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>SAGA Made Microservices Reliable. Agent Harness Makes AI Agents Reliable.</title>
      <dc:creator>Seenivasa Ramadurai</dc:creator>
      <pubDate>Sun, 14 Jun 2026 05:31:13 +0000</pubDate>
      <link>https://dev.to/sreeni5018/saga-made-microservices-reliable-agent-harness-makes-ai-agents-reliable-3d1k</link>
      <guid>https://dev.to/sreeni5018/saga-made-microservices-reliable-agent-harness-makes-ai-agents-reliable-3d1k</guid>
      <description>&lt;p&gt;&lt;em&gt;&lt;strong&gt;The distributed systems world solved long-running transactions with SAGA. The agentic AI world has a harder version of the same problem. Here's how Agent Harness answers it.&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwr7baaekdxt8bv6e1px5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fwr7baaekdxt8bv6e1px5.png" alt=" " width="800" height="186"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;I've been deep in agentic AI architecture for a while now &amp;amp; building &lt;strong&gt;Digital Workers&lt;/strong&gt;, designing &lt;strong&gt;multi-agent systems&lt;/strong&gt;, working through the messy production realities of agents that call tools, consult knowledge bases, and loop back on themselves when they're uncertain. And one question keeps coming up when I talk to engineers who come from a microservices background: "Can't we just use SAGA for this?"&lt;/p&gt;

&lt;p&gt;It's a fair question. &lt;strong&gt;SAGA is one of the more elegant patterns in distributed systems&lt;/strong&gt;. And on the surface, agentic workflows look similar enough that the analogy is tempting. Both involve coordinating multi-step processes. Both need state management and failure recovery. Both have to deal with partial completions.&lt;/p&gt;

&lt;p&gt;But the moment you dig into the details, you realize why SAGA alone isn't enough and why Agent Harness exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  What SAGA Was Built to Solve
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;If you've spent time in microservices land&lt;/strong&gt;, you've lived this problem. &lt;strong&gt;Service A completes, Service B completes, Service C fails&lt;/strong&gt; and now you have a half-committed distributed transaction with no clean rollback and no database level guarantee to save you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The SAGA pattern was invented for exactly this.&lt;/strong&gt; The break long-running transactions into a &lt;strong&gt;sequence of local steps&lt;/strong&gt;, and for every step that can succeed, write a compensating action in advance so that if something downstream fails, you can undo the damage cleanly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Figy3ref019ifr4dfaq1n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Figy3ref019ifr4dfaq1n.png" alt=" " width="800" height="660"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It works beautifully because microservices &lt;strong&gt;operate in a deterministic world. Every service has a known API contract&lt;/strong&gt;. Every response has a typed schema. Every failure is a status code or a typed exception. Every retry is predictable. The failure modes are knowable at design time, so you can write compensation logic at design time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI agents don't live in that world.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F756xte0qz9roovct41y2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F756xte0qz9roovct41y2.png" alt=" " width="800" height="658"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem Is Probabilistic, Not Deterministic
&lt;/h2&gt;

&lt;p&gt;Here's what &lt;strong&gt;fundamentally changes&lt;/strong&gt; when you move from &lt;strong&gt;microservices&lt;/strong&gt; to &lt;strong&gt;agentic AI systems&lt;/strong&gt;, your "&lt;strong&gt;services&lt;/strong&gt;" are now &lt;strong&gt;LLM calls&lt;/strong&gt;, &lt;strong&gt;tool invocations&lt;/strong&gt;, &lt;strong&gt;knowledge retrievals&lt;/strong&gt;, &lt;strong&gt;external APIs or MCP Server tool calls **, and **increasingly&lt;/strong&gt;  &lt;strong&gt;human approvals&lt;/strong&gt;. None of these behave like a well defined &lt;strong&gt;REST&lt;/strong&gt; endpoint with a contract you can write compensation logic against.&lt;/p&gt;

&lt;p&gt;An LLM call can return an answer that passes every syntax check but is semantically wrong confidently, fluently, plausibly wrong. A tool call might succeed at the HTTP layer but return data that sends the agent down an entirely incorrect reasoning path. A multi-step task might "complete" having taken three hallucinated intermediate steps before landing somewhere that superficially looks like the goal.&lt;/p&gt;

&lt;p&gt;And here's the part that should give you pause: &lt;strong&gt;a SAGA coordinator would mark all of that as success&lt;/strong&gt;. No exceptions. No compensation triggered. Workflow complete.&lt;/p&gt;

&lt;p&gt;Retrying won't fix it. Compensation logic won't fix it. You need something architecturally different: an &lt;strong&gt;Agent Harness&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What SAGA and Agent Harness Actually Share
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fb7olhfu6sqpym0c1hw3x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fb7olhfu6sqpym0c1hw3x.png" alt=" " width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before getting into where they diverge&lt;/strong&gt;, it's worth being honest about the parallel because &lt;strong&gt;it isn't just a clever analogy. It's structurally real.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both patterns exist to solve the same core problem: coordinating multi-step processes where individual steps can fail, state needs to be preserved across the lifecycle, and the overall system needs to recover gracefully when things go sideways.&lt;/p&gt;

&lt;p&gt;The SAGA Coordinator manages: &lt;strong&gt;state tracking&lt;/strong&gt;, &lt;strong&gt;retries&lt;/strong&gt;, &lt;strong&gt;compensation actions&lt;/strong&gt;, &lt;strong&gt;failure recovery&lt;/strong&gt;, workflow sequencing, and distributed reliability. The Agent Harness manages all of those same things just mapped to a completely different execution model.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;&lt;strong&gt;[The architecture maps cleanly. The implementation is night and day.]&lt;/strong&gt;&lt;/em&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9v973c2uj521fias69f0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9v973c2uj521fias69f0.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Agent Harness Does That SAGA Cannot
&lt;/h2&gt;

&lt;p&gt;SAGA assumes your workflow steps are atomic and deterministic. Agent Harness has to deal with steps that are neither. That's why it needs an entire category of capabilities that have no real SAGA equivalent:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory (Short &amp;amp; Long Term):&lt;/strong&gt; An agent working a multi-turn task needs to remember what it decided three steps ago, what the user said at the start, and what it already tried that didn't work. That's not transaction state. That's episodic memory and working context interleaved in a way that needs to survive tool calls, retries, and mid-task handoffs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Reflection &amp;amp; Critique:&lt;/strong&gt; Before committing to an action or an answer, a well designed harness routes the &lt;strong&gt;agent's proposed output through a&lt;/strong&gt; &lt;strong&gt;self critique step&lt;/strong&gt;. Did the answer actually address the stated goal? Does it contradict something established earlier in the session? Does it fall outside the policy boundaries? SAGA never needs to ask its services whether they feel confident about their output. Agent Harness does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Guardrails &amp;amp; Policies:&lt;/strong&gt; In production especially in regulated industries you &lt;strong&gt;don't want an agent calling a sensitive external API, accessing PII, or making a consequential decision without policy enforcement at the harness level.&lt;/strong&gt; This isn't exception handling after the fact. It's proactive constraint evaluation before execution. I've seen this matter enormously in healthcare projects where the consequences of an unguarded tool call are real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human-in-the-Loop:&lt;/strong&gt;  SAGA runs unattended by design. Agent Harness needs to know when to stop and ask a human and that decision happens at the semantic level, not the infrastructure level. &lt;strong&gt;"I'm not certain this is what the user intended" is a fundamentally different pause condition than "the API returned a 503."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evaluation &amp;amp; Validation:&lt;/strong&gt;  Did the &lt;strong&gt;agent's output actually achieve the goal? Not "did the tool call succeed"&lt;/strong&gt; did we actually do what we set out to do? This requires goal level evaluation, not just a &lt;strong&gt;success/failure&lt;/strong&gt; bit. It's one of the harder things to operationalize in practice, but skipping it is how you ship agents that complete tasks without accomplishing goals.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost &amp;amp; Token Monitoring:&lt;/strong&gt; LLM calls have &lt;strong&gt;variable cost depending on context length, model tier, and how deep the reasoning goes&lt;/strong&gt;. An agent running a complex multi-step task can burn through budget in ways that are invisible until you get the bill. A production Agent Harness needs token spend guardrails the way a microservices platform needs circuit breakers on latency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Durable Execution via Checkpointing:&lt;/strong&gt;  If an &lt;strong&gt;agent task runs for 40 minutes and the process crashes at minute 39, checkpointing lets you resume from the last stable state rather than starting over&lt;/strong&gt;. Philosophically similar to SAGA's compensating transactions but the implementation means serializing agent state, tool call history, memory contents, and intermediate reasoning. Substantially more complex, and substantially more necessary for long horizon tasks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fflchvwqq9l5j9efcf30j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fflchvwqq9l5j9efcf30j.png" alt=" " width="800" height="1274"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Concrete Scenario That Makes This Real
&lt;/h2&gt;

&lt;p&gt;Let me give you a specific example, because abstract architecture arguments only go so far.&lt;/p&gt;

&lt;p&gt;Imagine an agent tasked with: "&lt;strong&gt;Research our top three competitors&lt;/strong&gt;' pricing pages and prepare a comparison summary for the sales team."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A SAGA style system would model this as&lt;/strong&gt;: &lt;strong&gt;call tool to fetch Page A → call tool to fetch Page B → call tool to fetch Page C → call tool to generate summary → done.&lt;/strong&gt; If any fetch fails, compensate. If all fetches succeed, the workflow completes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But here's what can actually happen&lt;/strong&gt;: Page B returns a &lt;strong&gt;cached version from 2 months ago&lt;/strong&gt;. The agent doesn't know that it just sees valid HTML. It processes the outdated pricing as current. The summary it generates is factually wrong in a way that could embarrass your sales team.&lt;/p&gt;

&lt;p&gt;Every step "&lt;strong&gt;succeeded&lt;/strong&gt;." The SAGA coordinator marks it complete. No compensation triggered. &lt;strong&gt;And your sales team walks into a meeting with incorrect competitive data&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent Harness addresses this at multiple layers&lt;/strong&gt;. &lt;strong&gt;Reflection&lt;/strong&gt; &lt;strong&gt;catches&lt;/strong&gt; that the &lt;strong&gt;retrieved content has anomalous&lt;/strong&gt; date markers. Evaluation validates whether the output meets the quality criteria defined for the task. &lt;strong&gt;Guardrails can flag when retrieved content falls below a freshness threshold&lt;/strong&gt;. &lt;strong&gt;Human-in-the-loop&lt;/strong&gt; escalation routes the uncertainty to a person rather than silently proceeding.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's the gap. And it's not a small one.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Key Difference, Plainly Said
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;SAGA&lt;/strong&gt; manages &lt;strong&gt;deterministic workflows&lt;/strong&gt;. &lt;strong&gt;Agent&lt;/strong&gt; Harness manages &lt;strong&gt;probabilistic workflows&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fw9su9o1wxmabms1qlxun.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fw9su9o1wxmabms1qlxun.png" alt=" " width="800" height="581"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;In SAGA, failure modes are knowable at design time&lt;/strong&gt;. You write compensation logic once and trust it to cover the cases. In an Agent Harness, failure can mean: the tool returned a valid response that the agent misread. Or the agent completed every step correctly but arrived at a goal that doesn't satisfy what the user actually wanted. Or the agent is in a soft reasoning loop, &lt;strong&gt;re-checking the same condition because it's genuinely uncertain and nobody told it when to escalate.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Handling that requires reflection, self critique, goal validation, and graceful human escalation none of which exist in the SAGA vocabulary, because SAGA was never designed for an execution unit that reasons about the world.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Means If You're Building Today
&lt;/h2&gt;

&lt;p&gt;If you're designing an agentic system and you're thinking purely in SAGA terms, you're probably building something that's reliable at the infrastructure layer but brittle at the reasoning layer. Your agents will retry correctly. They'll compensate correctly. But they'll also confidently produce wrong answers, hallucinate tool results, and mark tasks complete that aren't — and your coordinator will have no way to know the difference.&lt;/p&gt;

&lt;p&gt;Agent Harness is the layer that closes that gap. It's not a replacement for orchestration. It sits above orchestration and asks: did we actually do the right thing, in the right way, within the right constraints, with the appropriate level of human oversight?&lt;/p&gt;

&lt;p&gt;The engineers who built SAGA were solving a genuinely hard distributed systems problem. The people building Agent Harness today are solving a harder version of it because the failure modes are less visible, the state is messier, and "success" is much harder to define when your execution unit is a language model reasoning about an open-ended goal.&lt;/p&gt;

&lt;p&gt;But the spirit is exactly the same: &lt;strong&gt;build systems that fail gracefully, recover intelligently, and complete what they started&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  SAGA made microservices reliable. Agent Harness is what makes AI agents reliable.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  One Question Worth Sitting With
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Of all the Agent Harness components&lt;/strong&gt;, I've found that &lt;strong&gt;Reflection&lt;/strong&gt; &amp;amp; &lt;strong&gt;Critique&lt;/strong&gt; and &lt;strong&gt;Human-in-the-Loop&lt;/strong&gt; are the &lt;strong&gt;two&lt;/strong&gt; that teams &lt;strong&gt;most consistently underinvest&lt;/strong&gt; in usually because they're harder to wire up than checkpointing or token monitoring, and the cost of skipping them isn't visible until something goes wrong in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which component do you find hardest to implement in practice  and how are you handling it?&lt;/strong&gt; I'm genuinely curious what patterns the community is landing on. Drop it in the comments.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnyzrd8344y3d2nhumj3j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnyzrd8344y3d2nhumj3j.png" alt=" " width="800" height="277"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thanks&lt;br&gt;
Sreeni Ramadorai&lt;/strong&gt;&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
