<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Pratik</title>
    <description>The latest articles on DEV Community by Pratik (@pratik_12b3f8bf3b50e48bae).</description>
    <link>https://dev.to/pratik_12b3f8bf3b50e48bae</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3353910%2F924541cd-88b5-47ef-8933-c8fbec3be59d.png</url>
      <title>DEV Community: Pratik</title>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pratik_12b3f8bf3b50e48bae"/>
    <language>en</language>
    <item>
      <title>How to Build Resilient AI Agents with Search Fallback Loops</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Tue, 06 Oct 2026 17:36:43 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/how-to-build-resilient-ai-agents-with-search-fallback-loops-4gd8</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/how-to-build-resilient-ai-agents-with-search-fallback-loops-4gd8</guid>
      <description>&lt;p&gt;Building autonomous AI agents is incredibly rewarding until you deploy them to production and real-world data breaks your clean pipelines.&lt;br&gt;
A common bottleneck is the tool execution layer. When your agent invokes a vector DB search or a live web API, it assumes it will receive relevant data. But out in the wild, APIs time out, rate limits get hit, and semantic searches frequently return empty arrays.&lt;br&gt;
If your agent treats tool calls as a linear path (Query -&amp;gt; Result -&amp;gt; Next Step), an empty or broken result causes the entire system to collapse or freeze.&lt;br&gt;
The solution is an Agentic Search Fallback Loop. Let's break down how it works and how to build one safely.&lt;br&gt;
The Problem: The Blind Retry Trap&lt;br&gt;
When developers first encounter tool failures in agents, the knee-jerk reaction is to add a simple while loop or a basic retry decorator.&lt;br&gt;
python&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# The Dangerous Way
&lt;/span&gt;&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_search_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use code with caution.&lt;br&gt;
If call_search_tool returns empty because the query keywords are too specific, running it three times changes absolutely nothing. You are simply burning API tokens and increasing latency for the exact same zero-value result.&lt;br&gt;
The Solution: The Strategic Pivot&lt;br&gt;
An Agentic Search Fallback Loop introduces an evaluation step between the failure and the retry. The agent changes its strategy based on why the tool failed.&lt;br&gt;
Architectural Upgrades: Circuit Breakers and Attribution Verification&lt;br&gt;
While a basic fallback loop improves resilience, true production environments require two critical safety mechanisms to avoid cascading failures and hallucinations:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Circuit Breaker Pattern: Simply routing to a backup provider during a systemic failure can cause a "thundering herd" problem, instantly hammering and crashing your backup system during a primary outage. A robust loop must open a circuit breaker, halting calls temporarily once a failure threshold is crossed.&lt;/li&gt;
&lt;li&gt;Attribution Verification: When you relax semantic constraints—such as lowering a vector similarity threshold—you introduce noise. This is exactly where wrong matches slip through, leading the LLM to generate highly confident hallucinations. To counter this, a post-retrieval verification step must validate that the final answer explicitly aligns with the source text.
Here is a full code implementation showing how to orchestrate a fallback loop that changes its internal parameters dynamically while using a circuit breaker and source attribution grading:
python
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;

&lt;span class="c1"&gt;# Simulating a Circuit Breaker State to prevent cascading failures
&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;CircuitBreaker&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;failure_threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failure_threshold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;failure_threshold&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failure_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_open&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record_failure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failure_count&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failure_count&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failure_threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_open&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;🚨 [CIRCUIT BREAKER] Tripped! Halting traffic to protect infrastructure.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record_success&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;failure_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_open&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;

&lt;span class="n"&gt;primary_breaker&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;CircuitBreaker&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Simulating an external search tool that fails on hyper-specific queries
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;mock_vector_search_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;similarity_threshold&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]]:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;primary_breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;is_open&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Circuit breaker is open. Request blocked.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Simulating an empty state for a highly restrictive query
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hyper-specific microservices architecture&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;similarity_threshold&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.75&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="c1"&gt;# Simulating a successful match once constraints relax
&lt;/span&gt;    &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;microservices architecture&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;similarity_threshold&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mf"&gt;0.75&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;primary_breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record_success&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Scalable Microservices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Production deployment strategies for Kubernetes...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="c1"&gt;# Simulating an LLM call that simplifies a failing query
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;llm_query_rewriter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;failed_query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;🔄 [LLM] Rewriting and broadening query: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;failed_query&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hyper-specific&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;failed_query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;microservices architecture&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;failed_query&lt;/span&gt;

&lt;span class="c1"&gt;# Attribution Grader to prevent hallucinations from relaxed queries
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify_source_attribution&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;retrieved_docs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;List&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;🔍 [VERIFIER] Checking if relaxed documents genuinely answer the original query...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# Simple deterministic validation strategy (In production, use a strict micro-LLM prompt)
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;retrieved_docs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;microservices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;doc&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;✅ [VERIFIER] Source attribution verified.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;❌ [VERIFIER] Source attribution failed. Context is irrelevant noise.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;execute_agentic_search_loop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;initial_query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;current_query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;initial_query&lt;/span&gt;
    &lt;span class="n"&gt;similarity_threshold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.85&lt;/span&gt;  &lt;span class="c1"&gt;# Strict initial threshold
&lt;/span&gt;
    &lt;span class="n"&gt;max_retries&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;
    &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;🚀 Starting agentic search for: &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;current_query&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;max_retries&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;retry_count&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;--- Iteration &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;retry_count&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (Threshold: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;similarity_threshold&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;) ---&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="c1"&gt;# Attempt retrieval
&lt;/span&gt;            &lt;span class="n"&gt;results&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;mock_vector_search_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;similarity_threshold&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

            &lt;span class="c1"&gt;# Check for Semantic Failure (Empty Data)
&lt;/span&gt;            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;⚠️ Search returned 0 documents. Initiating fallback logic...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

                &lt;span class="c1"&gt;# Tactic 1: Lower the vector search similarity threshold
&lt;/span&gt;                &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;similarity_threshold&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mf"&gt;0.70&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="n"&gt;similarity_threshold&lt;/span&gt; &lt;span class="o"&gt;-=&lt;/span&gt; &lt;span class="mf"&gt;0.10&lt;/span&gt;
                    &lt;span class="k"&gt;continue&lt;/span&gt;

                &lt;span class="c1"&gt;# Tactic 2: Leverage LLM to reformulate the text query
&lt;/span&gt;                &lt;span class="n"&gt;current_query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;llm_query_rewriter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;

            &lt;span class="c1"&gt;# Verify source attribution before passing data to the generation LLM
&lt;/span&gt;            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="nf"&gt;verify_source_attribution&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;initial_query&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;⚠️ Retrying due to failed attribution verification...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="n"&gt;similarity_threshold&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="mf"&gt;0.05&lt;/span&gt;  &lt;span class="c1"&gt;# Tighten threshold back up
&lt;/span&gt;                &lt;span class="n"&gt;current_query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;llm_query_rewriter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;

            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;✅ Valid data retrieved successfully!&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;success&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;attempts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;retry_count&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="c1"&gt;# Catching Systemic Failure (Network timeouts / API errors)
&lt;/span&gt;            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;💥 Systemic Error encountered: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;primary_breaker&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record_failure&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;🔄 Gracefully degrading to static local cache instead of hammering backup...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;degraded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;title&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Cached Docs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Static fallback context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}]}&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;🛑 Circuit breaker triggered. All fallback strategies exhausted.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[],&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reason&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Max retries reached without relevant matches.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;user_query&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hyper-specific microservices architecture patterns for Kubernetes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="n"&gt;final_output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;execute_agentic_search_loop&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Final Agent Output Summary:&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;final_output&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use code with caution.&lt;br&gt;
Breaking Down the Architecture&lt;br&gt;
This implementation works where simple retry counters fail due to two specific engineering design choices:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Separation of Concerns: The code handles semantic issues (if not results) completely differently from infrastructure failures (except Exception). If an API breaks, it prevents cascading backup infrastructure crashes using a circuit breaker. If the tool works but yields nothing, it isolates the problem to search syntax and adjusts parameters.&lt;/li&gt;
&lt;li&gt;Dynamic State Shift: Each retry uses unique state modifications. The loop alternates between lowering the similarity threshold and calling the query rewriter, maximizing the chance of a successful lookup on successive runs.&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Guardrails Against Hallucination: By running a verification pass on the expanded search space data, the system ensures that looser rules do not result in garbage &lt;br&gt;
context reaching the generation step.&lt;br&gt;
Critical Production Guardrails&lt;br&gt;
To prevent your agentic loops from running amok, you must hardcode deterministic limits directly into your tool-calling framework:&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Strict Iteration Limits: Never allow more than 2 or 3 loop cycles.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Token Budgets: Track the cumulative token usage inside the loop instance; abort immediately if it crosses a pre-set threshold.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Deterministic Safe-Fails: If the final fallback attempt yields nothing, bypass the LLM entirely and return a structured fallback message (e.g., {"status": "no_records_found"}). This prevents the agent from hallucinating an answer.&lt;br&gt;
&lt;strong&gt;The Interview Angle: System Design Focus&lt;/strong&gt;&lt;br&gt;
For engineers interviewing for advanced AI positions, understanding failure states is critical. You might face a system design question like this:&lt;br&gt;
Question: "How do you design a search agent to handle zero-document retrieval states without causing infinite loops or exploding costs?"&lt;br&gt;
Key points for your answer:&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Explain that you avoid raw looping mechanisms because they do not address semantic text mismatches.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Detail a dynamic strategy: if a query fails, your architecture decreases the vector search similarity threshold (e.g., moving cosine similarity from 0.85 to 0.70) or switches from a dense vector search to a keyword BM25 search.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Emphasize the inclusion of an automated circuit breaker to guarantee predictable runtime costs and system safety.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Building loops that adapt to empty or broken states transforms a fragile AI script into an enterprise-grade agent. How do you handle tool degradation in your production environments? Let's discuss in the comments below.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>architecture</category>
      <category>programming</category>
    </item>
    <item>
      <title># 🩺 I Built MediMind AI for Someone I Care About — An Open-Model Health Companion</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Sun, 04 Oct 2026 17:39:26 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/-i-built-medimind-ai-for-someone-i-care-about-an-open-model-health-companion-1dmg</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/-i-built-medimind-ai-for-someone-i-care-about-an-open-model-health-companion-1dmg</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;I built &lt;strong&gt;MediMind AI&lt;/strong&gt;, a personalized health companion designed to make medical information easier to understand.&lt;/p&gt;

&lt;p&gt;I built it for &lt;strong&gt;a real person close to me&lt;/strong&gt; who found it difficult to understand medical information and turn reports, symptoms, and medication information into clear next steps.&lt;/p&gt;

&lt;p&gt;Instead of creating another generic AI chatbot, I wanted to build something practical: a health-information assistant that combines personalization, retrieval, citations, and safety checks.&lt;/p&gt;

&lt;p&gt;MediMind includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI Health Assistant&lt;/li&gt;
&lt;li&gt;Medical report analysis&lt;/li&gt;
&lt;li&gt;Symptom analysis&lt;/li&gt;
&lt;li&gt;Medication information&lt;/li&gt;
&lt;li&gt;Medication reminders&lt;/li&gt;
&lt;li&gt;RAG-based medical knowledge retrieval&lt;/li&gt;
&lt;li&gt;Source and citation support&lt;/li&gt;
&lt;li&gt;Emergency/safety triage&lt;/li&gt;
&lt;li&gt;Health history and vitals&lt;/li&gt;
&lt;li&gt;Personalized care context&lt;/li&gt;
&lt;li&gt;Saved reports and health information&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The goal is not to diagnose someone. The goal is to make medical information easier to understand and help people prepare better questions for healthcare professionals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;🌐 &lt;strong&gt;Live Demo:&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://meed-mind-ai.vercel.app" rel="noopener noreferrer"&gt;https://meed-mind-ai.vercel.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The main demo flow is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Create Care Profile → Ask a health question → Retrieve relevant medical context → Generate an explanation using an open-weight model → Show supporting sources → Apply safety checks&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;💻 &lt;strong&gt;GitHub:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/pratikraut8889-max" rel="noopener noreferrer"&gt;
        pratikraut8889-max
      &lt;/a&gt; / &lt;a href="https://github.com/pratikraut8889-max/MeedMind-AI" rel="noopener noreferrer"&gt;
        MeedMind-AI
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      An AI-powered second opinion and accessibility tool helping patients understand medical reports in any language.
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;MediMind AI - Production Retrieval-Augmented Generation (RAG) Architecture&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;MediMind AI implements a production-quality, real clinical retrieval layer designed to ensure every medical answer is grounded in authoritative, peer-reviewed clinical guidelines, and that no sources or citations are ever fabricated.&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;1. End-to-End RAG Architecture&lt;/h2&gt;
&lt;/div&gt;
&lt;div class="snippet-clipboard-content notranslate position-relative overflow-auto"&gt;
&lt;pre class="notranslate"&gt;&lt;code&gt;User Question
      │
      ▼
1. Query Processing
   ├── Text Normalization &amp;amp; PII Sanitization
   ├── Medical Taxonomy &amp;amp; Synonym Expansion (AHA, CDC, NIH terms)
   ├── Intent Classification (emergency, diagnostic, medication, lifestyle)
   └── Dense Embedding Generation (gemini-embedding-2-preview)
      │
      ▼
2. Pluggable Medical Retriever (IMedicalRetriever)
   ├── Semantic Cosine Similarity (Dense Vector Search)
   ├── Lexical &amp;amp; Clinical Tag Overlap Matching
   └── Hybrid Scored Ranking: 0.65 * denseSimilarity + 0.35 * lexicalScore
      │
      ▼
3. Relevant Medical Documents
   └── Scored &amp;amp; Ranked Chunks with Document Identifiers &amp;amp; Metadata
      │
      ▼
4. Context Filtering Layer
   ├── Strict Relevance Thresholding (default minRelevanceScore = 0.58)
   ├── Low-Quality / Irrelevant Query Suppression
   ├── Token-Budget Aware Deduplication&lt;/code&gt;&lt;/pre&gt;…&lt;/div&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/pratikraut8889-max/MeedMind-AI" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;MediMind is built with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;React&lt;/li&gt;
&lt;li&gt;TypeScript&lt;/li&gt;
&lt;li&gt;Vite&lt;/li&gt;
&lt;li&gt;Express&lt;/li&gt;
&lt;li&gt;Firebase / Firestore&lt;/li&gt;
&lt;li&gt;RAG-based retrieval&lt;/li&gt;
&lt;li&gt;Open-weight AI&lt;/li&gt;
&lt;li&gt;Open embeddings&lt;/li&gt;
&lt;li&gt;Deterministic medical safety and triage logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The architecture is designed so that the AI model is not tightly coupled to a single vendor.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
  ↓
React UI
  ↓
Express API
  ↓
Safety / Triage
  ↓
RAG Retrieval
  ↓
Medical Context
  ↓
Open-Weight Model
  ↓
Structured Output Validation
  ↓
Citation Validation
  ↓
Safe Response
  ↓
User
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI is surrounded by safety and validation layers rather than being treated as an unquestionable medical authority.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Open Innovation Matters
&lt;/h2&gt;

&lt;p&gt;This project is especially interesting to me because the AI layer is built around an &lt;strong&gt;open-weight model&lt;/strong&gt; rather than making the entire application dependent on a single closed AI API.&lt;/p&gt;

&lt;p&gt;Using an open model gives the project a more flexible architecture.&lt;/p&gt;

&lt;p&gt;The model can be changed without rebuilding the whole application. Different inference environments can be evaluated, the AI layer can be adapted for different deployments, and developers have more control over how the model is integrated into the product.&lt;/p&gt;

&lt;p&gt;For a health-focused application, this flexibility matters because health information can be sensitive.&lt;/p&gt;

&lt;p&gt;MediMind also combines the model with retrieval and citation validation so that the AI is grounded in relevant medical information instead of simply being treated as the source of truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  The “Build for a Friend” Part
&lt;/h2&gt;

&lt;p&gt;The idea started from a simple problem: &lt;strong&gt;medical information is often difficult for non-medical people to understand.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I wanted to turn that problem into something useful for someone close to me rather than building an AI demo without a real user in mind.&lt;/p&gt;

&lt;p&gt;The project therefore focuses on making reports, symptoms, medications, and health information easier to understand while keeping medical safety at the center.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The person I built it for:&lt;/strong&gt;&lt;br&gt;
[Write their real relationship here, e.g. “my cousin”]&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The problem they actually have:&lt;/strong&gt;&lt;br&gt;
[Write their real problem here in one sentence]&lt;/p&gt;

&lt;p&gt;I intentionally kept the application focused on explanation and decision support rather than presenting it as a replacement for a doctor.&lt;/p&gt;
&lt;h2&gt;
  
  
  Safety
&lt;/h2&gt;

&lt;p&gt;MediMind is &lt;strong&gt;not a doctor and does not replace professional medical care&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The application includes a dedicated safety/triage layer for potentially urgent situations. High-risk inputs can be handled before the normal conversational response flow.&lt;/p&gt;

&lt;p&gt;The system is designed to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;recognize potentially urgent situations&lt;/li&gt;
&lt;li&gt;avoid presenting AI output as a diagnosis&lt;/li&gt;
&lt;li&gt;ground responses using retrieved information&lt;/li&gt;
&lt;li&gt;validate citations&lt;/li&gt;
&lt;li&gt;encourage professional medical care when appropriate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This project is a health-information and decision-support tool, not a medical device or substitute for a licensed healthcare professional.&lt;/p&gt;
&lt;h2&gt;
  
  
  What Open AI Made Possible
&lt;/h2&gt;

&lt;p&gt;The biggest architectural benefit of using an open-weight model is that MediMind is not built around one permanent AI vendor.&lt;/p&gt;

&lt;p&gt;The application can evolve from one open model to another without rewriting the entire product:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Open Model A
     ↓
Provider Abstraction
     ↓
   MediMind
     ↑
Open Model B
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This makes experimentation, model replacement, and future deployment options much easier.&lt;/p&gt;

&lt;p&gt;It also makes the AI layer a replaceable component rather than a hard-coded dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;I used AI-assisted development while building the project and focused on keeping the important decisions, safety logic, and application architecture understandable and reproducible.&lt;/p&gt;

&lt;p&gt;[Add your DevRelay / agent session link here if you have one.]&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Category
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Best Use of Gemma&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;MediMind uses an open-weight model as its core AI layer and combines it with RAG, personalization, and safety validation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;The current version focuses on making health information more understandable for a person close to me.&lt;/p&gt;

&lt;p&gt;The longer-term goal is to make MediMind a privacy-conscious, customizable health companion where the AI model, retrieval system, and safety layer can continue to evolve independently.&lt;/p&gt;




</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
    <item>
      <title>Research Dossier — Vibe-Coding a Case File, Then Teaching It to Hold Its Own Workflow</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Sat, 03 Oct 2026 17:47:13 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-vibe-coding-a-case-file-then-teaching-it-to-hold-its-own-workflow-4152</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-vibe-coding-a-case-file-then-teaching-it-to-hold-its-own-workflow-4152</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path Two: Vibe-Code Something Strange&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A note up front: this extends the same codebase as my &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2"&gt;Path One submission&lt;/a&gt; — one project, two angles. Path One is about the agent and the structured content it reasons over. This post is about the other half: how the app itself got built almost entirely through prompting, and what it took to push past a read-only frontend into something with a real Sanity-backed workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Research Dossier is a Next.js app, Sanity behind it, built through iterative conversational prompting rather than hand-writing the UI from scratch. The interface is deliberately not a blog-template look: a two-pane "case file" layout — a terminal-style live log on the left showing agent progress, a paper-dossier report on the right with a hand-stamped "CONTESTED" mark when sources disagree, custom ink/paper/gold color tokens, serif body type paired with monospace for the log — none of it from a component library's defaults.&lt;/p&gt;

&lt;p&gt;For this path specifically, I pushed further than the Path One version: the system now models its own approval workflow as real Sanity content, not just UI state.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Live app: &lt;a href="https://multi-agent-research-analyst.vercel.app/" rel="noopener noreferrer"&gt;https://multi-agent-research-analyst.vercel.app/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ask a question, watch the agents work in the case log, review the draft, click approve — then visit &lt;code&gt;/cases&lt;/code&gt; to see the same investigation listed with a status badge that flips from &lt;code&gt;draft&lt;/code&gt; to &lt;code&gt;approved&lt;/code&gt;. Open Sanity Studio and you'll find the same document, same transition, visible there too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/pratikdevelop/multi-agent-research-analyst" rel="noopener noreferrer"&gt;https://github.com/pratikdevelop/multi-agent-research-analyst&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  My Build Process
&lt;/h2&gt;

&lt;p&gt;This was built in long back-and-forth passes rather than one clean generation, and most of the real work was debugging what got generated, not writing it from scratch. A rough timeline of what actually broke and got fixed along the way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;State/node name collision&lt;/strong&gt; — LangGraph refused to compile because a node was named the same as a state channel (&lt;code&gt;research&lt;/code&gt;, &lt;code&gt;analysis&lt;/code&gt;). Not obvious from the error message's framing at first; fixed by renaming nodes to &lt;code&gt;research_agent&lt;/code&gt;, &lt;code&gt;analysis_agent&lt;/code&gt;, etc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model availability churn&lt;/strong&gt; — the first two Groq model names I tried (&lt;code&gt;llama-3.3-70b-versatile&lt;/code&gt;, then &lt;code&gt;llama-3.1-8b-instant&lt;/code&gt;) weren't actually available on my account's model list, discovered only by hitting &lt;code&gt;/v1/models&lt;/code&gt; directly and reading what my key could actually see. Landed on &lt;code&gt;openai/gpt-oss-20b&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A real empty-research bug&lt;/strong&gt; — the agent's tool-call loop wasn't feeding tool results back to the model, so research came back blank and every downstream agent had nothing to work with. Had to rewrite the loop to actually execute tool calls and continue the conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A Sanity auth error that turned out to be a project-ID/org-ID mixup&lt;/strong&gt; — "Session does not match project host" traced back to the request URL hitting the org ID instead of the project ID, both of which were sitting right next to each other on the same dashboard screen.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A null-crash in production&lt;/strong&gt; — &lt;code&gt;contradicts[]-&amp;gt;{...}&lt;/code&gt; returns &lt;code&gt;null&lt;/code&gt; in GROQ when the field was never set, not &lt;code&gt;[]&lt;/code&gt;. Crashed the Knowledge Base browser page until guarded.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vercel deploy failures, three different ones in a row&lt;/strong&gt; — a stray &lt;code&gt;output: 'export'&lt;/code&gt; in &lt;code&gt;next.config.js&lt;/code&gt; that forced static export and broke every API route; a peer-dependency conflict needing &lt;code&gt;--legacy-peer-deps&lt;/code&gt; set via &lt;code&gt;.npmrc&lt;/code&gt; AND the Vercel install-command override, since one alone didn't stick; then a stale production deployment pointing at an old build while the working one sat under "Preview."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The MCP swap&lt;/strong&gt; — once Sanity Context was enabled on my org, swapping the research agent from a direct GROQ query to the real Context MCP endpoint meant discovering the actual tool name it exposes (&lt;code&gt;knowledge_base_read&lt;/code&gt;) and that it serves a &lt;code&gt;/initial-context&lt;/code&gt; grounding endpoint worth injecting into the system prompt — neither of which is obvious without just connecting and inspecting what comes back. Kept the GROQ version as an automatic fallback rather than a hard cutover.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Workflow feature, for this path specifically&lt;/strong&gt; — added a &lt;code&gt;caseReport&lt;/code&gt; document type with a &lt;code&gt;status&lt;/code&gt; field (&lt;code&gt;draft&lt;/code&gt; → &lt;code&gt;approved&lt;/code&gt;), a write-scoped Sanity client kept deliberately separate from the read-only one used for research, and two API routes: one that persists the agent's draft the moment it's ready, one that patches the same document to &lt;code&gt;approved&lt;/code&gt; when a human closes the case. The &lt;code&gt;/cases&lt;/code&gt; page and Sanity Studio both show the same transition, because it's the same document.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How I Used Sanity — Workflow
&lt;/h2&gt;

&lt;p&gt;Three document types: &lt;code&gt;topic&lt;/code&gt;, &lt;code&gt;source&lt;/code&gt;, &lt;code&gt;claim&lt;/code&gt; (the Path One side), plus &lt;code&gt;caseReport&lt;/code&gt; for this path — &lt;code&gt;question&lt;/code&gt;, &lt;code&gt;draftText&lt;/code&gt;, &lt;code&gt;finalText&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt;, &lt;code&gt;createdAt&lt;/code&gt;, &lt;code&gt;approvedAt&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The workflow itself is two writes against one document: the research/writing/review pipeline creates it with &lt;code&gt;status: "draft"&lt;/code&gt;, and a human approving (or editing, then approving) in the UI patches the same document to &lt;code&gt;status: "approved"&lt;/code&gt; with the final text and a timestamp. No separate "approval" document, no external state machine — the content &lt;em&gt;is&lt;/em&gt; the workflow state, which is the same philosophy as the &lt;code&gt;contradicts&lt;/code&gt; field in the Path One schema: make the thing you care about tracking an explicit field, not something inferred.&lt;/p&gt;

&lt;p&gt;I didn't build an App SDK component for this submission — Workflows was the deeper of the two bonuses I had time to do properly, and the challenge notes doing one well beats doing both shallowly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;p&gt;Project ID: &lt;code&gt;4kagnnrl&lt;/code&gt;&lt;br&gt;
&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Stop Overpaying for APIs: When to Swap Your Cloud LLM for a Local SLM 🛠️</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Sat, 03 Oct 2026 07:24:20 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/stop-overpaying-for-apis-when-to-swap-your-cloud-llm-for-a-local-slm-2n67</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/stop-overpaying-for-apis-when-to-swap-your-cloud-llm-for-a-local-slm-2n67</guid>
      <description>&lt;p&gt;Let's face it: using an enterprise cloud LLM API to parse basic JSON, route support tickets, or clean up markdown is massive overkill. It's slow, expensive, and leaves your app vulnerable to third-party downtime.&lt;br&gt;
If you haven't looked at Small Language Models (SLMs) recently, it's time to check them out.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;+-------------------+-------------------------+-------------------------+

| Feature           | Cloud LLM               | Local SLM (&amp;lt;15B)        |
+-------------------+-------------------------+-------------------------+

| Deployment        | Cloud API Only          | Local, Edge, On-Prem    |
| Latency           | High (Network bound)    | Low (Local hardware)    |
| Data Privacy      | Third-party risk        | 100% Secure / Offline   |
| Cost Structure    | Pay-per-token           | Fixed Compute / Free    |
+-------------------+-------------------------+-------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  🧠 The Developer's Playbook: Where to Draw the Line
&lt;/h2&gt;

&lt;p&gt;🟩 When to stay with an LLM:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;You need deep, multi-step zero-shot reasoning.&lt;/li&gt;
&lt;li&gt;You are generating complex, multi-file code structures.&lt;/li&gt;
&lt;li&gt;You need massive, 100k+ token context windows.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  🚀 When to drop in an SLM:
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt; You are building specialized AI agents with fixed, repeatable tools.&lt;/li&gt;
&lt;li&gt; You need real-time, low-latency performance on edge devices or mobile.&lt;/li&gt;
&lt;li&gt; You handle sensitive user text / PII that cannot legally leave your server.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  🛠️ Setting It Up Locally: Asynchronous Log Parsing with FastAPI
&lt;/h2&gt;

&lt;p&gt;With ecosystem tools like Ollama, vLLM, and LangChain, spinning up a local SLM (like Llama-3-8B or Phi-3) takes minimal configuration.&lt;/p&gt;

&lt;p&gt;Instead of a basic script, let's build a production-ready asynchronous FastAPI endpoint. It consumes raw streaming application log entries, extracts entities using structured Pydantic schemas, and outputs clean JSON entirely offline.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# pip install fastapi uvicorn langchain-ollama langchain-core pydantic
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;uvicorn&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;HTTPException&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Field&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_ollama&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OllamaLLM&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_core.prompts&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ChatPromptTemplate&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_core.output_parsers&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;JsonOutputParser&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Local SLM Inference Gateway&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;*&lt;em&gt;1. Define input contract and expected structured output schema&lt;br&gt;
*&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;LogPayload&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;raw_log&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(...,&lt;/span&gt; &lt;span class="n"&gt;example&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;[ERROR] auth_service: JWT verification failed - Signature expired&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;LogAnalysis&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Must be exactly &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;SUCCESS&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;WARN&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, or &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;ERROR&lt;/span&gt;&lt;span class="sh"&gt;'"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;anomaly_detected&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;True if unexpected or malicious behavior is found&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;summary&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Field&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;A concise 1-sentence engineering breakdown of the issue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;2. Initialize local SLM (Requires Ollama running locally with the target model). Setting temperature=0.0 ensures highly deterministic JSON structures&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;local_slm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OllamaLLM&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llama3:8b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Warning: Ensure Ollama is running locally. Error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;*&lt;em&gt;3. Formulate strict extraction prompt instructions&lt;br&gt;
*&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ChatPromptTemplate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_template&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a specialized security agent. Analyze the following application log snippet. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Extract information matching the structural requirements schema.&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="s"&gt;Log: {log_entry}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;*&lt;em&gt;4. Chain components together using LCEL (LangChain Expression Language)&lt;br&gt;
*&lt;/em&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;log_chain&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="n"&gt;local_slm&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="nc"&gt;JsonOutputParser&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pydantic_object&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;LogAnalysis&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/v1/analyze-log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;LogAnalysis&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;analyze_application_log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;LogPayload&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
  &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Asynchronously swallows raw streaming logs, routes them to the local 
    SLM core loop, and yields structured JSON insights with near-zero latency.
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Await chain execution inside FastAPI's async execution loop
&lt;/span&gt;        &lt;span class="n"&gt;structured_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;log_chain&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;ainvoke&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;log_entry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;raw_log&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;structured_response&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;HTTPException&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SLM Engine Inference Failure: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;uvicorn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;host&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0.0.0.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;port&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your network latency drops to the floor, your third-party API billing statement hits exactly zero, and your monitoring microservice runs securely behind air-gapped on-prem environments.&lt;/p&gt;

&lt;p&gt;What's your go-to local model right now? Are you team Llama, Mistral, or running something even lighter on the edge? Drop your stack and your token-per-second benchmarks below! 👇&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>fastapi</category>
    </item>
    <item>
      <title>The Rise of AI Coding Agents: From Code Completion to Autonomous Software Engineering</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Fri, 02 Oct 2026 04:17:38 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/the-rise-of-ai-coding-agents-from-code-completion-to-autonomous-software-engineering-283e</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/the-rise-of-ai-coding-agents-from-code-completion-to-autonomous-software-engineering-283e</guid>
      <description>&lt;p&gt;Software development is entering a new phase.&lt;br&gt;
For years, AI coding tools were primarily focused on autocomplete, code generation, and answering programming questions. Today, products such as OpenAI Codex, Anthropic's Claude Code, and Google's Antigravity ecosystem are increasingly designed around a different idea: give an AI a software task and let it plan, modify files, use tools, run code, test its work, and iterate.&lt;br&gt;
The interesting shift is not simply that AI can write more code.&lt;br&gt;
It is that AI is becoming an active participant in the software development workflow.&lt;br&gt;
From Copilot to Coding&amp;nbsp;Agent&lt;br&gt;
The traditional AI coding workflow looks something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Developer → Prompt → Code Suggestion → Developer Review

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Agentic coding changes the loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Developer → Goal → Agent → Plan → Code → Run → Test → Iterate → Result

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That difference matters.&lt;br&gt;
A coding agent needs to understand the repository, maintain context across multiple steps, interact with development tools, and decide what to do next rather than simply generate a code snippet.&lt;br&gt;
OpenAI's current Codex product, for example, is designed to handle engineering tasks such as features, refactors, migrations, and other repository-level work, including parallel work across environments.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;OpenAI Codex: Moving Toward End-to-End Development
OpenAI has continued expanding Codex beyond individual coding requests.
With GPT-5.3-Codex, OpenAI described Codex as having stronger agentic capabilities across coding, web development, and broader computer-based work.
The Codex app also introduced workflows for managing multiple agents and running work in parallel, while later updates added features such as Goal mode, improved browser interactions, and longer-running workflows.
The important change is the abstraction level.
Instead of asking:
"Write this function."
Developers can increasingly ask:
"Implement this feature, run the tests, fix the failures, and prepare the change."
That is much closer to delegating a software task than generating code.&lt;/li&gt;
&lt;li&gt;Claude Code: The Terminal Becomes an Agent Workspace&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Anthropic has taken a similar approach with Claude Code.&lt;br&gt;
Claude Code is built around working directly with repositories, terminals, development tools, and longer-running tasks. Anthropic has continued expanding its autonomy and development capabilities, including support for longer sessions and autonomous operation.&lt;br&gt;
Anthropic has also introduced Claude Code Security, which can scan codebases for vulnerabilities and suggest targeted patches for human review.&lt;br&gt;
This is an important development because the coding agent is becoming involved not only in writing code, but also in reviewing and improving software quality.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Google: Agent-First Development with Antigravity
Google is approaching the same problem through its Antigravity development platform.
Google describes Antigravity as an agent-first development environment where agents can plan, execute, and verify complex tasks across the editor, terminal, and browser.
Google's 2026 developer updates expanded this approach with Antigravity 2.0 and Antigravity CLI, including support for specialized subagents and built-in controls such as terminal sandboxing, credential masking, and hardened Git policies.
Google also released Gemini 3.7 Flash, positioning it specifically for coding and agent workflows, with improvements in debugging, issue resolution, web development, and long-horizon engineering tasks.
The direction is clear: the IDE is becoming less of a place where humans manually write every line and more of a workspace where humans supervise AI-driven development workflows.&lt;/li&gt;
&lt;li&gt;The Real Innovation Is the Agent&amp;nbsp;Loop&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The most important part of these tools isn't raw code generation.&lt;br&gt;
It is the loop around generation:&lt;br&gt;
Understand → Plan → Execute → Observe → Test → Correct → Repeat&lt;br&gt;
A useful coding agent must be able to answer questions such as:&lt;br&gt;
What files are relevant?&lt;br&gt;
What is the existing architecture?&lt;br&gt;
What dependencies are required?&lt;br&gt;
Did the implementation actually work?&lt;br&gt;
Which tests failed?&lt;br&gt;
What should be changed next?&lt;/p&gt;

&lt;p&gt;This is why context, tool use, execution environments, testing, and memory are becoming as important as the underlying model.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Software Engineering Is Becoming More Supervisory&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This does not necessarily mean developers disappear.&lt;br&gt;
Instead, the developer's role can move higher up the abstraction stack.&lt;br&gt;
Rather than spending all of their time writing individual functions, developers may increasingly spend more time:&lt;br&gt;
Defining requirements → Designing architecture → Setting constraints → Reviewing agent output → Validating behavior&lt;br&gt;
That makes engineering judgment even more important.&lt;br&gt;
The question becomes not only:&lt;br&gt;
"Can the AI write this?"&lt;br&gt;
but also:&lt;br&gt;
"Did it build the right thing, within the right constraints?"&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The Production Challenge&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;As coding agents become more autonomous, new engineering challenges appear.&lt;br&gt;
How do you control what an agent can access?&lt;br&gt;
How do you prevent destructive commands?&lt;br&gt;
How do you track every action?&lt;br&gt;
How do you review changes created by multiple agents?&lt;br&gt;
How do you manage secrets and credentials?&lt;br&gt;
How do you know when an agent is stuck in a loop?&lt;br&gt;
OpenAI's own documentation on running Codex safely highlights the need for technical boundaries, approval requirements, access controls, and telemetry when deploying coding agents in real workflows.&lt;br&gt;
So the future of AI-assisted development isn't only about better models.&lt;br&gt;
It is also about better agent infrastructure.&lt;br&gt;
The Bigger&amp;nbsp;Picture&lt;br&gt;
The evolution looks something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Autocomplete
   ↓
AI Coding Assistant
   ↓
Repository-Aware Coding Agent
   ↓
Multi-Agent Development
   ↓
Autonomous Software Engineering Workflows

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This changes the economics and workflow of software development, but it also introduces a new layer of engineering complexity.&lt;br&gt;
The goal isn't simply to build an agent that can write code.&lt;br&gt;
The goal is to build a system that can reliably understand a task, execute it, verify the result, recover from failure, and remain within clearly defined boundaries.&lt;br&gt;
That is where AI coding is heading.&lt;br&gt;
And the next big question isn't:&lt;br&gt;
"Can AI write software?"&lt;br&gt;
It is:&lt;br&gt;
"How much of the software development lifecycle can an AI agent reliably own?"&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>machinelearning</category>
      <category>software</category>
    </item>
    <item>
      <title>🚀 OpenAI’s Agent Era: Agents API, Dots, GPT-6 Sol &amp; More</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Thu, 01 Oct 2026 19:15:44 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/openais-agent-era-agents-api-dots-gpt-6-sol-more-482a</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/openais-agent-era-agents-api-dots-gpt-6-sol-more-482a</guid>
      <description>&lt;p&gt;AI development is rapidly moving from chat-based assistants to autonomous agent systems.&lt;/p&gt;

&lt;p&gt;OpenAI's latest announcements provide a good example of this shift.&lt;/p&gt;

&lt;p&gt;At DevDay 2026, OpenAI announced more than 20 updates covering models, ChatGPT, Codex, APIs, and developer tools.&lt;/p&gt;

&lt;p&gt;Here are the major developments developers should know about.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Agents API 🤖&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Agents API is now available in public beta.&lt;/p&gt;

&lt;p&gt;It allows developers to build and run cloud agents using infrastructure designed for long-running agent workflows.&lt;/p&gt;

&lt;p&gt;Key capabilities include:&lt;/p&gt;

&lt;p&gt;Tool usage&lt;br&gt;
File handling&lt;br&gt;
Code execution&lt;br&gt;
Context management&lt;br&gt;
Subagents&lt;br&gt;
Long-running execution&lt;/p&gt;

&lt;p&gt;OpenAI says the infrastructure is based on the harness used to power Codex.&lt;/p&gt;

&lt;p&gt;The interesting part is that developers don't necessarily need to build the entire runtime layer themselves.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Dots: Always-On Agents 🔄&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;OpenAI introduced Dots, which it describes as always-on agents designed to work continuously on users' behalf.&lt;/p&gt;

&lt;p&gt;This introduces a different interaction model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Traditional AI

User → Prompt → AI → Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;versus:&lt;/p&gt;

&lt;p&gt;Agentic AI&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User → Goal
          ↓
       Agent
          ↓
   Tools + Context
          ↓
      Actions
          ↓
     Result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second model is much closer to software automation than traditional conversational AI.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;GPT-6 Sol and Luna 🧠&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;OpenAI has also introduced GPT-6 Sol and Luna, offering different balances of capability and cost.&lt;/p&gt;

&lt;p&gt;For developers, model selection increasingly becomes an architecture decision.&lt;/p&gt;

&lt;p&gt;The question isn't simply:&lt;/p&gt;

&lt;p&gt;Which model is smartest?&lt;/p&gt;

&lt;p&gt;It becomes:&lt;/p&gt;

&lt;p&gt;Which model provides the right capability, speed, and cost for this task?&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Codex + Developer Tooling 💻&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;OpenAI continues expanding Codex and its developer platform, with DevDay announcements covering improvements across coding workflows, APIs, and agent development.&lt;/p&gt;

&lt;p&gt;This points toward AI coding systems that can handle increasingly complete development workflows rather than only generating individual snippets.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What Developers Should Watch 🔍&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The emerging AI application stack increasingly looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
             AI Models
                 ↓
            Agent Runtime
                 ↓
       Tools + Code Execution
                 ↓
        Context + Memory
                 ↓
       Security + Permissions
                 ↓
       Monitoring + Evaluation
                 ↓
            Production

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difficult part of AI development is increasingly shifting from model access to reliable orchestration.&lt;/p&gt;

&lt;p&gt;Conclusion&lt;/p&gt;

&lt;p&gt;OpenAI's recent releases show how quickly agent infrastructure is developing.&lt;/p&gt;

&lt;p&gt;We're moving from:&lt;/p&gt;

&lt;p&gt;“Ask AI a question.”&lt;/p&gt;

&lt;p&gt;toward:&lt;/p&gt;

&lt;p&gt;“Give AI a goal and let it work.”&lt;/p&gt;

&lt;p&gt;That shift could have significant implications for software developers, startups, and enterprise applications.&lt;/p&gt;

&lt;p&gt;The key engineering challenge now is not just making agents capable.&lt;/p&gt;

&lt;p&gt;It's making them reliable, controllable, secure, and useful in production.&lt;/p&gt;

&lt;p&gt;What are you building with AI agents right now?&lt;/p&gt;

&lt;h1&gt;
  
  
  OpenAI #AIAgents #DevDay2026 #AgentsAPI #GPT6 #Codex #AI #Programming #WebDevelopment
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>openai</category>
      <category>agents</category>
    </item>
    <item>
      <title>🛠️ LangChain’s Latest Agent Infrastructure: Engine v2, Managed Deep Agents, SmithDB &amp; More</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Thu, 01 Oct 2026 19:09:27 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/langchains-latest-agent-infrastructure-engine-v2-managed-deep-agents-smithdb-more-5fab</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/langchains-latest-agent-infrastructure-engine-v2-managed-deep-agents-smithdb-more-5fab</guid>
      <description>&lt;p&gt;Building an AI agent is getting easier.&lt;/p&gt;

&lt;p&gt;Running one reliably in production is still the hard part.&lt;/p&gt;

&lt;p&gt;At Interrupt 2026, LangChain introduced a major set of updates focused on production agent infrastructure. Since then, the platform has expanded further with LangSmith Engine v2, public beta access for Managed Deep Agents, and additional governance capabilities.&lt;/p&gt;

&lt;p&gt;Here’s the current stack developers should know about.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;LangSmith Engine v2 🚀&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The original LangSmith Engine was introduced at Interrupt 2026 to help teams analyze production traces, identify recurring failures, diagnose root causes, and propose fixes.&lt;/p&gt;

&lt;p&gt;Engine v2, released in September 2026, goes further:&lt;/p&gt;

&lt;p&gt;🔴 Proactively red-teams agents to identify potential issues before they appear in production.&lt;br&gt;
📉 Detects inefficient agent trajectories and trends in error rate, latency, and cost.&lt;br&gt;
🧪 Automatically tests proposed prompt and code fixes before human review.&lt;br&gt;
🔄 Reproduces failures, tests fixes, evaluates the results, and then surfaces validated changes for review.&lt;/p&gt;

&lt;p&gt;That makes the development loop much more automated:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;gt; Production traces
&amp;gt;       ↓
&amp;gt; Issue detection
&amp;gt;       ↓
&amp;gt; Root-cause analysis
&amp;gt;       ↓
&amp;gt; Proposed fix
&amp;gt;       ↓
&amp;gt; Automated validation
&amp;gt;       ↓
&amp;gt; Human review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;SmithDB ⚡&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Agent traces are becoming much larger and more complex than traditional application logs.&lt;/p&gt;

&lt;p&gt;SmithDB is LangChain’s purpose-built database infrastructure for agent observability. It is designed for workloads involving deeply nested spans, long-running operations, full-text search, JSON filtering, and trace reconstruction.&lt;/p&gt;

&lt;p&gt;LangChain reports up to 15× faster performance on core LangSmith workloads compared with its previous infrastructure.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Managed Deep Agents 🤖&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Running long-lived autonomous agents yourself means managing:&lt;/p&gt;

&lt;p&gt;Runtime infrastructure&lt;br&gt;
Persistence&lt;br&gt;
Checkpointing&lt;br&gt;
Memory&lt;br&gt;
Tool execution&lt;br&gt;
Sandboxes&lt;br&gt;
Streaming&lt;br&gt;
Observability&lt;/p&gt;

&lt;p&gt;Managed Deep Agents moves much of that operational layer into a hosted LangSmith runtime.&lt;/p&gt;

&lt;p&gt;Developers can define agents using the Deep Agents framework and deploy them without building and maintaining their own agent server infrastructure. Managed Deep Agents became available in public beta in August 2026.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Production Sandboxes 🔒&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Giving an AI agent the ability to execute code introduces obvious security concerns.&lt;/p&gt;

&lt;p&gt;LangSmith Sandboxes, which reached GA at Interrupt 2026, provide isolated execution environments for workloads involving generated code, shell commands, files, dependencies, and data analysis.&lt;/p&gt;

&lt;p&gt;The environments use hardware-virtualized microVMs for stronger isolation from your main infrastructure.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Enterprise Governance 📊&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Production agents also need controls around cost, data, permissions, and context.&lt;/p&gt;

&lt;p&gt;LLM Gateway&lt;/p&gt;

&lt;p&gt;LangSmith’s LLM Gateway sits between agents and model providers and provides controls such as:&lt;/p&gt;

&lt;p&gt;Spend limits&lt;br&gt;
Rate limits&lt;br&gt;
Model fallbacks&lt;br&gt;
PII and secret redaction&lt;br&gt;
Centralized usage visibility&lt;/p&gt;

&lt;p&gt;The Gateway entered public beta in July 2026.&lt;/p&gt;

&lt;p&gt;Context Hub&lt;/p&gt;

&lt;p&gt;Context Hub provides a centralized way to manage and version the files and information that shape agent behavior, including instructions, policies, examples, and skills.&lt;/p&gt;

&lt;p&gt;Messages View&lt;/p&gt;

&lt;p&gt;For complex agent workflows, Messages View makes multi-turn traces easier to read and understand, reducing the friction of debugging long agent sessions.&lt;/p&gt;

&lt;p&gt;🔍 Why This Matters&lt;/p&gt;

&lt;p&gt;The interesting part isn't any single feature.&lt;/p&gt;

&lt;p&gt;The bigger story is that agent infrastructure is becoming a full production stack:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   AI Models
            ↓
     Agent Runtime
            ↓
   Tools + Memory + Files
            ↓
      Secure Sandbox
            ↓
 Observability + Evaluation
            ↓
 Governance + Cost Controls
            ↓
      Production Scale
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We're moving from “build an agent” to “operate an autonomous system.”&lt;/p&gt;

&lt;p&gt;That changes the engineering questions too.&lt;/p&gt;

&lt;p&gt;It's no longer only:&lt;/p&gt;

&lt;p&gt;Which model should I use?&lt;/p&gt;

&lt;p&gt;It's also:&lt;/p&gt;

&lt;p&gt;How do I monitor the agent?&lt;br&gt;
How do I test it?&lt;br&gt;
How do I control its cost?&lt;br&gt;
How do I secure tool execution?&lt;br&gt;
How do I recover from failures?&lt;br&gt;
How do I improve it continuously?&lt;/p&gt;

&lt;p&gt;That is where the next generation of agent infrastructure is being built.&lt;/p&gt;

&lt;p&gt;What’s your take? Are you moving toward managed agent runtimes, or do you prefer to keep the runtime infrastructure self-hosted?&lt;/p&gt;

</description>
      <category>langchain</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>automation</category>
    </item>
    <item>
      <title>Jev vs LLMs: Why AI Agents May Need a Decision Layer</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Thu, 24 Sep 2026 02:56:52 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/jev-vs-llms-why-ai-agents-may-need-a-decision-layer-338a</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/jev-vs-llms-why-ai-agents-may-need-a-decision-layer-338a</guid>
      <description>&lt;h1&gt;
  
  
  Jev vs LLMs: Why AI Agents May Need a Decision Layer
&lt;/h1&gt;

&lt;p&gt;Here's an uncomfortable pattern in modern AI applications:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User input
   ↓
LLM
   ↓
generated text
   ↓
parser
   ↓
application logic
   ↓
action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We're often using a general-purpose language model to make a tiny decision.&lt;/p&gt;

&lt;p&gt;Should we retry?&lt;/p&gt;

&lt;p&gt;Should we escalate?&lt;/p&gt;

&lt;p&gt;Which tool should we call?&lt;/p&gt;

&lt;p&gt;Which model should handle this?&lt;/p&gt;

&lt;p&gt;Should this request be blocked?&lt;/p&gt;

&lt;p&gt;Those are not necessarily generation problems.&lt;/p&gt;

&lt;p&gt;They're &lt;strong&gt;decision problems&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That's where &lt;strong&gt;Jev&lt;/strong&gt;, TypeSafe AI's first System One model, gets interesting.&lt;/p&gt;

&lt;p&gt;TypeSafe introduced Jev in September 2026 as a model designed around structured decisions rather than open-ended string generation.&lt;/p&gt;




&lt;h2&gt;
  
  
  The core idea
&lt;/h2&gt;

&lt;p&gt;The simplest mental model is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Traditional LLM:

state → generated string


System One:

state → typed decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev's developer documentation describes the interface as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input:
Text / JSON / text arrays

Output:
Choice / Score / Noul

Control flow:
Your application
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model doesn't own your application flow.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Your code does.&lt;/p&gt;




&lt;h1&gt;
  
  
  Jev vs LLM
&lt;/h1&gt;

&lt;p&gt;Let's make the difference concrete.&lt;/p&gt;

&lt;h2&gt;
  
  
  LLM approach
&lt;/h2&gt;

&lt;p&gt;Suppose an agent receives:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The customer says:

"I was charged twice for the same order.
Please refund the duplicate payment."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You might ask an LLM:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Classify this request and return JSON.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then receive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"team"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"urgent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.96&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looks great.&lt;/p&gt;

&lt;p&gt;But your application is now depending on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prompt instructions&lt;/li&gt;
&lt;li&gt;generated output&lt;/li&gt;
&lt;li&gt;schema adherence&lt;/li&gt;
&lt;li&gt;parsing&lt;/li&gt;
&lt;li&gt;interpretation&lt;/li&gt;
&lt;li&gt;potentially unnecessary text generation&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  Jev approach
&lt;/h1&gt;

&lt;p&gt;Define the decisions your application actually needs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Question 1:
Which team should handle this?

Choices:
billing
technical
general
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Question 2:
Should this request be considered urgent?

Noul:
yes / no
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And perhaps:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Question 3:
How frustrated is the customer?

Score:
0 = low
1 = medium
2 = high
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result can be consumed directly by application code.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;state
  │
  ├── Choice → billing
  │
  ├── Noul   → 0.88
  │
  └── Score  → 1.7
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev's current developer materials document these three output types and probability/confidence information.&lt;/p&gt;




&lt;h1&gt;
  
  
  A practical agent architecture
&lt;/h1&gt;

&lt;p&gt;Here's where this becomes more interesting.&lt;/p&gt;

&lt;p&gt;Imagine an agent proposes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tool"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"delete_project"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"project"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"production"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't let the model directly execute it.&lt;/p&gt;

&lt;p&gt;Instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ┌───────────────┐
                 │   AI Agent    │
                 │   proposes    │
                 │   tool call   │
                 └───────┬───────┘
                         │
                         ▼
                 ┌───────────────┐
                 │     Jev       │
                 │   Decision    │
                 └───────┬───────┘
                         │
             ┌───────────┼───────────┐
             ▼           ▼           ▼
           ALLOW       REVIEW       BLOCK
             │           │           │
             ▼           ▼           ▼
           execute      human        stop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important architectural rule is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev decides. Code controls.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Your deterministic application layer should still own authorization, thresholds, audit logs, and side effects.&lt;/p&gt;




&lt;h1&gt;
  
  
  Code example
&lt;/h1&gt;

&lt;p&gt;The exact SDK syntax can change, so treat this as an architectural example rather than a copy-paste contract:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;decision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;jev&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;decide&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;agent_state&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;questions&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_policy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choice&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;instructions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Should this tool call execute?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allow&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Safe and authorized&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Human approval required&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;block&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Do not execute&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;choice&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_policy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choice&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_policy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;confidence&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;choice&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;allow&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.90&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;execute_tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;choice&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;request_human_approval&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;block_tool&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice something important:&lt;/p&gt;

&lt;p&gt;The model doesn't get to decide what &lt;code&gt;0.90&lt;/code&gt; means.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The developer does.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the difference between an AI prediction and an application policy.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why probabilities matter
&lt;/h1&gt;

&lt;p&gt;Suppose Jev returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;allow  = 0.94
review = 0.04
block  = 0.02
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your application might decide:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.90&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;human_review&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Another application might require:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;confidence&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.995&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;human_review&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same model.&lt;/p&gt;

&lt;p&gt;Different risk tolerance.&lt;/p&gt;

&lt;p&gt;This makes the model a component inside a larger control system rather than the system itself.&lt;/p&gt;

&lt;p&gt;And that's exactly the kind of workflow TypeSafe describes for System One models.&lt;/p&gt;




&lt;h1&gt;
  
  
  Choice vs Score vs Noul
&lt;/h1&gt;

&lt;p&gt;A useful way to think about Jev's interface is:&lt;/p&gt;

&lt;h3&gt;
  
  
  Choice
&lt;/h3&gt;

&lt;p&gt;Use when you need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A / B / C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Which model should process this request?

fast
deep
human
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Score
&lt;/h3&gt;

&lt;p&gt;Use when you need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;How much?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;How urgent is this request?

0 = low
1 = medium
2 = high
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Noul
&lt;/h3&gt;

&lt;p&gt;Use when you need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Yes / No
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Should this request be escalated?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Jev's current documentation describes Noul as a value from 0 to 1 for binary questions.&lt;/p&gt;




&lt;h1&gt;
  
  
  Where Jev could fit
&lt;/h1&gt;

&lt;p&gt;This architecture opens up some interesting use cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Agent routing
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;request
   ↓
Jev
   ↓
simple ──────→ cheap model
complex ─────→ reasoning model
uncertain ───→ human
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  2. Tool verification
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;proposed tool call
       ↓
      Jev
       ↓
allow / review / block
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  3. Retry control
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;failed request
      ↓
     Jev
      ↓
retry / change strategy / stop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  4. Support automation
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;message
   ↓
Jev
   ├── billing
   ├── technical
   └── general
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  5. Search ranking
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;query + result
       ↓
      Jev
       ↓
relevance score
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Jev community is already experimenting with agent routing, browser automation, compaction, MCP tools and other integrations.&lt;/p&gt;




&lt;h1&gt;
  
  
  But Jev isn't a replacement for an LLM
&lt;/h1&gt;

&lt;p&gt;This is probably the most important point.&lt;/p&gt;

&lt;p&gt;Don't think:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Jev &amp;gt; LLM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Think:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Jev + LLM + code
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A general-purpose LLM is still the natural component for things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;writing&lt;/li&gt;
&lt;li&gt;conversation&lt;/li&gt;
&lt;li&gt;code generation&lt;/li&gt;
&lt;li&gt;open-ended reasoning&lt;/li&gt;
&lt;li&gt;planning&lt;/li&gt;
&lt;li&gt;summarization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A decision model is useful when the application already knows the possible decisions.&lt;/p&gt;

&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM:
"Write a response to the customer."

Jev:
"Which queue should handle this?"

Code:
"Actually execute the routing."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Different problems.&lt;/p&gt;

&lt;p&gt;Different interfaces.&lt;/p&gt;




&lt;h1&gt;
  
  
  What makes this technically interesting?
&lt;/h1&gt;

&lt;p&gt;TypeSafe describes Jev as using a different architecture and training approach called &lt;strong&gt;Reinforcement Learning for Calibrated Decisions (RLCD)&lt;/strong&gt;. The company says Jev produces probabilities in parallel rather than autoregressively generating a string token by token.&lt;/p&gt;

&lt;p&gt;That's a fundamentally different optimization target.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;maximize useful generated sequence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the goal becomes closer to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;produce useful + calibrated decisions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tradeoff is obvious too:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You give up general string generation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In exchange, the model is specialized for the decision interface.&lt;/p&gt;




&lt;h1&gt;
  
  
  The performance claim
&lt;/h1&gt;

&lt;p&gt;TypeSafe currently advertises Jev as dramatically faster and cheaper than LLMs for its System One workflows, including a headline comparison of &lt;strong&gt;193.6× faster and 444.6× cheaper&lt;/strong&gt; on its site.&lt;/p&gt;

&lt;p&gt;Those are &lt;strong&gt;TypeSafe's reported results&lt;/strong&gt;, not an independent benchmark.&lt;/p&gt;

&lt;p&gt;That's an important distinction.&lt;/p&gt;

&lt;p&gt;Before putting Jev in a production workflow, I'd measure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;latency
accuracy
calibration
cost
failure modes
distribution shift
human escalation rate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;on your own data.&lt;/p&gt;

&lt;p&gt;The Jev developer materials also recommend representative testing and human review for uncertain/high-impact cases.&lt;/p&gt;




&lt;h1&gt;
  
  
  The bigger idea
&lt;/h1&gt;

&lt;p&gt;We've spent years making AI models increasingly good at producing text.&lt;/p&gt;

&lt;p&gt;But production software isn't made entirely of text.&lt;/p&gt;

&lt;p&gt;It's made of decisions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;route
retry
approve
reject
escalate
rank
stop
continue
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Maybe the next evolution of AI applications isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One giant model that does everything.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Maybe it's:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             ┌────────────┐
             │     LLM    │
             │  Generate  │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │    Jev     │
             │   Decide   │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │    Code    │
             │   Control  │
             └─────┬──────┘
                   │
                   ▼
                 ACTION
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;LLMs generate.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decision models decide.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code controls.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's a much more interesting architecture for AI agents than simply throwing a bigger prompt at a bigger model.&lt;/p&gt;

&lt;p&gt;And that's why Jev is worth experimenting with.&lt;/p&gt;

&lt;p&gt;Try it, benchmark it, break it, and see where the decision primitive actually belongs in your stack.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>development</category>
      <category>agents</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Wed, 23 Sep 2026 17:51:31 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/-2eln</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/-2eln</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2" class="crayons-story__hidden-navigation-link"&gt;Research Dossier: An Agent That Shows Its Disagreements Instead of Hiding Them&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
      &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2" class="crayons-article__context-note crayons-article__context-note__feed"&gt;&lt;p&gt;Sanity Challenge Path One Submission&lt;/p&gt;

&lt;/a&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/pratik_12b3f8bf3b50e48bae" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3353910%2F924541cd-88b5-47ef-8933-c8fbec3be59d.png" alt="pratik_12b3f8bf3b50e48bae profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/pratik_12b3f8bf3b50e48bae" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Pratik
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Pratik
                
                
              
              &lt;div id="story-author-preview-content-4701266" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/pratik_12b3f8bf3b50e48bae" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3353910%2F924541cd-88b5-47ef-8933-c8fbec3be59d.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Pratik&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 20&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2" id="article-link-4701266"&gt;
          Research Dossier: An Agent That Shows Its Disagreements Instead of Hiding Them
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devchallenge"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devchallenge&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/sanitychallenge"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;sanitychallenge&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/sanity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;sanity&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;15&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              4&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Research Dossier: An Agent That Shows Its Disagreements Instead of Hiding Them</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Sun, 20 Sep 2026 18:09:29 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/research-dossier-an-agent-that-shows-its-disagreements-instead-of-hiding-them-5fe2</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/sanity-2026-09-16"&gt;Sanity Challenge, Path One: Ship an Agent That Queries Real Content&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;p&gt;Research Dossier is a multi-agent research analyst built with LangGraph.&lt;/p&gt;

&lt;p&gt;Instead of asking one model to answer a research question from its own knowledge, the system routes the question through four stages:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Research → Analysis → Writing → Review&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The research stage is grounded in a Sanity Knowledge Base, queried directly via GROQ. Sanity Context wasn't yet enabled on my org during the build window, so the research agent's tool talks to Sanity's Content API directly rather than through the managed Context/MCP layer. The tool is isolated behind a single function, so swapping it for the Sanity Context MCP endpoint later is a contained change, not a rewrite of the agent's reasoning logic.&lt;/p&gt;

&lt;p&gt;The interesting problem I wanted to solve is not simply finding information. It is handling situations where the sources themselves disagree.&lt;/p&gt;

&lt;p&gt;When conflicting claims are found, Research Dossier does not silently merge them into one confident answer. It preserves the disagreement, shows the sources behind the claims, and marks the final report:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CONTESTED&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem
&lt;/h2&gt;

&lt;p&gt;My Knowledge Base contains multiple pieces of evidence about LangGraph checkpoint deserialization.&lt;/p&gt;

&lt;p&gt;One source is the LangGraph documentation itself, which implies checkpointing "just works" safely out of the box. Another is the official security advisory for CVE-2026-28277, which documents unsafe msgpack deserialization by default and describes strict-mode and allowlist-based hardening. A third source is a second advisory database's framing of the same CVE, which tempers the risk with essential context — it's classified as a defense-in-depth issue requiring an attacker to already have privileged write access, not a standalone remote exploit. A fourth is a community checkpointer implementation, which raises the separate question of whether that hardening even reliably extends to third-party backends.&lt;/p&gt;

&lt;p&gt;The system keeps all of these claims and their provenance distinct, rather than flattening them into one answer.&lt;/p&gt;

&lt;p&gt;That matters because a normal keyword search can find all of these pieces of information without preserving the relationships between them.&lt;/p&gt;

&lt;p&gt;Research Dossier treats disagreement as structured information instead of noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Live app: &lt;a href="https://multi-agent-research-analyst.vercel.app/" rel="noopener noreferrer"&gt;https://multi-agent-research-analyst.vercel.app/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Try one of the built-in example questions, or ask:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Does LangGraph handle checkpoint deserialization safely by default?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Watch the case log as the system progresses through Research → Analysis → Writing → Review, with each step's output expandable if you want to see the full reasoning. Once a draft is ready, it enters a human-approval step — you can review or edit it before the case is closed and the final dossier is rendered. The report preserves the conflicting evidence and distinguishes stronger sources from lower-trust material instead of flattening everything into one conclusion, and stamps the report CONTESTED when sources disagreed.&lt;/p&gt;

&lt;p&gt;You can also browse the raw Knowledge Base directly at &lt;code&gt;/sources&lt;/code&gt; — every claim, its source, and what it contradicts, without needing to ask a question first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Keyword Search Isn't Enough
&lt;/h2&gt;

&lt;p&gt;A keyword search can return documentation, security advisories, and community discussions that mention checkpoint serialization.&lt;/p&gt;

&lt;p&gt;The problem is that matching text does not tell the agent how those pieces of information relate to each other.&lt;/p&gt;

&lt;p&gt;Research Dossier retrieves structured claims together with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;their sources,&lt;/li&gt;
&lt;li&gt;source trust information,&lt;/li&gt;
&lt;li&gt;and relationships between claims.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That lets the analysis stage reason about disagreement rather than simply presenting a list of matching passages.&lt;/p&gt;

&lt;p&gt;The result can be explicitly marked CONTESTED when the evidence remains in conflict.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Used Sanity
&lt;/h2&gt;

&lt;p&gt;I modeled the Knowledge Base around three core document types:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;topic&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;source&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;claim&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key relationship is:&lt;br&gt;
claim&lt;br&gt;
└── contradicts → claim&lt;/p&gt;

&lt;p&gt;This makes disagreement machine-readable. Instead of asking an LLM to infer whether two unrelated passages appear to disagree, the content model explicitly represents that relationship.&lt;/p&gt;

&lt;p&gt;A claim also references its source, allowing the research pipeline to retain provenance while moving from retrieval to analysis to writing and review.&lt;/p&gt;
&lt;h2&gt;
  
  
  Retrieval Layer
&lt;/h2&gt;

&lt;p&gt;The research agent's tool queries the Knowledge Base directly via GROQ, expanding &lt;code&gt;topic&lt;/code&gt;, &lt;code&gt;source&lt;/code&gt;, and &lt;code&gt;contradicts[]-&amp;gt;source&lt;/code&gt; in a single request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LangGraph
│
▼
Research Agent
│
▼
GROQ query (@sanity/client)
│
▼
Sanity Knowledge Base
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This retrieval layer is deliberately isolated behind a single tool function — swapping it for the Sanity Context MCP endpoint is a contained change, not a rewrite of the agent's reasoning logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 User Question&lt;br&gt;
                      │&lt;br&gt;
                      ▼&lt;br&gt;
               ┌──────────────┐&lt;br&gt;
               │   Research   │&lt;br&gt;
               │    Agent     │&lt;br&gt;
               └──────┬───────┘&lt;br&gt;
                      │&lt;br&gt;
                      ▼&lt;br&gt;
            ┌────────────────────┐&lt;br&gt;
            │   GROQ Query Layer │&lt;br&gt;
            └─────────┬──────────┘&lt;br&gt;
                      │&lt;br&gt;
                      ▼&lt;br&gt;
            ┌────────────────────┐&lt;br&gt;
            │ Sanity Knowledge   │&lt;br&gt;
            │       Base         │&lt;br&gt;
            │                    │&lt;br&gt;
            │ claims + sources + │&lt;br&gt;
            │ contradictions     │&lt;br&gt;
            └─────────┬──────────┘&lt;br&gt;
                      │&lt;br&gt;
                      ▼&lt;br&gt;
               ┌──────────────┐&lt;br&gt;
               │   Analysis   │&lt;br&gt;
               └──────┬───────┘&lt;br&gt;
                      ▼&lt;br&gt;
               ┌──────────────┐&lt;br&gt;
               │   Writing    │&lt;br&gt;
               └──────┬───────┘&lt;br&gt;
                      ▼&lt;br&gt;
               ┌──────────────┐&lt;br&gt;
               │    Review    │&lt;br&gt;
               └──────┬───────┘&lt;br&gt;
                      │&lt;br&gt;
             APPROVED / REVISE&lt;br&gt;
                      │&lt;br&gt;
                      ▼&lt;br&gt;
            Human Approval / Edit&lt;br&gt;
                      │&lt;br&gt;
                      ▼&lt;br&gt;
                Final Report&lt;br&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h2&gt;
&lt;br&gt;
  &lt;br&gt;
  &lt;br&gt;
  What Each Agent Does&lt;br&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Research&lt;/strong&gt;&lt;br&gt;
The research agent is responsible for retrieval. It queries the Sanity Knowledge Base directly via GROQ and is explicitly instructed not to rely on general knowledge alone. When contradictory evidence is retrieved, it keeps both sides and their sources.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analysis&lt;/strong&gt;&lt;br&gt;
The analysis agent compares the retrieved claims. It considers source provenance, trust level, and recency, and distinguishes well-supported evidence from weaker or single-source claims.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Writing&lt;/strong&gt;&lt;br&gt;
The writing agent converts the analysis into a source-linked report. It is instructed not to introduce factual claims that were not present in the research findings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Review&lt;/strong&gt;&lt;br&gt;
The review agent acts as a hallucination gate. It checks the draft against the original research findings. Unsupported claims trigger a revision pass instead of being silently accepted. The workflow allows bounded revision before producing the final report — which then goes to a human approval step before being marked closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Sanity?
&lt;/h2&gt;

&lt;p&gt;The project could have been built as a conventional search application. That would miss the important part of the problem.&lt;/p&gt;

&lt;p&gt;The Knowledge Base stores claims, sources, and relationships between claims. In particular, the &lt;code&gt;contradicts&lt;/code&gt; relationship makes disagreement part of the data model.&lt;/p&gt;

&lt;p&gt;That structure changes what the agent can do. It is not simply retrieving text that matches a query. It is retrieving structured knowledge that can be compared, traced back to sources, and carried through a multi-agent reasoning and review pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Example Sources
&lt;/h2&gt;

&lt;p&gt;The checkpoint-deserialization investigation uses sources with different levels of authority, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LangGraph checkpoint documentation&lt;/li&gt;
&lt;li&gt;LangGraph security advisory — CVE-2026-28277&lt;/li&gt;
&lt;li&gt;A second advisory database's framing of the same CVE, with mitigating context&lt;/li&gt;
&lt;li&gt;A community checkpointer implementation, stored in the Knowledge Base&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent preserves those provenance differences rather than treating every retrieved claim as equally authoritative.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sanity Project Details
&lt;/h2&gt;

&lt;p&gt;Project ID: &lt;code&gt;4kagnnrl&lt;/code&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/pratikdevelop/multi-agent-research-analyst" rel="noopener noreferrer"&gt;https://github.com/pratikdevelop/multi-agent-research-analyst&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The repository contains the LangGraph agent, Sanity schemas, and the Next.js interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Wanted to Demonstrate
&lt;/h2&gt;

&lt;p&gt;The interesting part of this project isn't simply that multiple agents can call a CMS.&lt;/p&gt;

&lt;p&gt;It is that structured content can make disagreement explicit.&lt;/p&gt;

&lt;p&gt;Instead of forcing conflicting evidence into one confident answer, the Knowledge Base preserves the claims and their relationships, the analysis stage compares them, the review stage checks that the final report remains grounded in the retrieved evidence, and a human gets the final say before the case closes.&lt;/p&gt;

&lt;p&gt;Research Dossier doesn't try to make disagreement disappear. It makes disagreement visible.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>sanitychallenge</category>
      <category>sanity</category>
      <category>ai</category>
    </item>
    <item>
      <title>How to Build a Self-Healing CI/CD Pipeline with AI Agents (2026 Guide)</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Wed, 09 Sep 2026 07:00:13 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/how-to-build-a-self-healing-cicd-pipeline-with-ai-agents-2026-guide-g00</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/how-to-build-a-self-healing-cicd-pipeline-with-ai-agents-2026-guide-g00</guid>
      <description>&lt;p&gt;Let’s be real: nothing kills developer flow state quite like a failed CI pipeline. You push your code, context-switch to grab a coffee, and come back to a wall of red text and a cryptic NullPointerException in a module you didn’t even touch.&lt;br&gt;
In 2026, we don't have to do this anymore.&lt;br&gt;
With the maturity of agentic frameworks and local code-generation models, we can now build self-healing pipelines. Instead of just alerting you that a build failed, the pipeline intercepts the failure, spins up an AI agent, generates a fix, verifies it, and opens a PR.&lt;br&gt;
Here is a pragmatic guide to implementing this in your workflow today.&lt;br&gt;
The Architecture&lt;br&gt;
To build this, we need three things:&lt;br&gt;
A webhook listener for our CI provider (GitHub Actions, GitLab, etc.).&lt;br&gt;
An agentic framework (like LangGraph or AutoGen) to handle the reasoning loop.&lt;br&gt;
An ephemeral sandbox to safely test the agent's fix.&lt;br&gt;
&lt;strong&gt;Step 1: Catching the Failure Context&lt;/strong&gt;&lt;br&gt;
The biggest mistake devs make when building AI debuggers is just passing the raw error log to the LLM. Context is king. You need to pass the error log, the specific file that failed, and the recent git diff.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# webhook_handler.py
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Request&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;DebugAgent&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/webhook/ci-failure&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ci_failure_webhook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="c1"&gt;# Extract the exact context the agent needs
&lt;/span&gt;    &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;error_log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;logs&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;stderr&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failing_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;logs&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;failing_file_path&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;recent_diff&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;repository&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;last_commit_diff&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;test_command&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;config&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;test_script&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;# Hand off to the agent
&lt;/span&gt;    &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;DebugAgent&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;heal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2: The Agentic Reasoning Loop&lt;/strong&gt;&lt;br&gt;
We use a graph-based agent so it can iterate. If the first patch doesn't fix the test, the agent needs to read the new error log and try again.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# agent.py
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langgraph.graph&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;StateGraph&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;END&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;get_model&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;sandbox&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;run_tests_in_sandbox&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;DebugAgent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;llm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qwen-coder-local&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# Keep it local for speed/security!
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;graph&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_build_graph&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_build_graph&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;workflow&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;StateGraph&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="c1"&gt;# Nodes
&lt;/span&gt;        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analyze&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;analyze_error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;patch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;generate_patch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;verify_fix&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="c1"&gt;# Edges
&lt;/span&gt;        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_entry_point&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analyze&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analyze&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;patch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;patch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

        &lt;span class="c1"&gt;# Conditional edge: If tests pass, end. If fail, loop back to analyze (max 3 times)
&lt;/span&gt;        &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_conditional_edges&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;verify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;should_retry&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;analyze&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;success&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;END&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;escalate&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;END&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;workflow&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;analyze_error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Prompt the LLM to understand the root cause based on logs + diff
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt; 

    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate_patch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Generate a unified diff patch
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt;

    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;verify_fix&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Apply patch to ephemeral docker container and run tests
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  🚨 Pro-Tips &amp;amp; Gotchas from the Trenches
&lt;/h2&gt;

&lt;p&gt;If you are building this in production, watch out for these three things:&lt;br&gt;
&lt;strong&gt;Log Truncation&lt;/strong&gt;: LLMs will hallucinate if you feed them 50,000 lines of logs. Write a pre-processor that extracts only the fatal error blocks and the surrounding 50 lines of context.&lt;br&gt;
Infinite Loops: Always set a max_iterations limit (I use 3). If the agent can't fix it in 3 tries, it's a complex architectural issue. Escalate to a human.&lt;br&gt;
&lt;strong&gt;Security&lt;/strong&gt;: Never give the agent write access to your main branch. The agent should only ever push to a temporary branch like ai-fix/issue-123. Let human reviewers merge it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrap Up
&lt;/h2&gt;

&lt;p&gt;Building a self-healing pipeline takes a weekend of setup, but it saves hundreds of hours of context-switching over the year. Start small: hook an agent up to your linting failures first, then move to unit test failures.&lt;br&gt;
Have you implemented agentic CI/CD in your stack yet? What framework are you using for the reasoning loop? Drop your setups in the comments below!&lt;/p&gt;

</description>
      <category>devops</category>
      <category>ai</category>
      <category>cicd</category>
      <category>githubactions</category>
    </item>
    <item>
      <title>What I Learned from Reading the PyCharm PyTorch Tutorial</title>
      <dc:creator>Pratik</dc:creator>
      <pubDate>Sat, 01 Aug 2026 02:59:13 +0000</pubDate>
      <link>https://dev.to/pratik_12b3f8bf3b50e48bae/what-i-learned-from-reading-the-pycharm-pytorch-tutorial-4fjo</link>
      <guid>https://dev.to/pratik_12b3f8bf3b50e48bae/what-i-learned-from-reading-the-pycharm-pytorch-tutorial-4fjo</guid>
      <description>&lt;p&gt;I recently read the &lt;strong&gt;"PyTorch Tutorial for Deep Learning"&lt;/strong&gt; on the PyCharm Blog by JetBrains as part of my AI learning journey.&lt;/p&gt;

&lt;p&gt;Although this wasn't a hands-on coding session, it helped me build a stronger understanding of the fundamentals behind PyTorch and modern deep learning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Highlights from the Tutorial
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;An introduction to PyTorch and why it's one of the most popular deep learning frameworks.&lt;/li&gt;
&lt;li&gt;Understanding tensors and their role in neural networks.&lt;/li&gt;
&lt;li&gt;Creating, reshaping, and manipulating tensors.&lt;/li&gt;
&lt;li&gt;Basic tensor arithmetic and NumPy interoperability.&lt;/li&gt;
&lt;li&gt;Using CUDA for GPU acceleration.&lt;/li&gt;
&lt;li&gt;How &lt;code&gt;torch.nn.Module&lt;/code&gt; is used to build neural networks.&lt;/li&gt;
&lt;li&gt;Understanding the purpose of the &lt;code&gt;forward()&lt;/code&gt; method.&lt;/li&gt;
&lt;li&gt;Learning how Autograd automatically computes gradients during training.&lt;/li&gt;
&lt;li&gt;Setting up a PyTorch development environment in PyCharm.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why This Article Stood Out
&lt;/h2&gt;

&lt;p&gt;What I appreciated most is that the tutorial focuses on explaining &lt;em&gt;why&lt;/em&gt; these concepts matter before jumping into building models. It provides a clear roadmap for beginners who want to understand the building blocks of deep learning.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's Next?
&lt;/h2&gt;

&lt;p&gt;My next step is to put these concepts into practice by building beginner-friendly PyTorch projects, starting with the MNIST handwritten digit classifier mentioned in the tutorial.&lt;/p&gt;

&lt;p&gt;Learning is most valuable when theory is followed by practice, and that's exactly what I plan to do next.&lt;/p&gt;

&lt;p&gt;If you're beginning your journey into AI or deep learning, this tutorial is a great place to start.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Original article:&lt;/strong&gt; &lt;a href="https://blog.jetbrains.com/pycharm/2026/07/pytorch-tutorial-for-deep-learning/" rel="noopener noreferrer"&gt;https://blog.jetbrains.com/pycharm/2026/07/pytorch-tutorial-for-deep-learning/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Happy learning! 🚀&lt;/p&gt;

</description>
      <category>pytorch</category>
      <category>python</category>
      <category>ai</category>
      <category>learning</category>
    </item>
  </channel>
</rss>
