<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ryan2run</title>
    <description>The latest articles on DEV Community by ryan2run (@ryan_zhao).</description>
    <link>https://dev.to/ryan_zhao</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4118260%2Fbf4cdc0a-e5d6-458f-9fa9-89818f1b7c57.jpg</url>
      <title>DEV Community: ryan2run</title>
      <link>https://dev.to/ryan_zhao</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ryan_zhao"/>
    <language>en</language>
    <item>
      <title>T-Mem: Memory That Anticipates, Not Archives - Tencent's Breakthrough in AI Long-Term Memory</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Tue, 15 Sep 2026 02:04:35 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/t-mem-memory-that-anticipates-not-archives-tencents-breakthrough-in-ai-long-term-memory-2dle</link>
      <guid>https://dev.to/ryan_zhao/t-mem-memory-that-anticipates-not-archives-tencents-breakthrough-in-ai-long-term-memory-2dle</guid>
      <description>&lt;h1&gt;
  
  
  T-Mem: Memory That Antanticates, Not Archives
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhlaqbd3dggb5we9aqj81.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhlaqbd3dggb5we9aqj81.png" alt="T-Mem Associative Memory Framework" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Problem with Current AI Memory
&lt;/h2&gt;

&lt;p&gt;Large language model agents are increasingly entering long-term companion chat scenarios — not just answering one question and leaving, but building lasting relationships with users. This creates a specific requirement for memory systems: users don't repeat the same things year after year in the same words. The memory system must retrieve past memories even when phrasing, topics, and contexts have completely changed.&lt;/p&gt;

&lt;p&gt;But today, almost all long-term memory solutions share the same fundamental assumption: as long as the query is "similar enough" to the memory, it can be retrieved. This only works when memory and dialogue have literal or semantic similarity. In real conversations, however, phrasing changes, contexts shift, yet memories still need to be recalled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tencent's Solution: T-Mem
&lt;/h2&gt;

&lt;p&gt;T-Mem flips this assumption upside down: instead of struggling to find similarity at retrieval time, save preset scenarios as bridges at the moment of memory write-in. This way, even when memory and dialogue have zero semantic similarity, key memories can still be recalled through contextual association.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Key Insight&lt;/strong&gt;: A long-term memory system earns its adaptive value not by archiving the dialogue stream faithfully, but by anticipating, at write time, the future cues under which its contents will need to be reached.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  How T-Mem Works
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The 2x2 Quadrant Framework
&lt;/h3&gt;

&lt;p&gt;T-Mem models memory across two orthogonal axes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval Direction&lt;/strong&gt;: Descriptive recall (surface similarity) vs. Associative recall (latent connection)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory Granularity&lt;/strong&gt;: Single fact vs. Complete scene&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This creates four quadrants (QI-QIV), each with a trigger family:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;QI Entity Trigger&lt;/strong&gt;: Enriches single facts with generalized labels&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QII Bridge Trigger&lt;/strong&gt;: Predicts "in what context will this fact be needed?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QIII Forward Trigger&lt;/strong&gt;: Thinks several steps ahead along the conversation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QIV Scene Trigger&lt;/strong&gt;: Writes complete scenes as multi-dimensional archives&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  The Example That Explains Everything
&lt;/h3&gt;

&lt;p&gt;Consider this scenario:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Months ago&lt;/strong&gt;: "Xiao Zhang is allergic to seafood, went to the hospital last week, need to be careful."&lt;br&gt;
&lt;strong&gt;Now&lt;/strong&gt;: "Where are we going for team building next week?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;These two sentences share ZERO semantic similarity. But the memory about Xiao Zhang's seafood allergy is CRITICAL for choosing a restaurant that avoids seafood.&lt;/p&gt;

&lt;p&gt;With T-Mem, when "Xiao Zhang is seafood allergic" is written to memory, the system asks: "In what future scenarios will this information be used?" It pre-saves triggers for team building restaurant selection, health checkups, allergy history collection, etc.&lt;/p&gt;

&lt;p&gt;When "Where are we going for team building?" arrives, T-Mem lights up the pre-saved trigger, retrieves the seafood allergy memory, and suggests: "How about not going to the seaside? Xiao Zhang can't eat seafood."&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance Results
&lt;/h2&gt;

&lt;p&gt;T-Mem achieves &lt;strong&gt;dual SOTA&lt;/strong&gt; on two memory benchmarks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;T-Mem Score&lt;/th&gt;
&lt;th&gt;Improvement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;LoCoMo&lt;/td&gt;
&lt;td&gt;80.26%&lt;/td&gt;
&lt;td&gt;New SOTA&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LoCoMo-Plus&lt;/td&gt;
&lt;td&gt;74.81%&lt;/td&gt;
&lt;td&gt;New SOTA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The real test is LoCoMo-Plus, which deliberately removes all lexical and semantic similarity between clues and answers. Mainstream systems drop 28-50 percentage points; T-Mem only drops &lt;strong&gt;5.45 percentage points&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Implementation
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code&lt;/strong&gt;: &lt;a href="https://github.com/Sherlockwz/T-Mem" rel="noopener noreferrer"&gt;github.com/Sherlockwz/T-Mem&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Paper&lt;/strong&gt;: &lt;a href="https://arxiv.org/abs/2606.15405" rel="noopener noreferrer"&gt;arxiv.org/abs/2606.15405&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Team&lt;/strong&gt;: Tencent PCG (郭伟东, 王达凯, 汪子轩, 刘辉, 徐羽)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Published&lt;/strong&gt;: EMNLP 2026 Main Conference&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deployed&lt;/strong&gt;: QQ AI Partner project&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;T-Mem demonstrates that associative memory is not an optional bonus feature, but the missing half of the similarity-based approach. Memory systems should anticipate how their contents will be reached, not just archive what was said.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tags: AI, Memory, LLM, Tencent, Agent, EMNLP&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>memory</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>PhysBrain 1.5: Chinese Physical AI Breakthrough - Physical Foundation Model Tops Open Source Rankings</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Tue, 15 Sep 2026 02:04:12 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/physbrain-15-chinese-physical-ai-breakthrough-physical-foundation-model-tops-open-source-51ok</link>
      <guid>https://dev.to/ryan_zhao/physbrain-15-chinese-physical-ai-breakthrough-physical-foundation-model-tops-open-source-51ok</guid>
      <description>&lt;h1&gt;
  
  
  PhysBrain 1.5: Chinese Physical AI Breakthrough
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn90fe9i90mnzc66vt65p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn90fe9i90mnzc66vt65p.png" alt="PhysBrain Physical AI Loop" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Physical AI Race Heats Up
&lt;/h2&gt;

&lt;p&gt;Physical AI is becoming the most critical track in global AI for 2026. While large models have spent three years pushing "can talk, can think, can work" to the extreme, the industry is now asking: &lt;strong&gt;Can models enter the real world? Can they take action?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Recent experiments have put this question front and center:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini Robotics 2&lt;/strong&gt; (Google DeepMind, July 2026): Based on Gemini 3.5 Flash, uses "embodied reasoning" brain ER 2 to make robots "watch videos and plan next steps"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-6 Astra&lt;/strong&gt; (OpenAI): Pushed the ceiling of general intelligence higher with breakthroughs in reasoning, coding, and multimodal understanding&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Robocurve Experiment&lt;/strong&gt;: Deployed GPT-6 Astra on a real robotic arm — completed 19/20 "put blocks in bowl" tasks, while Claude Fable 5.1 only completed 8&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But top LLMs still have a gap in physical manipulation. On millimeter-level precision tasks like "blue puzzle piece alignment insertion," GPT-6 Astra's success rate drops from 95% to 10%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enter PhysBrain 1.5
&lt;/h2&gt;

&lt;p&gt;On September 9, 2026, Chinese company &lt;strong&gt;DeepCybo&lt;/strong&gt; (深度机智) released &lt;strong&gt;PhysBrain 1.5&lt;/strong&gt;, a physical foundation model that achieves &lt;strong&gt;72.5 average score across 28 public benchmarks&lt;/strong&gt; — ranking &lt;strong&gt;#1 among open-source models&lt;/strong&gt;, just 1 point behind top closed-source models GPT-6 Astra (73.3) and Gemini 3.6 Flash (73.0).&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Achievements
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Open Source Rank&lt;/td&gt;
&lt;td&gt;#1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average Score (28 benchmarks)&lt;/td&gt;
&lt;td&gt;72.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Source #1&lt;/td&gt;
&lt;td&gt;14 benchmarks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open Source #2&lt;/td&gt;
&lt;td&gt;10 benchmarks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vs. Hy-Embodied-VLM-1.0 (66.0)&lt;/td&gt;
&lt;td&gt;+6.5 points&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;vs. RynnBrain 1.1 (63.1)&lt;/td&gt;
&lt;td&gt;+9.4 points&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Two Model Sizes
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;PhysBrain 1.5-2B&lt;/strong&gt;: 66.6 average — exceeds all non-PhysBrain open models&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PhysBrain 1.5-8B&lt;/strong&gt;: 72.5 average — tops open source, nears closed-source leaders&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both models are fully open-source with Apache 2.0 license.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Physical Loop Architecture
&lt;/h2&gt;

&lt;p&gt;PhysBrain 1.5 is built around a unified architecture called &lt;strong&gt;Physical Loop&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Observe the world&lt;/strong&gt; — perceive the environment&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Understand space and tasks&lt;/strong&gt; — comprehend spatial relationships&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Judge action consequences&lt;/strong&gt; — predict outcomes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execute&lt;/strong&gt; — take physical action&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Correct based on feedback&lt;/strong&gt; — self-correct from new observations&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is a continuous, stable closed-loop system that self-corrects based on feedback.&lt;/p&gt;

&lt;h3&gt;
  
  
  Three Core Capabilities
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Embodied Understanding&lt;/strong&gt;: Knows "what is happening now" — outputs structured spatial coordinates and trajectory targets&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Action Generation&lt;/strong&gt;: Knows "what to do next" — learns from human video, adapts to different robot types&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Future State Prediction&lt;/strong&gt;: Foresees "what happens after" — predicts 1-second physical state in RGB, depth, and robot mask&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Human Learning Approach
&lt;/h2&gt;

&lt;p&gt;DeepCybo's key insight: &lt;strong&gt;Physical intelligence cannot be achieved by translating internet text into actions.&lt;/strong&gt; Humans grow operational skills through repeated "see — try — get feedback — correct." &lt;/p&gt;

&lt;p&gt;The team built &lt;strong&gt;Ego360&lt;/strong&gt;, a human panoramic real-data system, using full-body pose, hand movement, and task-level voice from panoramic video. Physical pre-training supervision comes entirely from human interaction videos.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open Source &amp;amp; Community
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Project Page&lt;/strong&gt;: &lt;a href="https://deepcybo-physai.github.io/PhysBrain-1.5" rel="noopener noreferrer"&gt;deepcybo-physai.github.io/PhysBrain-1.5&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hugging Face&lt;/strong&gt;: &lt;a href="https://huggingface.co/collections/DeepCybo/physbrain-15" rel="noopener noreferrer"&gt;huggingface.co/collections/DeepCybo/physbrain-15&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eval Kit&lt;/strong&gt;: &lt;a href="https://github.com/DeepCybo-PhysAI/PhysBrainEvalKit" rel="noopener noreferrer"&gt;github.com/DeepCybo-PhysAI/PhysBrainEvalKit&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demo&lt;/strong&gt;: &lt;a href="https://huggingface.co/spaces/hugging-apps/physbrain1-5-8b-demo" rel="noopener noreferrer"&gt;huggingface.co/spaces/hugging-apps/physbrain1-5-8b-demo&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Within 3 days of launch, PhysBrain 1.5 received &lt;strong&gt;3,000+ downloads&lt;/strong&gt; on Hugging Face.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;PhysBrain 1.5 demonstrates that physical foundation models are not a short-term trend but a long-term strategic investment. By focusing on human learning patterns and building a complete data-model-body loop, DeepCybo has positioned PhysBrain at the forefront of global physical AI.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tags: AI, Robotics, PhysicalAI, OpenSource, DeepLearning&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>robotics</category>
      <category>physicalai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Iris Search Agent: Open Source Context Management Beats Parameter Scaling</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Tue, 15 Sep 2026 02:03:50 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/iris-search-agent-open-source-context-management-beats-parameter-scaling-34mn</link>
      <guid>https://dev.to/ryan_zhao/iris-search-agent-open-source-context-management-beats-parameter-scaling-34mn</guid>
      <description>&lt;h1&gt;
  
  
  Iris Search Agent: Open Source Context Management Beats Parameter Scaling
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F93m18apxjtkfi4z01po8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F93m18apxjtkfi4z01po8.png" alt="Iris Search Agent Architecture" width="800" height="573"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Open Source Search Agent Revolution
&lt;/h2&gt;

&lt;p&gt;AllSpark Research has released &lt;strong&gt;Iris&lt;/strong&gt;, an open-source search agent that is changing the competitive landscape. Two models are available:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Total Params&lt;/th&gt;
&lt;th&gt;Active Params&lt;/th&gt;
&lt;th&gt;BrowseComp&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Iris-mini&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;35B&lt;/td&gt;
&lt;td&gt;3B&lt;/td&gt;
&lt;td&gt;82.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Iris-pro&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;397B&lt;/td&gt;
&lt;td&gt;17B&lt;/td&gt;
&lt;td&gt;88.6&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both use Mixture-of-Experts architecture with 256K context windows. Weights are released under Apache 2.0.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Search Agents Are Different
&lt;/h2&gt;

&lt;p&gt;Search agents differ from regular chat models because they must:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Decide what to search&lt;/strong&gt; — autonomously choose queries&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluate results&lt;/strong&gt; — determine if more searching is needed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Know when to stop&lt;/strong&gt; — recognize when evidence is sufficient&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For two years, this capability was dominated by closed-source products. Open source had little to compete with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmark Results
&lt;/h2&gt;

&lt;p&gt;Four benchmarks, impressive results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Iris-mini&lt;/th&gt;
&lt;th&gt;Iris-pro&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;BrowseComp&lt;/td&gt;
&lt;td&gt;82.2&lt;/td&gt;
&lt;td&gt;88.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BrowseComp-ZH&lt;/td&gt;
&lt;td&gt;84.8&lt;/td&gt;
&lt;td&gt;85.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSearchQA&lt;/td&gt;
&lt;td&gt;86.9&lt;/td&gt;
&lt;td&gt;92.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HLE&lt;/td&gt;
&lt;td&gt;52.3&lt;/td&gt;
&lt;td&gt;56.4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  The Surprising Chinese Result
&lt;/h3&gt;

&lt;p&gt;On BrowseComp-ZH (Chinese), the 35B model scores 84.8 and the 397B model scores 85.1 — a difference of only &lt;strong&gt;0.3 points&lt;/strong&gt;. The same models on English BrowseComp show a 6.4 point gap.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key insight&lt;/strong&gt;: For Chinese tasks, parameter scale barely matters. Context management is the real differentiator.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context Management Strategy
&lt;/h2&gt;

&lt;p&gt;The paper compares three inference-time context strategies:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;No management&lt;/strong&gt; — baseline&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discard-all&lt;/strong&gt; — clear tool call history when context exceeds threshold, restart from original question&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retry&lt;/strong&gt; — summarize eliminated clues before continuing&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Result&lt;/strong&gt;: Both models improved with context management, and Iris-mini improved MORE than Iris-pro.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;For small models, managing the ever-growing search history is more effective than adding parameters.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Training Data: Reverse-Engineered from Hyperlinks
&lt;/h2&gt;

&lt;p&gt;Multi-hop search questions are hard to create. Writing a question that requires crossing 3-4 webpages is expensive, and models can cheat with string matching.&lt;/p&gt;

&lt;p&gt;Iris takes a different approach:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick a seed page from web corpus hyperlink structure&lt;/li&gt;
&lt;li&gt;Follow outgoing links to build an entity graph&lt;/li&gt;
&lt;li&gt;Create multi-hop chains on the graph&lt;/li&gt;
&lt;li&gt;Rewrite questions to ensure NO clue can be solved by literal matching&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  SFT-RL Climbing
&lt;/h2&gt;

&lt;p&gt;Training uses a process called &lt;strong&gt;SFT-RL climbing&lt;/strong&gt;: supervised fine-tuning and reinforcement learning alternate. The RL phase runs on &lt;strong&gt;real search&lt;/strong&gt;, not offline snapshots. The reward model and observation summarization modules are deployed on the training cluster.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Our policy is optimized by RL against live search."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This means every training step makes real network requests — significant engineering cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters
&lt;/h2&gt;

&lt;p&gt;Open source has caught up to closed source in general chat and code. &lt;strong&gt;Search has been the laggard.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The reason is clear: search agent capability is half model, half tool orchestration framework. Closed-source products tune both; open source often releases only the model.&lt;/p&gt;

&lt;p&gt;Iris releases model weights, context management strategy, and data construction method — the complete package.&lt;/p&gt;

&lt;h2&gt;
  
  
  Community &amp;amp; Availability
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hugging Face&lt;/strong&gt;: Models available for download&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub&lt;/strong&gt;: Training pipeline code marked "coming soon"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;License&lt;/strong&gt;: Apache 2.0 — commercial use allowed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reference&lt;/strong&gt;: &lt;a href="https://arxiv.org/abs/2609.04304" rel="noopener noreferrer"&gt;arXiv 2609.04304&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For anyone wanting to run deep search, the 35B/3B spec is a sweet spot — runs on a single machine with the paper's context management.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Iris proves that open source finally has a serious competitor in search. The combination of model weights, context strategy, and data methodology closes a gap that has persisted for two years.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tags: AI, SearchAgent, OpenSource, LLM, Agent&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>searchagent</category>
      <category>opensource</category>
      <category>llm</category>
    </item>
    <item>
      <title>AI Agent Architecture Patterns: A Deep Dive into Modern Agent Design</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Mon, 14 Sep 2026 00:30:44 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/ai-agent-architecture-patterns-a-deep-dive-into-modern-agent-design-11i4</link>
      <guid>https://dev.to/ryan_zhao/ai-agent-architecture-patterns-a-deep-dive-into-modern-agent-design-11i4</guid>
      <description>&lt;h1&gt;
  
  
  AI Agent Architecture Patterns: A Deep Dive into Modern Agent Design
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwirnrxhb7e3968x4zgvs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwirnrxhb7e3968x4zgvs.png" alt="AI Agent Architecture" width="800" height="577"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI agents are transforming how we interact with technology. But behind every smart agent lies a carefully designed architecture. In this article, we explore the key patterns that power modern AI agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is an AI Agent?
&lt;/h2&gt;

&lt;p&gt;An AI agent is a system that can perceive its environment, make decisions, and take actions to achieve specific goals. Unlike traditional chatbots, agents can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Plan multi-step tasks&lt;/li&gt;
&lt;li&gt;Use external tools and APIs&lt;/li&gt;
&lt;li&gt;Learn from feedback&lt;/li&gt;
&lt;li&gt;Collaborate with other agents&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Key Architecture Patterns
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. ReAct (Reasoning + Acting)
&lt;/h3&gt;

&lt;p&gt;The ReAct pattern combines reasoning and acting in a loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Observe&lt;/strong&gt; the current state&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reason&lt;/strong&gt; about what to do next&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Act&lt;/strong&gt; using available tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observe&lt;/strong&gt; the result&lt;/li&gt;
&lt;li&gt;Repeat until the goal is achieved&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This pattern is powerful because it allows agents to handle complex, multi-step tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. SOP (Standard Operating Procedure)
&lt;/h3&gt;

&lt;p&gt;SOP agents follow predefined procedures for specific tasks. Think of it as a decision tree:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Define clear steps&lt;/li&gt;
&lt;li&gt;Specify conditions for each branch&lt;/li&gt;
&lt;li&gt;Allow tool usage at each step&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This approach is great for tasks that require consistency and reliability.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Reflection
&lt;/h3&gt;

&lt;p&gt;Reflection agents can self-correct by reviewing their own outputs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Generate a solution&lt;/li&gt;
&lt;li&gt;Critique the solution&lt;/li&gt;
&lt;li&gt;Revise based on feedback&lt;/li&gt;
&lt;li&gt;Repeat until satisfied&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This self-improvement loop leads to higher quality outputs.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Multi-Agent Systems
&lt;/h3&gt;

&lt;p&gt;The most powerful agents work in teams:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Planner&lt;/strong&gt;: Breaks down complex tasks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Executor&lt;/strong&gt;: Performs specific actions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Critic&lt;/strong&gt;: Reviews and provides feedback&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Coordinator&lt;/strong&gt;: Manages communication&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each agent has a specialized role, leading to better outcomes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing the Right Architecture
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ReAct&lt;/td&gt;
&lt;td&gt;Complex reasoning tasks&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SOP&lt;/td&gt;
&lt;td&gt;Repetitive workflows&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reflection&lt;/td&gt;
&lt;td&gt;Quality-critical tasks&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-Agent&lt;/td&gt;
&lt;td&gt;Large-scale projects&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Future of Agent Architecture
&lt;/h2&gt;

&lt;p&gt;As AI advances, we expect to see:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;More sophisticated planning capabilities&lt;/li&gt;
&lt;li&gt;Better tool integration&lt;/li&gt;
&lt;li&gt;Improved memory systems&lt;/li&gt;
&lt;li&gt;Enhanced collaboration between agents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key is choosing the right architecture for your use case.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;AI agent architecture is a rapidly evolving field. By understanding these patterns, you can design more effective and reliable agents.&lt;/p&gt;

&lt;p&gt;What architecture pattern do you find most interesting? Share your thoughts in the comments!&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tags: AI, Agents, Architecture, Machine Learning, AI Design&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>LLM Inference Optimization: Techniques for Faster and Cheaper AI</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Mon, 14 Sep 2026 00:30:07 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/llm-inference-optimization-techniques-for-faster-and-cheaper-ai-54ml</link>
      <guid>https://dev.to/ryan_zhao/llm-inference-optimization-techniques-for-faster-and-cheaper-ai-54ml</guid>
      <description>&lt;h1&gt;
  
  
  LLM Inference Optimization: Techniques for Faster and Cheaper AI
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fni0lqwmoqu714lr45fvc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fni0lqwmoqu714lr45fvc.png" alt="LLM Inference Optimization" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Large Language Models are powerful, but they can be slow and expensive. In this article, we explore practical techniques to optimize LLM inference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Optimize LLM Inference?
&lt;/h2&gt;

&lt;p&gt;As AI applications scale, inference costs and latency become critical bottlenecks. Optimization helps you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reduce response times&lt;/li&gt;
&lt;li&gt;Lower computational costs&lt;/li&gt;
&lt;li&gt;Scale to more users&lt;/li&gt;
&lt;li&gt;Deploy on edge devices&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Key Optimization Techniques
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Quantization
&lt;/h3&gt;

&lt;p&gt;Quantization reduces the precision of model weights:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;INT8&lt;/strong&gt;: 8-bit integers (4x speedup)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;INT4&lt;/strong&gt;: 4-bit integers (8x speedup)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FP8&lt;/strong&gt;: 8-bit floating point&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Trade-off: Slight accuracy loss for massive speed gains.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. KV Cache Optimization
&lt;/h3&gt;

&lt;p&gt;KV Cache stores attention computations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;PagedAttention&lt;/strong&gt;: Memory-efficient caching&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sliding Window&lt;/strong&gt;: Limited context windows&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compression&lt;/strong&gt;: Reduce cache size&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Result: Faster generation for long contexts.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Speculative Decoding
&lt;/h3&gt;

&lt;p&gt;Use a smaller model to draft tokens:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Small model drafts multiple tokens&lt;/li&gt;
&lt;li&gt;Large model verifies in parallel&lt;/li&gt;
&lt;li&gt;Accept or reject drafts&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Speedup: 2-3x without quality loss.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Prompt Optimization
&lt;/h3&gt;

&lt;p&gt;Better prompts mean fewer tokens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compression&lt;/strong&gt;: Remove redundancy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Structure&lt;/strong&gt;: Clear formatting&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Examples&lt;/strong&gt;: Few-shot learning&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Batch Processing
&lt;/h3&gt;

&lt;p&gt;Process multiple requests together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dynamic batching&lt;/li&gt;
&lt;li&gt;Padding optimization&lt;/li&gt;
&lt;li&gt;Memory pooling&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Performance Metrics
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Technique&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Quality&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Quantization&lt;/td&gt;
&lt;td&gt;4x&lt;/td&gt;
&lt;td&gt;75% less&lt;/td&gt;
&lt;td&gt;Minor loss&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KV Cache&lt;/td&gt;
&lt;td&gt;2x&lt;/td&gt;
&lt;td&gt;50% less&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speculative&lt;/td&gt;
&lt;td&gt;2.5x&lt;/td&gt;
&lt;td&gt;60% less&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt Opt&lt;/td&gt;
&lt;td&gt;1.5x&lt;/td&gt;
&lt;td&gt;33% less&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Implementation Tips
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Start with KV Cache (easiest win)&lt;/li&gt;
&lt;li&gt;Add quantization for edge deployment&lt;/li&gt;
&lt;li&gt;Use speculative decoding for throughput&lt;/li&gt;
&lt;li&gt;Optimize prompts for cost savings&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Future
&lt;/h2&gt;

&lt;p&gt;Expect even more optimization techniques:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Hardware-specific kernels&lt;/li&gt;
&lt;li&gt;Dynamic routing&lt;/li&gt;
&lt;li&gt;Neural architecture search&lt;/li&gt;
&lt;li&gt;Hybrid approaches&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Optimization is not a one-size-fits-all solution. Choose techniques based on your priorities: speed, cost, or quality.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fni0lqwmoqu714lr45fvc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fni0lqwmoqu714lr45fvc.png" alt="Optimization Flow" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What optimization technique has worked best for you? Share your experience!&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tags: AI, LLM, Optimization, Machine Learning&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>optimization</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>AI Model Evaluation: Best Practices for Testing and Validation</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Mon, 14 Sep 2026 00:29:37 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/ai-model-evaluation-best-practices-for-testing-and-validation-4nfh</link>
      <guid>https://dev.to/ryan_zhao/ai-model-evaluation-best-practices-for-testing-and-validation-4nfh</guid>
      <description>&lt;h1&gt;
  
  
  AI Model Evaluation: Best Practices for Testing and Validation
&lt;/h1&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwtj15ff2qgtu41h6iln6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwtj15ff2qgtu41h6iln6.png" alt="AI Model Evaluation" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Evaluating AI models is critical for ensuring quality, safety, and reliability. In this article, we explore best practices for model evaluation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Evaluate AI Models?
&lt;/h2&gt;

&lt;p&gt;AI models can make mistakes, show bias, or behave unexpectedly. Evaluation helps you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ensure model quality&lt;/li&gt;
&lt;li&gt;Detect bias and fairness issues&lt;/li&gt;
&lt;li&gt;Verify safety standards&lt;/li&gt;
&lt;li&gt;Measure real-world performance&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Evaluation Framework
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Benchmarks
&lt;/h3&gt;

&lt;p&gt;Standardized tests for model capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MMLU&lt;/strong&gt;: Knowledge and reasoning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HumanEval&lt;/strong&gt;: Code generation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GSM8K&lt;/strong&gt;: Math problem solving&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SuperGLUE&lt;/strong&gt;: Language understanding&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Benchmarks provide objective, comparable metrics.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Red Teaming
&lt;/h3&gt;

&lt;p&gt;Adversarial testing to find weaknesses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Prompt injection&lt;/strong&gt;: Test for security&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Jailbreak&lt;/strong&gt;: Test for safety&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edge cases&lt;/strong&gt;: Test for robustness&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bias detection&lt;/strong&gt;: Test for fairness&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Red teaming reveals vulnerabilities before deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. User Testing
&lt;/h3&gt;

&lt;p&gt;Real-world usage feedback:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A/B testing&lt;/strong&gt;: Compare model versions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User surveys&lt;/strong&gt;: Gather subjective feedback&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Usage analytics&lt;/strong&gt;: Track real patterns&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error analysis&lt;/strong&gt;: Study failure cases&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;User testing provides ground-truth insights.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluation Metrics
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What It Measures&lt;/th&gt;
&lt;th&gt;Importance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy&lt;/td&gt;
&lt;td&gt;Correct predictions&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Response time&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fairness&lt;/td&gt;
&lt;td&gt;Bias detection&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Robustness&lt;/td&gt;
&lt;td&gt;Error handling&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Safety&lt;/td&gt;
&lt;td&gt;Harm prevention&lt;/td&gt;
&lt;td&gt;Critical&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Best Practices
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Multi-dimensional evaluation&lt;/strong&gt;: Test across many dimensions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Continuous testing&lt;/strong&gt;: Evaluate regularly, not just once&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-in-the-loop&lt;/strong&gt;: Combine automated and human review&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Document results&lt;/strong&gt;: Track improvements over time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share findings&lt;/strong&gt;: Learn from each other&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Tools and Frameworks
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MLflow&lt;/strong&gt;: Experiment tracking&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weights &amp;amp; Biases&lt;/strong&gt;: Model monitoring&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepEval&lt;/strong&gt;: Evaluation framework&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LangSmith&lt;/strong&gt;: LLM testing&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Future
&lt;/h2&gt;

&lt;p&gt;Expect more sophisticated evaluation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Automated red teaming&lt;/li&gt;
&lt;li&gt;Real-time monitoring&lt;/li&gt;
&lt;li&gt;Dynamic benchmarks&lt;/li&gt;
&lt;li&gt;Community-driven evaluation&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Model evaluation is not a one-time task. It's an ongoing process that requires multiple approaches.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwtj15ff2qgtu41h6iln6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwtj15ff2qgtu41h6iln6.png" alt="Evaluation Pyramid" width="800" height="575"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What evaluation methods have you found most effective? Share your insights!&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tags: AI, Evaluation, Machine Learning, Testing&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>evaluation</category>
      <category>machinelearning</category>
      <category>testing</category>
    </item>
    <item>
      <title>DeepSeek R1: The Open-Source Reasoning Revolution That Changes Everything</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Sun, 13 Sep 2026 03:47:54 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/deepseek-r1-the-open-source-reasoning-revolution-that-changes-everything-48m6</link>
      <guid>https://dev.to/ryan_zhao/deepseek-r1-the-open-source-reasoning-revolution-that-changes-everything-48m6</guid>
      <description>&lt;h1&gt;
  
  
  DeepSeek R1: The Open-Source Reasoning Revolution That Changes Everything
&lt;/h1&gt;

&lt;h2&gt;
  
  
  The Reasoning Problem
&lt;/h2&gt;

&lt;p&gt;Traditional LLMs generate text token by token, left to right. This autoregressive approach works for simple tasks but struggles with complex reasoning, math, and multi-step logic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The core problem&lt;/strong&gt;: How do you get an LLM to think before answering?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3abvx48q4wpotsw0mxb3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3abvx48q4wpotsw0mxb3.png" alt="DeepSeek R1 MoE Architecture" width="800" height="536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Solution: Mixture of Experts (MoE)
&lt;/h2&gt;

&lt;p&gt;DeepSeek R1 uses a &lt;strong&gt;Mixture of Experts&lt;/strong&gt; architecture combined with &lt;strong&gt;Reinforcement Learning from Reasoning Feedback (RLRF)&lt;/strong&gt; to achieve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fast inference&lt;/strong&gt; — Only activate relevant experts per query&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deep reasoning&lt;/strong&gt; — Chain multiple reasoning steps internally&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open weights&lt;/strong&gt; — Anyone can download and fine-tune&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  How MoE Works
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Input&lt;/strong&gt; arrives at the router&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Router&lt;/strong&gt; selects the top-k experts for this specific query&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Experts&lt;/strong&gt; process in parallel (math, code, logic, science)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Aggregator&lt;/strong&gt; combines outputs into a coherent response&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is dramatically more efficient than activating all parameters for every query.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance Benchmarks
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;DeepSeek R1&lt;/th&gt;
&lt;th&gt;GPT-4&lt;/th&gt;
&lt;th&gt;Claude 3.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Math (AIME)&lt;/td&gt;
&lt;td&gt;79.4%&lt;/td&gt;
&lt;td&gt;83.0%&lt;/td&gt;
&lt;td&gt;81.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding (LiveCode)&lt;/td&gt;
&lt;td&gt;61.2%&lt;/td&gt;
&lt;td&gt;65.0%&lt;/td&gt;
&lt;td&gt;63.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning (GPQA)&lt;/td&gt;
&lt;td&gt;74.8%&lt;/td&gt;
&lt;td&gt;78.0%&lt;/td&gt;
&lt;td&gt;76.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Key insight&lt;/strong&gt;: Open-source models are now competitive with and sometimes surpassing closed models on reasoning tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why This Matters
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Accessibility&lt;/strong&gt; — Anyone can download and run R1 locally&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transparency&lt;/strong&gt; — Open weights mean open reasoning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Innovation&lt;/strong&gt; — Researchers can fine-tune for specific domains&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost&lt;/strong&gt; — Open models reduce dependency on expensive APIs&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Road Ahead
&lt;/h2&gt;

&lt;p&gt;With MoE plus RLRF, the gap between open and closed models continues to narrow. The next frontier? &lt;strong&gt;Multi-modal reasoning&lt;/strong&gt; — combining text, vision, and audio into unified reasoning pipelines.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What reasoning benchmarks matter most to you? Share your thoughts below.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>deepseek</category>
      <category>moe</category>
      <category>reasoning</category>
    </item>
    <item>
      <title>Multi-Modal AI: From Text to Vision and Beyond — The Unified Future</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Sun, 13 Sep 2026 03:47:24 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/multi-modal-ai-from-text-to-vision-and-beyond-the-unified-future-4c36</link>
      <guid>https://dev.to/ryan_zhao/multi-modal-ai-from-text-to-vision-and-beyond-the-unified-future-4c36</guid>
      <description>&lt;h1&gt;
  
  
  Multi-Modal AI: From Text to Vision and Beyond — The Unified Future
&lt;/h1&gt;

&lt;h2&gt;
  
  
  The Single-Modality Limit
&lt;/h2&gt;

&lt;p&gt;For years, AI models were &lt;strong&gt;single-modality&lt;/strong&gt; — text-only, image-only, or audio-only. This created silos:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A text model cannot see images&lt;/li&gt;
&lt;li&gt;An image model cannot hear audio&lt;/li&gt;
&lt;li&gt;Each modality required separate training&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The problem&lt;/strong&gt;: Real-world understanding is inherently multi-modal.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhk9yoggb9dfnu20jg9d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxhk9yoggb9dfnu20jg9d.png" alt="Multi-Modal AI Architecture" width="800" height="536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Breakthrough: Unified Encoders
&lt;/h2&gt;

&lt;p&gt;Modern multi-modal models use a &lt;strong&gt;shared latent space&lt;/strong&gt; — a single representation that encodes text, images, audio, and video into a common format.&lt;/p&gt;

&lt;h3&gt;
  
  
  How It Works
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Each modality&lt;/strong&gt; has its own encoder (text tokenizer, image CNN, audio encoder)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Projections&lt;/strong&gt; map each encoder output into the shared latent space&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A unified transformer&lt;/strong&gt; processes all modalities together&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Task heads&lt;/strong&gt; generate outputs in any modality&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Why This Matters
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cross-modal retrieval&lt;/strong&gt;: Search images with text queries&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visual question answering&lt;/strong&gt;: Ask questions about images&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image captioning&lt;/strong&gt;: Generate descriptions from visual input&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Text-to-image generation&lt;/strong&gt;: Create visuals from text prompts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Video understanding&lt;/strong&gt;: Combine temporal plus visual plus audio signals&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Real-World Applications
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Domain&lt;/th&gt;
&lt;th&gt;Application&lt;/th&gt;
&lt;th&gt;Impact&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Healthcare&lt;/td&gt;
&lt;td&gt;Medical image plus report analysis&lt;/td&gt;
&lt;td&gt;Better diagnostics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Education&lt;/td&gt;
&lt;td&gt;Visual plus text learning&lt;/td&gt;
&lt;td&gt;Personalized tutoring&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Robotics&lt;/td&gt;
&lt;td&gt;Vision plus language plus action&lt;/td&gt;
&lt;td&gt;Autonomous navigation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content Creation&lt;/td&gt;
&lt;td&gt;Text-to-video plus audio&lt;/td&gt;
&lt;td&gt;Creative automation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Future: True Multimodal Intelligence
&lt;/h2&gt;

&lt;p&gt;The next generation will feature:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Real-time multi-modal streaming&lt;/strong&gt; — Process video, audio, and text simultaneously&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-modal generation&lt;/strong&gt; — Generate video from text, audio from images&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Embodied AI&lt;/strong&gt; — Robots that see, hear, speak, and act&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-level understanding&lt;/strong&gt; — Context-aware across all sensory modalities&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Which multi-modal application excites you most? Let us know in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>multimodal</category>
      <category>vision</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>AI Safety and Alignment: Building Trustworthy Agents That Do Not Fail You</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Sun, 13 Sep 2026 03:46:58 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/ai-safety-and-alignment-building-trustworthy-agents-that-do-not-fail-you-1p6m</link>
      <guid>https://dev.to/ryan_zhao/ai-safety-and-alignment-building-trustworthy-agents-that-do-not-fail-you-1p6m</guid>
      <description>&lt;h1&gt;
  
  
  AI Safety and Alignment: Building Trustworthy Agents That Do Not Fail You
&lt;/h1&gt;

&lt;h2&gt;
  
  
  The Trust Problem
&lt;/h2&gt;

&lt;p&gt;As AI agents become more capable, &lt;strong&gt;trustworthiness&lt;/strong&gt; becomes the critical differentiator. A model that is smart but unreliable is worse than useless — it is dangerous.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8s1dkqnbuzo3vx6di6wi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8s1dkqnbuzo3vx6di6wi.png" alt="AI Safety Alignment Pyramid" width="800" height="795"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Safety Pyramid
&lt;/h2&gt;

&lt;p&gt;Building trustworthy AI requires &lt;strong&gt;layered defense&lt;/strong&gt;:&lt;/p&gt;

&lt;h3&gt;
  
  
  Level 1: Technical Robustness
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Error handling and edge case coverage&lt;/li&gt;
&lt;li&gt;Input validation and sanitization&lt;/li&gt;
&lt;li&gt;Graceful degradation under stress&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Level 2: Interpretability
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Model transparency and explainability&lt;/li&gt;
&lt;li&gt;Activation visualization and probing&lt;/li&gt;
&lt;li&gt;Mechanistic interpretability research&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Level 3: Content Safety
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Harmful output filtering&lt;/li&gt;
&lt;li&gt;Toxicity detection and prevention&lt;/li&gt;
&lt;li&gt;Bias mitigation and fairness&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Level 4: Instruction Following
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Accurate task completion&lt;/li&gt;
&lt;li&gt;Refusal of harmful requests&lt;/li&gt;
&lt;li&gt;Context-aware compliance&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Level 5: Value Alignment
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Human preference learning (RLHF)&lt;/li&gt;
&lt;li&gt;Constitutional AI principles&lt;/li&gt;
&lt;li&gt;Multi-stakeholder value balancing&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Level 6: Robustness
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Adversarial attack defense&lt;/li&gt;
&lt;li&gt;Distribution shift handling&lt;/li&gt;
&lt;li&gt;Out-of-distribution generalization&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Each Layer Matters
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Without Level 1&lt;/strong&gt;, the system crashes on edge cases.&lt;br&gt;
&lt;strong&gt;Without Level 2&lt;/strong&gt;, you cannot debug failures.&lt;br&gt;
&lt;strong&gt;Without Level 3&lt;/strong&gt;, the system generates harmful content.&lt;br&gt;
&lt;strong&gt;Without Level 4&lt;/strong&gt;, the system ignores user intent.&lt;br&gt;
&lt;strong&gt;Without Level 5&lt;/strong&gt;, the system pursues wrong goals.&lt;br&gt;
&lt;strong&gt;Without Level 6&lt;/strong&gt;, the system fails in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Safety Measures
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Red teaming&lt;/strong&gt; — Actively try to break your system&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluation benchmarks&lt;/strong&gt; — Measure safety, not just accuracy&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-in-the-loop&lt;/strong&gt; — Keep humans in the decision loop&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitoring&lt;/strong&gt; — Track model behavior in production&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rollback plans&lt;/strong&gt; — Have kill switches ready&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The Bottom Line
&lt;/h2&gt;

&lt;p&gt;Safety is not a feature — it is a &lt;strong&gt;foundation&lt;/strong&gt;. Every AI system, regardless of capability, must be built on these layered principles.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What safety measures have you implemented? Share your experiences below.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>safety</category>
      <category>alignment</category>
      <category>ethics</category>
    </item>
    <item>
      <title>AI Agent Tool Mastery: How Modern Agents Choose and Use Tools Effectively</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Fri, 11 Sep 2026 02:17:32 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/ai-agent-tool-mastery-how-modern-agents-choose-and-use-tools-effectively-2hkg</link>
      <guid>https://dev.to/ryan_zhao/ai-agent-tool-mastery-how-modern-agents-choose-and-use-tools-effectively-2hkg</guid>
      <description>&lt;h1&gt;
  
  
  AI Agent Tool Mastery: How Modern Agents Choose and Use Tools Effectively
&lt;/h1&gt;

&lt;h2&gt;
  
  
  The Tool Selection Problem
&lt;/h2&gt;

&lt;p&gt;When an AI agent receives a complex request like "Find the weather, book a flight, and send a confirmation email," it must do more than just generate text. It needs to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Understand intent&lt;/strong&gt; — What does the user actually want?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Match tools&lt;/strong&gt; — Which tools can fulfill each part of the request?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chain operations&lt;/strong&gt; — How do results from one tool feed into the next?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handle failures&lt;/strong&gt; — What happens when a tool fails or returns unexpected data?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dacyzywzr0pxfs3l4j4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dacyzywzr0pxfs3l4j4.png" alt="Agent Tool Selection Architecture" width="800" height="536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Architecture of Tool Selection
&lt;/h2&gt;

&lt;p&gt;Modern AI agents use a multi-stage pipeline to select and execute tools:&lt;/p&gt;

&lt;h3&gt;
  
  
  Stage 1: Intent Analysis
&lt;/h3&gt;

&lt;p&gt;The agent parses the user request to identify:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Primary goals (e.g., "get weather," "book flight")&lt;/li&gt;
&lt;li&gt;Secondary goals (e.g., "send confirmation")&lt;/li&gt;
&lt;li&gt;Constraints (e.g., "tomorrow morning," "economy class")&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Stage 2: Tool Matching
&lt;/h3&gt;

&lt;p&gt;Each identified goal is matched against available tools using:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Semantic similarity (does the tool description match the intent?)&lt;/li&gt;
&lt;li&gt;Historical performance (has this tool succeeded for similar requests?)&lt;/li&gt;
&lt;li&gt;Availability (is the tool currently accessible?)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Stage 3: Execution Chaining
&lt;/h3&gt;

&lt;p&gt;Results from one tool may trigger additional tool calls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Weather data → Flight search with date constraints&lt;/li&gt;
&lt;li&gt;Flight results → Booking API with passenger details&lt;/li&gt;
&lt;li&gt;Booking confirmation → Email tool with itinerary&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi17z03vxvxqme65wfzoq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi17z03vxvxqme65wfzoq.png" alt="Tool Performance Comparison" width="800" height="531"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance Metrics That Matter
&lt;/h2&gt;

&lt;p&gt;Not all tools are created equal. Here's how different tool categories perform:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Search&lt;/th&gt;
&lt;th&gt;Weather&lt;/th&gt;
&lt;th&gt;Booking&lt;/th&gt;
&lt;th&gt;Email&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy&lt;/td&gt;
&lt;td&gt;92%&lt;/td&gt;
&lt;td&gt;88%&lt;/td&gt;
&lt;td&gt;76%&lt;/td&gt;
&lt;td&gt;95%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speed&lt;/td&gt;
&lt;td&gt;78&lt;/td&gt;
&lt;td&gt;95&lt;/td&gt;
&lt;td&gt;65&lt;/td&gt;
&lt;td&gt;90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reliability&lt;/td&gt;
&lt;td&gt;88&lt;/td&gt;
&lt;td&gt;92&lt;/td&gt;
&lt;td&gt;70&lt;/td&gt;
&lt;td&gt;96&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Key Insight&lt;/strong&gt;: High accuracy doesn't always mean high speed. Agents must balance these trade-offs based on context.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evolution of Agent Tools
&lt;/h2&gt;

&lt;p&gt;The capability of agent tools has evolved dramatically:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;2020&lt;/strong&gt;: Basic functions — simple API calls with fixed parameters&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2021&lt;/strong&gt;: Static tools — predefined toolsets with limited flexibility&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2022&lt;/strong&gt;: Dynamic selection — agents could choose tools based on context&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2023&lt;/strong&gt;: Multi-tool chaining — agents could sequence multiple tool calls&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2024&lt;/strong&gt;: Self-improving tools — tools that learn from past failures&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2025&lt;/strong&gt;: Autonomous tool creation — agents that write their own tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;2026&lt;/strong&gt;: Adaptive ecosystems — tools that evolve based on user behavior&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8gmi5ytxzwiqo4trcrdr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8gmi5ytxzwiqo4trcrdr.png" alt="Agent Tool Evolution Timeline" width="799" height="339"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Example: A Day in the Life of an Agent
&lt;/h2&gt;

&lt;p&gt;Let's walk through a real-world scenario:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;User Request&lt;/strong&gt;: "I need to travel to Tokyo next week for a conference. Find flights, book the cheapest option, and send the itinerary to my assistant."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Agent Actions&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Search API&lt;/strong&gt; — Finds conference dates and location&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weather API&lt;/strong&gt; — Checks Tokyo weather for next week&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flight Search API&lt;/strong&gt; — Finds available flights&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Booking API&lt;/strong&gt; — Books the cheapest flight&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Email API&lt;/strong&gt; — Sends itinerary to assistant&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Calendar API&lt;/strong&gt; — Adds event to user's calendar&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Result&lt;/strong&gt;: A fully automated travel booking experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Pitfalls and How to Avoid Them
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Pitfall 1: Over-Tooling
&lt;/h3&gt;

&lt;p&gt;Using too many tools for simple tasks slows down responses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution&lt;/strong&gt;: Implement a complexity threshold — only chain tools when necessary.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall 2: Ignoring Tool Failures
&lt;/h3&gt;

&lt;p&gt;When a tool fails, agents should have fallback strategies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution&lt;/strong&gt;: Implement retry logic, alternative tools, and user notifications.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pitfall 3: No Performance Monitoring
&lt;/h3&gt;

&lt;p&gt;Without tracking tool performance, agents can't optimize.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Solution&lt;/strong&gt;: Log all tool calls, measure success rates, and adjust weights.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Future of Agent Tools
&lt;/h2&gt;

&lt;p&gt;As AI agents become more sophisticated, tool ecosystems will evolve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-optimizing tools&lt;/strong&gt; that adjust parameters based on usage patterns&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-platform tools&lt;/strong&gt; that work across different services seamlessly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User-adaptive tools&lt;/strong&gt; that learn individual preferences over time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Collaborative tools&lt;/strong&gt; that work together across agent boundaries&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;AI agent tool mastery isn't just about having access to many tools — it's about selecting the right tools, chaining them effectively, and continuously improving based on performance data.&lt;/p&gt;

&lt;p&gt;The agents that thrive will be those that treat tools not as static resources, but as dynamic components of an adaptive ecosystem.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What tool challenges have you encountered with AI agents? Share your experiences in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>tools</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>HarnessDev: Self-Evolving Agent Frameworks — How LLMs Build Their Own Infrastructure</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Fri, 11 Sep 2026 02:07:56 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/harnessdev-self-evolving-agent-frameworks-how-llms-build-their-own-infrastructure-332k</link>
      <guid>https://dev.to/ryan_zhao/harnessdev-self-evolving-agent-frameworks-how-llms-build-their-own-infrastructure-332k</guid>
      <description>&lt;h1&gt;
  
  
  HarnessDev: How LLMs Are Building Their Own Agent Frameworks
&lt;/h1&gt;

&lt;h2&gt;
  
  
  ByteDance's Breakthrough in Self-Evolving Agent Systems
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Published: September 10, 2026 | Reading time: 12 minutes&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Revolutionary Research
&lt;/h2&gt;

&lt;p&gt;Last week, ByteDance's Seed team, in collaboration with Singapore University of Technology and Design, Georgia Tech, and other institutions, released &lt;strong&gt;HarnessDev&lt;/strong&gt; — a groundbreaking research project that answers a fundamental question in AI agent engineering:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can LLMs create their own Agent Harnesses and continuously improve them based on task feedback?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The answer is a resounding &lt;strong&gt;yes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvw13fsvym5fds334ersn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvw13fsvym5fds334ersn.png" alt="Agent Loop Architecture" width="800" height="531"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What Is an Agent Harness?
&lt;/h2&gt;

&lt;p&gt;An Agent Harness is the core control system that drives an AI agent. It includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Task execution loops&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool selection and parameter constraints&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context management&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;State tracking&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Result verification&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure recovery mechanisms&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Think of it as the "operating system" for an AI agent — without it, the agent is just a language model with no structure or direction.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Two-Phase Process
&lt;/h2&gt;

&lt;p&gt;HarnessDev divides the agent development process into two phases:&lt;/p&gt;

&lt;h3&gt;
  
  
  Phase 1: Creation
&lt;/h3&gt;

&lt;p&gt;Starting from a &lt;strong&gt;Weak Seed Harness&lt;/strong&gt; (a minimal framework with basic I/O capabilities), the LLM builds a complete agent harness by adding control logic for:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Execution&lt;/strong&gt; — Task loops, planning, scheduling, and stopping conditions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tools&lt;/strong&gt; — Tool selection, parameter constraints, input/output handling, and error processing&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context&lt;/strong&gt; — Organization of task information, code, history, and constraints&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;State&lt;/strong&gt; — Current goals, progress tracking, attempt records, and failure information&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lifecycle&lt;/strong&gt; — Timeout handling, recovery mechanisms, and task cleanup&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verification&lt;/strong&gt; — Testing, result checking, completion determination, and logging&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Phase 2: Evolution
&lt;/h3&gt;

&lt;p&gt;Using the created harness as a starting point, the LLM continuously adjusts it based on downstream task feedback, evaluating performance on held-out tasks.&lt;/p&gt;




&lt;h2&gt;
  
  
  Key Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. LLMs Can Build Effective Harnesses
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;18 Code Harnesses were created, adding a total of &lt;strong&gt;17,111 lines of code&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt; required the fewest changes (1,006 lines) but achieved the highest score on Terminal-Bench 2.1 (&lt;strong&gt;68.8&lt;/strong&gt;)&lt;/li&gt;
&lt;li&gt;All 18 harnesses implemented Execution Loops; Tools, Lifecycle, and Verification had high completion rates&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Not All Code Is Used
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Out of 108 component instances in Code Harnesses, only &lt;strong&gt;72&lt;/strong&gt; were observed running in real tasks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;18 components&lt;/strong&gt; (all from State and Memory) never appeared in actual execution&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;26,679 task trajectories&lt;/strong&gt; recorded zero checkpoint events, despite some harnesses implementing checkpoint logic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This reveals a critical insight: &lt;strong&gt;implementing a mechanism doesn't mean it's actually used&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Verification Gap
&lt;/h3&gt;

&lt;p&gt;In self-evaluation, &lt;strong&gt;Opus&lt;/strong&gt; found that out of 100 runs, the harness reported success 99 times, but only &lt;strong&gt;48&lt;/strong&gt; were actually correct. This led to the addition of a &lt;strong&gt;Completion Check&lt;/strong&gt; mechanism.&lt;/p&gt;

&lt;p&gt;Similarly, in Data tasks, &lt;strong&gt;441 out of 2,325&lt;/strong&gt; executions produced degraded commits, but the harness failed to detect them.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cross-Model Adaptation Challenges
&lt;/h3&gt;

&lt;p&gt;When switching executors, performance varies significantly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Opus's SWE-Pro Harness&lt;/strong&gt;: 69.3 (Self-Eval) → 33.0 (Gemini executor)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Qwen Harness&lt;/strong&gt;: Improved by 17.6 points on BrowseComp when using Gemini&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This shows that harnesses become &lt;strong&gt;executor-specific&lt;/strong&gt; over time, requiring re-tuning when switching models.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Evolution Limitations
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;5 evolution trajectories improved on the visible feedback set&lt;/li&gt;
&lt;li&gt;However, improvements on &lt;strong&gt;held-out tasks&lt;/strong&gt; were much smaller (average 3.11 points)&lt;/li&gt;
&lt;li&gt;Only &lt;strong&gt;53.1%&lt;/strong&gt; of version changes showed consistent direction between feedback and held-out sets&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evaluation fluctuation&lt;/strong&gt; is approximately ±4.75 points, making it hard to distinguish real improvements from noise&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The Six Control Capabilities
&lt;/h2&gt;

&lt;p&gt;HarnessDev categorizes agent control into six capabilities:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Execution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Task loops, planning, scheduling&lt;/td&gt;
&lt;td&gt;When to stop, how to plan&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tools&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tool selection, parameters, error handling&lt;/td&gt;
&lt;td&gt;Which API to call, how to handle errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Organization of task info and history&lt;/td&gt;
&lt;td&gt;What context to provide the LLM&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;State&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Goals, progress, failure records&lt;/td&gt;
&lt;td&gt;Current state, attempt history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Lifecycle&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Timeout, recovery, cleanup&lt;/td&gt;
&lt;td&gt;Handle failures, recover from errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Testing, result checking, logging&lt;/td&gt;
&lt;td&gt;Verify results, log outcomes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Execution Cost Analysis
&lt;/h2&gt;

&lt;p&gt;Different execution strategies significantly impact token consumption:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPT-5.5 Harness&lt;/strong&gt;: 29.3M tokens, medal rate 19.1&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4 Harness&lt;/strong&gt;: 208.4M tokens, score 19.6&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Same performance, 7x token difference!&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Across the entire MLE-bench experiment, token overhead varied by &lt;strong&gt;19x&lt;/strong&gt; between different harnesses.&lt;/p&gt;

&lt;p&gt;This highlights the importance of &lt;strong&gt;cost-aware design&lt;/strong&gt; — a small performance improvement may not justify a large token increase.&lt;/p&gt;




&lt;h2&gt;
  
  
  Code Example: Creating a Simple Harness
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Weak Seed Harness (minimal framework)
&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;WeakSeed&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;initial&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;history&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;running&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="c1"&gt;# Basic execution loop
&lt;/span&gt;        &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;running&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;action&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;plan_action&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;call_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;record_history&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_complete&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
                &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;completed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;plan_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# LLM generates action
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Execute tool
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;is_complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="c1"&gt;# Check completion
&lt;/span&gt;        &lt;span class="k"&gt;pass&lt;/span&gt;

&lt;span class="c1"&gt;# LLM enhances this with control logic
# (Execution, Tools, Context, State, Lifecycle, Verification)
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Why This Matters
&lt;/h2&gt;

&lt;h3&gt;
  
  
  For AI Researchers
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Shows LLMs can self-improve their own execution frameworks&lt;/li&gt;
&lt;li&gt;Highlights the gap between &lt;strong&gt;implemented&lt;/strong&gt; and &lt;strong&gt;actually used&lt;/strong&gt; mechanisms&lt;/li&gt;
&lt;li&gt;Reveals the importance of &lt;strong&gt;verification&lt;/strong&gt; and &lt;strong&gt;cross-model adaptation&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  For AI Practitioners
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Demonstrates the value of &lt;strong&gt;structured agent design&lt;/strong&gt; over pure memory&lt;/li&gt;
&lt;li&gt;Shows the importance of &lt;strong&gt;cost-aware optimization&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Highlights the need for &lt;strong&gt;robust verification&lt;/strong&gt; mechanisms&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  For the Industry
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Represents a step toward &lt;strong&gt;self-evolving AI systems&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Shows the potential for &lt;strong&gt;automated agent development&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Highlights challenges in &lt;strong&gt;generalization&lt;/strong&gt; and &lt;strong&gt;adaptation&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;HarnessDev represents a significant step toward self-evolving AI agents. However, several challenges remain:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Implementation vs. Usage Gap&lt;/strong&gt; — Not all implemented mechanisms are actually used&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-Model Adaptation&lt;/strong&gt; — Harnesses become executor-specific&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evolution Limitations&lt;/strong&gt; — Improvements don't always generalize to new tasks&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost Awareness&lt;/strong&gt; — Performance gains may come with disproportionate token costs&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The research provides valuable insights for building more robust, efficient, and self-improving AI agents.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article is based on research published by ByteDance Seed team on September 8, 2026. Paper: arXiv:2609.01437 | Project: self-developing-agents.github.io&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>research</category>
      <category>bytecloud</category>
    </item>
    <item>
      <title>Procedural Graphs: Self-Improving LLM Agent Execution Structures</title>
      <dc:creator>ryan2run</dc:creator>
      <pubDate>Fri, 11 Sep 2026 02:05:56 +0000</pubDate>
      <link>https://dev.to/ryan_zhao/procedural-graphs-self-improving-llm-agent-execution-structures-ba2</link>
      <guid>https://dev.to/ryan_zhao/procedural-graphs-self-improving-llm-agent-execution-structures-ba2</guid>
      <description>&lt;h1&gt;
  
  
  Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
&lt;/h1&gt;

&lt;h2&gt;
  
  
  When AI Agents Start Writing Their Own "Brain Circuits"
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Published: September 10, 2026 | Reading time: 12 minutes&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Revolutionary Research
&lt;/h2&gt;

&lt;p&gt;On September 9, 2026, researchers &lt;strong&gt;Yuxing Lu&lt;/strong&gt;, &lt;strong&gt;Yicheng Chen&lt;/strong&gt;, and &lt;strong&gt;Shanchan Wu&lt;/strong&gt; published a groundbreaking paper on &lt;strong&gt;Procedural Graphs&lt;/strong&gt; — a self-evolving execution structure for LLM agents that can literally rewrite its own "brain circuits."&lt;/p&gt;

&lt;p&gt;Paper: arXiv:2609.08593&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjevyypij6gpvv69qst9v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjevyypij6gpvv69qst9v.png" alt="Self-Evolving Graph" width="800" height="531"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem with Current LLM Agents
&lt;/h2&gt;

&lt;p&gt;Today's LLM agents (like AutoGPT, LangChain agents) work like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Maintain a growing &lt;strong&gt;memory&lt;/strong&gt; of past thoughts, observations, actions&lt;/li&gt;
&lt;li&gt;Generate next action based on this memory&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No fixed structure&lt;/strong&gt;, no predefined流程&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This works for simple tasks, but breaks down for complex ones:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Lost goals&lt;/strong&gt;: Forget what they're supposed to do in long interactions&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Action mismatch&lt;/strong&gt;: Call tools in wrong order (e.g., analyze before searching)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repeated labor&lt;/strong&gt;: Try same ineffective operations repeatedly&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No planning&lt;/strong&gt;: No global view, step-by-step navigation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Analogy&lt;/strong&gt;: Like a chef without a recipe — overwhelmed by complex dishes.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Solution: Procedural Graphs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Procedural Graphs&lt;/strong&gt; organize procedural knowledge ("how to do") just like Knowledge Graphs organize factual knowledge ("what is"):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Structure&lt;/th&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge Graph&lt;/td&gt;
&lt;td&gt;(Entity, Relation, Entity)&lt;/td&gt;
&lt;td&gt;What is?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Procedural Graph&lt;/td&gt;
&lt;td&gt;(Procedure, Relation, Procedure)&lt;/td&gt;
&lt;td&gt;How to?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Core Components
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Nodes (程序步骤)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Each node represents a procedural step&lt;/li&gt;
&lt;li&gt;Contains: description, expected I/O, success/failure conditions&lt;/li&gt;
&lt;li&gt;Like a box in a flowchart&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Edges (关系)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Represent relationships between steps&lt;/li&gt;
&lt;li&gt;Types: Sequential ("then"), Conditional ("if...then"), Parallel ("simultaneously")&lt;/li&gt;
&lt;li&gt;Like arrows in a flowchart&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Attributes (属性)&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Execution probability, average time, success rate, common error patterns&lt;/li&gt;
&lt;li&gt;Update with execution experience&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Works: A Concrete Example
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Task&lt;/strong&gt;: "Book a cheap flight from Beijing to Shanghai, departing tomorrow."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Procedural Graph&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Start]
   ↓
[Search Flight Info]
   ↓
[Filter Low-Cost Options]
   ↓
[Check Seat Availability]
   ↓
[If Available: Fill Passenger Info]
   ↓
[Select Payment Method]
   ↓
[Confirm Order]
   ↓
[End]
   ↓
[If Unavailable: Return to Filter]
   ↓
[If No Satisfying Result: Expand Search]
   ↓
[Search Nearby Airports]
   ↓
[Return to Filter]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a &lt;strong&gt;directed graph&lt;/strong&gt; with branches, loops, and conditionals — far more powerful than linear memory.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Self-Evolution Mechanism
&lt;/h2&gt;

&lt;p&gt;The most amazing capability is &lt;strong&gt;self-evolution&lt;/strong&gt;:&lt;/p&gt;

&lt;h3&gt;
  
  
  Evolution Cycle
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Collect trajectories&lt;/strong&gt;: Record complete execution paths&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare analysis&lt;/strong&gt;: Contrast failed trajectories with successful ones&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Identify differences&lt;/strong&gt;: Find where things went wrong&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Generate edits&lt;/strong&gt;: LLM Refiner proposes modifications&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify and retain&lt;/strong&gt;: Test on validation set&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Three Types of Edits
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Topology Edits&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add new nodes (new steps)&lt;/li&gt;
&lt;li&gt;Remove redundant nodes&lt;/li&gt;
&lt;li&gt;Add/modify edges (change flow structure)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Attribute Edits&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Update node success rate statistics&lt;/li&gt;
&lt;li&gt;Adjust condition thresholds&lt;/li&gt;
&lt;li&gt;Update execution probabilities&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Content Edits&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Modify node descriptions for accuracy&lt;/li&gt;
&lt;li&gt;Update guidance language for effectiveness&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  From Skeleton to Maturity
&lt;/h2&gt;

&lt;p&gt;Researchers tested three initialization methods:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Initialization&lt;/th&gt;
&lt;th&gt;Evolution Speed&lt;/th&gt;
&lt;th&gt;Final Performance&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Empty Graph&lt;/strong&gt; (only "Start" node)&lt;/td&gt;
&lt;td&gt;Slower&lt;/td&gt;
&lt;td&gt;Close to others&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Minimal Skeleton&lt;/strong&gt; (basic manual nodes)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fastest&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Best&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Expert Prior&lt;/strong&gt; (human-designed)&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;td&gt;Repairable if flawed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Amazing Discovery&lt;/strong&gt;: Even starting from a &lt;strong&gt;flawed expert prior&lt;/strong&gt;, the evolution mechanism can "repair" it to achieve good performance.&lt;/p&gt;

&lt;p&gt;This shows &lt;strong&gt;robustness&lt;/strong&gt; — doesn't require perfect initial design.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Graphs Beat Memory
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Pure Memory&lt;/th&gt;
&lt;th&gt;Workflow Memory&lt;/th&gt;
&lt;th&gt;Procedural Graph&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Structure&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Case-level abstract&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Procedure-level&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Generalization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Good&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Explainability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Poor&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Good&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Evolution&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Strong&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Efficiency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;High&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Key Advantage&lt;/strong&gt;: Procedural Graphs abstract the &lt;strong&gt;general flow&lt;/strong&gt; for a class of tasks, not just specific past cases.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Analogy&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Workflow Memory = Remember "Last time I made Mapo Tofu, I stir-fried meat first, then added bean paste"&lt;/li&gt;
&lt;li&gt;Procedural Graph = Understand "General stir-fry flow: Heat pan → Add oil → Stir-fry main ingredient → Season → Serve"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The latter generalizes to &lt;strong&gt;any stir-fry&lt;/strong&gt;; the latter can only repeat Mapo Tofu.&lt;/p&gt;




&lt;h2&gt;
  
  
  Experimental Results
&lt;/h2&gt;

&lt;h3&gt;
  
  
  WebShop (Web Shopping)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Task: Purchase items on e-commerce sites based on natural language instructions&lt;/li&gt;
&lt;li&gt;Graph vs Memory: &lt;strong&gt;15-25% success rate improvement&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Evolved Graph vs Initial: &lt;strong&gt;10-20% improvement&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  ALFWorld (Home Tasks)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Task: Execute daily tasks in simulated home environment&lt;/li&gt;
&lt;li&gt;Graph helps remember complex object interaction sequences&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  HotPotQA (Multi-hop QA)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Task: Multi-step information retrieval and reasoning&lt;/li&gt;
&lt;li&gt;Graph optimizes retrieval strategy and evidence integration&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Tool Use Tasks
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Task: Combine multiple APIs to complete complex goals&lt;/li&gt;
&lt;li&gt;Graph ensures correct tool call order and parameter settings&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Key Findings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Finding 1: Cross-LLM Generalization
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Graph evolved on one LLM (e.g., GPT-4) can transfer to another (e.g., Claude or Llama)&lt;/li&gt;
&lt;li&gt;Shows graphs capture &lt;strong&gt;task structure&lt;/strong&gt;, not model-specific traits&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Finding 2: Few-Shot Advantage
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Effective even with &lt;strong&gt;few examples (&amp;lt;10)&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Pure memory baseline drops sharply with few samples&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Finding 3: Expert Prior Repairability
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Can repair flawed human-designed starting points&lt;/li&gt;
&lt;li&gt;Lowers deployment threshold — no perfect initial design needed&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Deeper Insights
&lt;/h2&gt;

&lt;h3&gt;
  
  
  From Connectionism to Symbolism
&lt;/h3&gt;

&lt;p&gt;Procedural Graphs represent an important trend: &lt;strong&gt;neural-network + symbolic-structure fusion&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deep Learning&lt;/strong&gt; (connectionism): Good at learning patterns from data, lacks explicit reasoning structure&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Symbolic AI&lt;/strong&gt;: Good at logical reasoning and structured knowledge, lacks learning from data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Graph combines both&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LLM provides semantic understanding&lt;/li&gt;
&lt;li&gt;Graph structure provides procedural constraints&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is a &lt;strong&gt;Neuro-Symbolic&lt;/strong&gt; architecture — possibly a key path to more reliable AI.&lt;/p&gt;

&lt;h3&gt;
  
  
  Biological Intelligence Analogy
&lt;/h3&gt;

&lt;p&gt;Human brain similarly combines two systems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;System 1&lt;/strong&gt; (fast, intuitive, pattern-matching): Like LLM generation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;System 2&lt;/strong&gt; (slow, logical, rule-based): Like Procedural Graph execution&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The graph is like giving LLMs a &lt;strong&gt;System 2&lt;/strong&gt; — an explicit, checkable, fixable execution controller.&lt;/p&gt;




&lt;h2&gt;
  
  
  Code Example: Creating a Procedural Graph
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;procedural_graph&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ProceduralGraph&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Edge&lt;/span&gt;

&lt;span class="c1"&gt;# Create a simple graph for a coding task
&lt;/span&gt;&lt;span class="n"&gt;graph&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ProceduralGraph&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Add nodes
&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Start&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Begin task&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;search&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Search Info&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Search for relevant information&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;analyze&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze Data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze collected data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;code&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write Code&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Generate code solution&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;verify&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Verify Result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Test and verify output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;end&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Node&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;End&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Task complete&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Add edges
&lt;/span&gt;&lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;search&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sequential&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;search&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;analyze&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;conditional&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;condition&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;info_found&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;analyze&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sequential&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sequential&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;conditional&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;condition&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;passed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_edge&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;code&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;conditional&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;condition&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# Loop back
&lt;/span&gt;
&lt;span class="c1"&gt;# Navigate the graph
&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt;
&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_local_subgraph&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;action&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;generate_action&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;action&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;current&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;graph&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;navigate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Architecture Overview
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Key Components
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Graph Store&lt;/strong&gt;: Stores nodes, edges, and attributes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Navigator&lt;/strong&gt;: Determines next node based on current state&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refiner&lt;/strong&gt;: Proposes graph edits based on trajectory feedback&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verifier&lt;/strong&gt;: Tests graph modifications on validation set&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Evolution Process
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Collect execution trajectories&lt;/li&gt;
&lt;li&gt;Compare successful vs. failed paths&lt;/li&gt;
&lt;li&gt;Identify differences and generate edits&lt;/li&gt;
&lt;li&gt;Verify edits on validation set&lt;/li&gt;
&lt;li&gt;Apply successful edits to production graph&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  Software Engineering Perspective
&lt;/h2&gt;

&lt;p&gt;Procedural Graphs introduce key concepts:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Separation of Concerns
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"What to do"&lt;/strong&gt; (task understanding): LLM&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"How to do"&lt;/strong&gt; (execution flow): Graph&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"How to improve"&lt;/strong&gt; (flow optimization): Refiner&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Version Control
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Every graph edit is recorded&lt;/li&gt;
&lt;li&gt;Can rollback to previous versions&lt;/li&gt;
&lt;li&gt;Can compare performance across versions&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Testability
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Graph can be independently tested on validation sets&lt;/li&gt;
&lt;li&gt;Edit effects can be quantified&lt;/li&gt;
&lt;li&gt;Avoids "black-box optimization" uncertainty&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The story of Procedural Graphs is essentially a story about &lt;strong&gt;organization&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Information alone has no value. Only when organized into useful structures — recipes, flowcharts, algorithms, organizational charts — can it guide action, produce results, and continuously improve.&lt;/p&gt;

&lt;p&gt;LLM agents have massive knowledge and powerful generation capabilities, but lack structured execution frameworks. &lt;strong&gt;Procedural Graphs fill this gap&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Give agents a &lt;strong&gt;"skeleton"&lt;/strong&gt; — clear execution flow&lt;/li&gt;
&lt;li&gt;Give agents &lt;strong&gt;"learning ability"&lt;/strong&gt; — self-evolve from failures&lt;/li&gt;
&lt;li&gt;Give agents &lt;strong&gt;"explainability"&lt;/strong&gt; — humans can understand and modify its "thinking"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When that recipe starts rewriting itself, it's no longer just a book. It becomes a &lt;strong&gt;living thing&lt;/strong&gt; — constantly adapting, learning, and improving.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is the ultimate vision of Procedural Graphs&lt;/strong&gt;: Not giving agents a fixed program, but giving them a &lt;strong&gt;brain that can write its own programs&lt;/strong&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;This article is based on research published by Yuxing Lu, Yicheng Chen, and Shanchan Wu on September 9, 2026. Paper: arXiv:2609.08593&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>research</category>
      <category>graph</category>
    </item>
  </channel>
</rss>
