<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: joinwell52</title>
    <description>The latest articles on DEV Community by joinwell52 (@joinwell52).</description>
    <link>https://dev.to/joinwell52</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3887527%2F9ce60f79-7027-4ecd-8c9b-bf495e53c9b6.png</url>
      <title>DEV Community: joinwell52</title>
      <link>https://dev.to/joinwell52</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/joinwell52"/>
    <language>en</language>
    <item>
      <title>Farewell to the “Black-Box Myth”: Engineering Reflections on Mainstream Agent Frameworks and the Rise of Physical Externalization</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Tue, 22 Sep 2026 14:37:46 +0000</pubDate>
      <link>https://dev.to/joinwell52/farewell-to-the-black-box-myth-engineering-reflections-on-mainstream-agent-frameworks-and-the-25lg</link>
      <guid>https://dev.to/joinwell52/farewell-to-the-black-box-myth-engineering-reflections-on-mainstream-agent-frameworks-and-the-25lg</guid>
      <description>&lt;h2&gt;
  
  
  Introduction: The Recurring Cycle of Software Engineering and the Agent Illusion
&lt;/h2&gt;

&lt;p&gt;Looking back at the evolution of computing, every breakthrough primitive seems to trigger a familiar cycle: &lt;strong&gt;the primitive is first mythologized → oversized black-box middleware is built around it → reality pushes back through physical and engineering constraints → the industry eventually rediscovers the low-level boundaries that abstraction could never erase.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Distributed computing went through this in the 1990s. CORBA and DCOM tried to make remote invocation resemble a local function call as closely as possible, softening the visible boundaries of latency, partial failure, and partitioning. The Web took a different route: it accepted the network as a physical boundary and used HTTP, uniform interfaces, and looser resource semantics to reduce coupling. Enterprise integration followed a similar pattern. Heavyweight ESBs once tried to absorb message orchestration, transformation, and global coordination inside a central bus; later architectures redistributed many of those responsibilities across event streams, commit logs, APIs, and lighter contracts.&lt;/p&gt;

&lt;p&gt;The AI Agent industry is now replaying a remarkably similar architectural temptation. A large language model begins as a &lt;strong&gt;probabilistic generator&lt;/strong&gt;, yet once it can plan, call tools, write code, and coordinate with other models, we keep assigning it more roles: scheduler, memory system, message bus, state interpreter, and sometimes even the final judge of whether a task is complete.&lt;/p&gt;

&lt;p&gt;When production systems repeatedly encounter &lt;strong&gt;task loops, state drift, irreproducible execution, and opaque debugging&lt;/strong&gt;, the right question is no longer merely “can the model become smarter?” It is a more fundamental systems question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Are we assigning too many system responsibilities to a component that is excellent at probabilistic cognition but poorly suited to being the sole source of durable truth?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Over the past three years, the most important change in Agent engineering has not been that models are becoming more human-like. It is that &lt;strong&gt;systems are becoming less willing to trust models to remember everything, coordinate everything, and explain everything by themselves&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In 2023, prominent Agent systems emphasized iterative reasoning, role play, and natural-language collaboration. By 2025–2026, mainstream engineering frameworks had shifted more responsibility toward durable execution, checkpoints, state inspection, workflow state, and external artifacts. At the same time, EvoGit, tap, PheroPath, SwarmWorld, FCoP, and related work pushed coordination further out of messages and private runtimes and into Git, files, shared environments, and protocolized artifacts.&lt;/p&gt;

&lt;p&gt;This does not mean that “the filesystem will replace LangGraph,” nor that stigmergy has been proven superior to centralized orchestration. The more restrained and more important principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Agents may be probabilistic, but collaboration facts should not exist only inside probabilistic context.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;From Context to Messages, from Checkpointed State to Artifacts and then Protocolized Artifacts, the real change is a gradual migration of factual authority away from the model’s private context.&lt;/p&gt;

&lt;p&gt;The rest of this article examines what papers, frameworks, and open-source projects actually support—and what they do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Background: What Is Actually Changing in Agent Engineering?
&lt;/h2&gt;

&lt;p&gt;The strongest idea in the original argument is the refusal to keep treating the LLM as an omniscient state bus. But to make that claim technically defensible, three misconceptions must be corrected.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the problem is not that LLMs are incapable of self-correction. The problem is that they do not possess a built-in fact-type system.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Autoregressive models can notice and fix earlier errors. The danger is that once an erroneous output enters subsequent context, it is conditioned together with verified tool observations, user facts, and model speculation. Unless the surrounding system records provenance, validation status, re-fetch rules, and rollback boundaries, there is no native marker saying: “this sentence was only your earlier guess; do not treat it as a database fact.”&lt;/p&gt;

&lt;p&gt;Long-context research reinforces this caution. &lt;em&gt;Lost in the Middle&lt;/em&gt; showed that moving relevant evidence to the middle of a long context can reduce performance for many models. &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;[7]&lt;/a&gt; Later work has associated these effects with positional attention biases. Newer models have improved in some retrieval settings, so “middle information is always lost” should not be treated as a permanent law.&lt;/p&gt;

&lt;p&gt;The engineering principle should therefore be stated as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not require the model to maintain, over long periods, which historical statement is the current fact. Put identity, version, validity, and verification status into the system.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Second, natural language is not “absurd IPC,” but using open-ended language directly as the control protocol is expensive.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;IPC can be shared memory, pipes, sockets, HTTP, JSON, protobuf, or text. The real distinction is not binary versus textual. Control-plane protocols usually seek low ambiguity, explicit fields, typed errors, idempotency, and machine-verifiable boundaries. Free-form language pushes semantic interpretation back onto a probabilistic model.&lt;/p&gt;

&lt;p&gt;AutoGen illustrates the conversation-centric route well: its paper models conversable agents and flexible conversation patterns. &lt;a href="https://arxiv.org/html/2308.08155v2" rel="noopener noreferrer"&gt;[8]&lt;/a&gt; Its historical GroupChat design also reflects managed speaker selection and conversation-context passing. By 2026, AutoGen had entered maintenance mode and Microsoft was recommending the Microsoft Agent Framework to new users—another sign that “group chat as architecture” is no longer the only evolutionary path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, modern Agent frameworks have hybridized. They cannot be classified by a frozen 2023-era product image.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A better classification asks &lt;strong&gt;where coordination primarily lives&lt;/strong&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Primary coordination locus&lt;/th&gt;
&lt;th&gt;Typical mechanism&lt;/th&gt;
&lt;th&gt;Representative examples&lt;/th&gt;
&lt;th&gt;Main risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Context-centric&lt;/td&gt;
&lt;td&gt;Model repeatedly reads and extends context&lt;/td&gt;
&lt;td&gt;AutoGPT Classic, original BabyAGI&lt;/td&gt;
&lt;td&gt;Polluted history, token cost, weak recovery boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conversation-centric&lt;/td&gt;
&lt;td&gt;Agents coordinate through messages, group chat, and handoff&lt;/td&gt;
&lt;td&gt;AutoGen, MetaGPT, CrewAI Crews&lt;/td&gt;
&lt;td&gt;Semantic ambiguity, message growth, complex termination/routing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime-state-centric&lt;/td&gt;
&lt;td&gt;Graph/event/workflow + checkpoint + deterministic routing&lt;/td&gt;
&lt;td&gt;LangGraph, LlamaIndex Workflows, CrewAI Flows, current AutoGPT&lt;/td&gt;
&lt;td&gt;Framework coupling, state-schema evolution, runtime-specific recovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Artifact/environment-centric&lt;/td&gt;
&lt;td&gt;Git, files, worktrees, shared artifacts, environmental traces&lt;/td&gt;
&lt;td&gt;Aider, EvoGit, tap, PheroPath, FCoP, Govcraft, SwarmWorld&lt;/td&gt;
&lt;td&gt;Storage semantics, concurrency, metadata scale, artifact governance&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These styles are not mutually exclusive. Aider still uses model context but anchors code changes in Git; CrewAI combines autonomous Crews with controlled Flows; OpenHands combines conversations/events with a real workspace and code artifacts; LangGraph ecosystems increasingly use filesystem-backed tools.&lt;/p&gt;

&lt;p&gt;So the evolution is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;conversation frameworks → file frameworks&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is closer to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;model memory → framework memory → durable artifacts → open governance contracts.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Research Scope and Evidence Principles
&lt;/h2&gt;

&lt;p&gt;This article primarily draws on papers and open-source projects from &lt;strong&gt;2021–2026&lt;/strong&gt;, focusing on Agent state management, multi-Agent coordination, shared artifacts, Git/file collaboration, externalized memory, and stigmergic coordination. Papers are cited from original or formally published versions whenever possible; engineering claims are grounded primarily in official GitHub repositories, READMEs, and project documentation.&lt;/p&gt;

&lt;p&gt;Where implementation details cannot be fully verified from public materials, the article avoids inference and states the verification boundary explicitly. Judgments about reliability, auditability, and engineering complexity are qualitative architectural analyses based on public evidence, not a unified benchmark ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evidence Summary: Papers and Open-Source Projects
&lt;/h2&gt;

&lt;p&gt;The papers most relevant to the article’s thesis include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Work&lt;/th&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Evidence relevant to this article&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;Lost in the Middle&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2023; TACL 2024&lt;/td&gt;
&lt;td&gt;Long-context use is position-sensitive; evidence placed in the middle can reduce performance. This does not prove that models “forget,” but it undermines the idea that a context window is a reliable state database.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2308.00352" rel="noopener noreferrer"&gt;MetaGPT&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2023; ICLR 2024&lt;/td&gt;
&lt;td&gt;Organizes multi-Agent software work around roles and SOPs, and explicitly discusses cascading hallucination in naive chaining.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2308.08155" rel="noopener noreferrer"&gt;AutoGen&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2023&lt;/td&gt;
&lt;td&gt;Makes conversational agents and conversation patterns first-class abstractions for multi-Agent applications.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2307.07924" rel="noopener noreferrer"&gt;ChatDev&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2023; ACL 2024&lt;/td&gt;
&lt;td&gt;Decomposes software development into chat chains and adds communication/dehallucination mechanisms.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2506.02049" rel="noopener noreferrer"&gt;EvoGit&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2025-06&lt;/td&gt;
&lt;td&gt;Independent coding agents evolve shared code asynchronously through Git lineage without direct messaging or shared memory.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2604.08224" rel="noopener noreferrer"&gt;Externalization in LLM Agents&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-04&lt;/td&gt;
&lt;td&gt;Frames memory, skills, protocols, and harness engineering as externalization of cognitive burden.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2606.14445" rel="noopener noreferrer"&gt;tap&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-06&lt;/td&gt;
&lt;td&gt;Uses a file-first protocol, persistent Markdown messages, and Git worktree isolation for heterogeneous LLM-agent collaboration.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://arxiv.org/abs/2608.26081" rel="noopener noreferrer"&gt;SwarmWorld&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;2026-08&lt;/td&gt;
&lt;td&gt;Shows role differentiation and artifact reuse among initially homogeneous agents in a persistent shared environment, while preserving important limits on the superiority of interaction.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These papers support the claim that &lt;strong&gt;external structure is becoming more important&lt;/strong&gt;, not the stronger claim that central orchestration has been proven obsolete.&lt;/p&gt;

&lt;p&gt;The open-source landscape is similarly hybrid. AutoGPT has evolved beyond its Classic autonomous-loop image; CrewAI combines Crews and Flows; LangGraph makes checkpoints and durable execution core infrastructure; Aider anchors model edits in Git; PheroPath attaches coordination signals to files; FCoP formalizes typed collaboration artifacts; Govcraft experiments with pressure-field coordination.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technical Analysis: From Where State Lives to Where System Facts Live
&lt;/h2&gt;

&lt;p&gt;What makes an Agent system a “black box” is not whether it uses a graph, a database, or files. It is &lt;strong&gt;which layer owns final factual authority&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  State Carriers Define Failure Boundaries
&lt;/h3&gt;

&lt;p&gt;A context window has almost no infrastructure overhead: the model can read it directly. But guesses, observations, tool results, constraints, and protocol text can collapse into one token stream. As the history grows, cost and positional retrieval risk also grow. &lt;em&gt;Lost in the Middle&lt;/em&gt; is enough to show that “fits in context” and “can be reliably retrieved” are not equivalent. &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;[7]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Chat history is more structured than a monolithic context but is still primarily an event transcript, not authoritative application state. “Agent A said the task is complete” is not the same fact as “the task passed validation and was committed.”&lt;/p&gt;

&lt;p&gt;Checkpoint databases improve this substantially. LangGraph makes durable execution and state resume core capabilities; LlamaIndex Workflows can persist workflow state to files or databases. &lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[5]&lt;/a&gt; &lt;a href="https://github.com/run-llama/workflows-py" rel="noopener noreferrer"&gt;[29]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The artifact-oriented alternative changes the access relationship:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Framework-mediated state:
Human -&amp;gt; SDK -&amp;gt; serializer -&amp;gt; checkpoint -&amp;gt; state

Artifact-mediated state:
Human/tool -&amp;gt; file/git/schema -&amp;gt; state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second structure is easier to inspect out-of-band, but it is not automatically transparent. A 200 MB opaque JSON blob on disk is still a black box. Conversely, a PostgreSQL checkpoint with a stable open schema, versioning rules, and audit API can be highly auditable.&lt;/p&gt;

&lt;p&gt;“Physical externalization” should therefore mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Critical collaboration facts are persisted as artifacts with independent identity, stable semantics, open read paths, and verifiable provenance, so that their lifecycle does not depend on the model session that created them.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should not mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“written to a file = externalization complete.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Durability, Portability, and Auditability Are Different Properties
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Carrier&lt;/th&gt;
&lt;th&gt;Durability&lt;/th&gt;
&lt;th&gt;Portability&lt;/th&gt;
&lt;th&gt;Auditability&lt;/th&gt;
&lt;th&gt;Main note&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt / context&lt;/td&gt;
&lt;td&gt;low–medium&lt;/td&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;Transcript can be preserved, but runtime fact boundaries are weak&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chat-history DB&lt;/td&gt;
&lt;td&gt;medium–high&lt;/td&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;Has chronology, but message ≠ verified state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Private checkpoint&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;medium–low&lt;/td&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;Strong recovery; cross-framework interpretation depends on schema/API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;JSON / Markdown artifacts&lt;/td&gt;
&lt;td&gt;high*&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;high*&lt;/td&gt;
&lt;td&gt;Depends on storage, schema, and provenance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Git objects / commits&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;Excellent for code/text; not ideal for every high-frequency mutable state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;xattr&lt;/td&gt;
&lt;td&gt;medium&lt;/td&gt;
&lt;td&gt;low–medium&lt;/td&gt;
&lt;td&gt;medium–low&lt;/td&gt;
&lt;td&gt;Bound to files, but Git/copy/cross-platform visibility varies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Object-store artifacts&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;Namespace and atomic-transition semantics differ from POSIX&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;PheroPath exposes an important paradox: &lt;strong&gt;physical existence does not automatically imply universal visibility.&lt;/strong&gt; Hidden xattrs can be elegant, but an explicit sidecar JSON file may be easier for Git, CI, backups, and cross-language tools to consume.&lt;/p&gt;

&lt;h3&gt;
  
  
  Control Flow Is Moving from Language Semantics Toward Deterministic Boundaries
&lt;/h3&gt;

&lt;p&gt;CrewAI now pairs role-based Crews with event-driven Flows that support structured state, branching, and routing. &lt;a href="https://github.com/crewAIInc/crewAI" rel="noopener noreferrer"&gt;[44]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LangGraph makes durable workflow and state transition part of its orchestration layer. &lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[5]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;LlamaIndex Workflows uses async functions that produce and consume events while supporting loops, parallelism, persistence, and recovery. &lt;a href="https://github.com/run-llama/workflows-py" rel="noopener noreferrer"&gt;[29]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These systems are converging on a common principle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“What happens next” should not always be decided by asking the model to say one more thing.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model is well suited to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Unstructured input
    ↓
Candidate structured decision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The workflow layer should own:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Candidate decision
    ↓
Validation
    ↓
Authorized state transition
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to make every model call literally stateless. Real Agents have tool sessions, caches, memory, and side effects. The goal is to prevent any one model call from having exclusive authority over global system facts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stigmergy Reduces Direct Coordination Dependence; It Does Not Eliminate Coordination
&lt;/h3&gt;

&lt;p&gt;SwarmWorld, EvoGit, and Govcraft support the idea that a shared environment, Git genealogy, or shared-state pressure field can become a &lt;strong&gt;first-class coordination medium&lt;/strong&gt;, rather than merely passive storage after a conversation. &lt;a href="https://arxiv.org/abs/2506.02049" rel="noopener noreferrer"&gt;[14]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2608.26081" rel="noopener noreferrer"&gt;[22]&lt;/a&gt; &lt;a href="https://github.com/Govcraft/pressure-field-experiment" rel="noopener noreferrer"&gt;[35]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That does not prove that “stigmergy always beats conversation.”&lt;/p&gt;

&lt;p&gt;A more defensible conclusion is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;For problems with a shared constraint surface, where local actions leave measurable environmental changes and solutions can accumulate through local improvement, stigmergic coordination deserves to be compared directly with conversation and hierarchy as an independent architectural baseline.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Strongly transactional workflows, approval chains, and regulated processes still require explicit accountability and global invariants.&lt;/p&gt;

&lt;h3&gt;
  
  
  Atomicity Is Where Physical Externalization Is Most Easily Romanticized
&lt;/h3&gt;

&lt;p&gt;FCoP treats &lt;code&gt;os.rename()&lt;/code&gt; as an important synchronization primitive on shared directory trees, but FCoP 4.0.3 has evolved far beyond “directory + Markdown + rename.” Its public protocol semantics include TASK / REPORT / ISSUE / REVIEW, Attempt, relations, authorization, Branch Family convergence, idempotency, and crash-safe recovery.&lt;/p&gt;

&lt;p&gt;The attraction of filesystem state transitions is obvious:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;inbox/TASK-123.md
      |
      | atomic namespace transition
      v
active/TASK-123.md
      |
      | result validated
      v
review/TASK-123.md
      |
      | accepted
      v
done/TASK-123.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But four concepts must remain separate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;atomic visibility
≠
exclusive ownership
≠
crash durability
≠
distributed consensus
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A same-filesystem rename can help guarantee that observers do not see a half-renamed path. It does not solve every distributed-consistency problem.&lt;/p&gt;

&lt;p&gt;Object storage makes the distinction even clearer. General-purpose S3 historically implements renaming through copy plus delete, while S3 Express One Zone directory buckets now expose atomic &lt;code&gt;RenameObject&lt;/code&gt;. &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/directory-buckets-objects-rename.html" rel="noopener noreferrer"&gt;[42]&lt;/a&gt; &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html" rel="noopener noreferrer"&gt;[43]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The engineering lesson is not to force every backend to behave like POSIX. It is to separate protocol invariants from backend-specific storage primitives.&lt;/p&gt;

&lt;h3&gt;
  
  
  Small Files, Orphaned State, and Recovery Do Not Disappear by Declaration
&lt;/h3&gt;

&lt;p&gt;There is no universal file-count threshold at which every filesystem “fails.” The general problem is that, at scale, cost shifts from payload I/O toward namespace and metadata operations, directory scanning, indexing, watching, backup, retention, antivirus, and synchronization.&lt;/p&gt;

&lt;p&gt;A better design uses hot/cold tiers, immutable event segments, compacted snapshots, and archival storage.&lt;/p&gt;

&lt;p&gt;Likewise, “orphan lock” is only one ownership pattern. With immutable attempts, process death can leave an incomplete attempt instead of a permanent blocking lock:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;task-42/
    attempt-001/  failed
    attempt-002/  expired
    attempt-003/  accepted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The recovery principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Do not recover by erasing the accident. Recover by making the accident a legitimate part of history.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  The Real Advantage of Externalization Is Not “It Cannot Fail”
&lt;/h3&gt;

&lt;p&gt;Exposing state through a filesystem or object store does not automatically solve concurrency, identity, authorization, schema migration, or recovery.&lt;/p&gt;

&lt;p&gt;Its greatest advantage is not that it “never fails.” It is this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When it fails, there is a better chance that the failure leaves facts a third party can inspect.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzdajld9xcc5lnj96fpaj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzdajld9xcc5lnj96fpaj.png" alt="Figure 2 | From Checkpoints and Git to Filesystem Protocols" width="800" height="566"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2 | From Checkpoints and Git to Filesystem Protocols: collaboration facts progressively externalize.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Case Deep Dive: From Checkpoints and Git to Filesystem Protocols
&lt;/h2&gt;

&lt;h3&gt;
  
  
  LangGraph: Beyond the “In-Memory Black Box,” but Still Runtime-Defined
&lt;/h3&gt;

&lt;p&gt;LangGraph explicitly positions itself as a low-level orchestration framework for long-running, stateful agents and makes durable execution, failure resume, human state inspection, and persistent memory core capabilities. &lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[5]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the criticism “LangGraph crashes and all state is lost” is obsolete.&lt;/p&gt;

&lt;p&gt;The more accurate structure is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM / Tool Node
      ↓
Graph Transition
      ↓
Framework State
      ↓
Checkpoint
      ↓
Persistent Backend
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The deeper question is whether a checkpoint can be understood by another language or governance tool without importing LangGraph, starting the original Python application, or reconstructing the original node semantics.&lt;/p&gt;

&lt;p&gt;If not, the state remains partly enclosed by the runtime. This is not a defect so much as the natural price of a runtime abstraction.&lt;/p&gt;

&lt;p&gt;A pragmatic design keeps both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LangGraph internal checkpoint
           +
external canonical artifacts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Internal checkpoint state serves recovery; canonical artifacts serve cross-framework fact exchange.&lt;/p&gt;

&lt;h3&gt;
  
  
  Aider: The Important Lesson Is Git as a Fact Anchor, Not SPEC.md
&lt;/h3&gt;

&lt;p&gt;The claim that Aider requires every task to be materialized into &lt;code&gt;SPEC.md&lt;/code&gt; or checkbox files is not supported by its official architecture.&lt;/p&gt;

&lt;p&gt;What Aider explicitly does is map the repository, make Git commits, allow ordinary Git diff/manage/undo workflows, and run lint/test after changes. &lt;a href="https://github.com/Aider-AI/aider" rel="noopener noreferrer"&gt;[9]&lt;/a&gt; &lt;a href="https://github.com/Aider-AI/aider" rel="noopener noreferrer"&gt;[30]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The important pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Prompt / conversation
        ↓
candidate edit
        ↓
working tree
        ↓
git diff
        ↓
lint / test
        ↓
commit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I fixed it.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;is not the fact.&lt;/p&gt;

&lt;p&gt;The facts are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;diff exists
test exit code = 0
commit hash = ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a migration from linguistic claims to externally verifiable engineering facts.&lt;/p&gt;

&lt;h3&gt;
  
  
  PheroPath: A Powerful Idea—and an Auditability Paradox
&lt;/h3&gt;

&lt;p&gt;PheroPath is one of the most literal “digital pheromone” experiments. &lt;a href="https://github.com/zi-yue-1129/PheroPath" rel="noopener noreferrer"&gt;[32]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It stores signals such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DANGER
TODO
SAFE
INSIGHT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;through filesystem extended attributes or sidecar JSON, with CLI operations such as &lt;code&gt;sniff&lt;/code&gt;, &lt;code&gt;secrete&lt;/code&gt;, and &lt;code&gt;cleanse&lt;/code&gt;, plus temporal decay and editor visualization.&lt;/p&gt;

&lt;p&gt;The key insight is simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Context should be attached to the object being acted upon, not only to one conversation with the model.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But PheroPath also reveals a hierarchy inside “externalization”:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exists on disk
     ≠
visible to ordinary tools
     ≠
tracked by Git
     ≠
portable across platforms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For governance-critical data, an explicit sidecar artifact may be less elegant than xattr but easier for Git review, CI, backup, and object-store migration.&lt;/p&gt;

&lt;h3&gt;
  
  
  FCoP: From “Files as an Implementation Detail” to “Files Carry Protocol Semantics”
&lt;/h3&gt;

&lt;p&gt;FCoP is a particularly useful example of &lt;strong&gt;Protocolized Artifacts&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Its formal name is &lt;strong&gt;File-based Coordination Protocol&lt;/strong&gt;. &lt;code&gt;Filename as Protocol&lt;/code&gt; is one of its core design ideas. Its paper/research record is published at DOI &lt;code&gt;10.5281/zenodo.22855630&lt;/code&gt;; the latest code archive is Zenodo Record &lt;code&gt;22746175&lt;/code&gt;, version 4.0.3. &lt;a href="https://joinwell52-ai.github.io/FCoP/" rel="noopener noreferrer"&gt;[33]&lt;/a&gt; &lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;[37]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;FCoP 4.0.3 publicly defines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK / REPORT / ISSUE / REVIEW
        ↓
explicit lifecycle
        ↓
Attempt / Relation / Authorization
        ↓
Root TASK + Branch Family
        ↓
family inspection + explicit convergence
        ↓
durable REVIEW / merge record
        ↓
idempotency + crash-safe recovery
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Its significance is no longer merely “state mapped to folders.” It attempts to make collaboration facts carry explicit type, identity, state transition, authorization, relation, and recovery semantics.&lt;/p&gt;

&lt;p&gt;FCoP also states its boundary: it is not itself a scheduler, Agent runtime, LLM SDK, database, or automatic merge-decision engine. Its value in this discussion lies in trying to make collaboration facts outlive any one model session or private runtime.&lt;/p&gt;

&lt;p&gt;That does not make the filesystem a universal distributed-systems solution. Atomicity, backend adaptation, identity, schema evolution, binary artifacts, and security remain real engineering boundaries.&lt;/p&gt;

&lt;h3&gt;
  
  
  Govcraft and SwarmWorld: Valuable Because They Do Not Fully Support the Strongest Claim
&lt;/h3&gt;

&lt;p&gt;Govcraft reports 270 meeting-room scheduling trials with the following solve rates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Solve rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pressure field&lt;/td&gt;
&lt;td&gt;48.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Conversation-style&lt;/td&gt;
&lt;td&gt;12.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hierarchical&lt;/td&gt;
&lt;td&gt;1.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sequential&lt;/td&gt;
&lt;td&gt;0.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random&lt;/td&gt;
&lt;td&gt;0.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://github.com/Govcraft/pressure-field-experiment" rel="noopener noreferrer"&gt;[35]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This shows that, on that benchmark, the stigmergy-inspired shared-state pressure field clearly outperformed the authors’ baselines.&lt;/p&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; show that hierarchy generally succeeds only 1.5% of the time in Agent systems.&lt;/p&gt;

&lt;p&gt;SwarmWorld provides richer but similarly bounded evidence: shared societies build broader and more resilient technology portfolios; roles emerge; much reuse begins with environmental observation; yet isolated best-of-N search can still produce competitive best single artifacts. &lt;a href="https://arxiv.org/abs/2608.26081" rel="noopener noreferrer"&gt;[22]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Together they support this narrower conclusion:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The environment can be a first-class coordination medium, not merely passive storage written after conversation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  A Physical-Externalization Workflow
&lt;/h3&gt;

&lt;p&gt;A practical production pattern looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Task Queue / Workflow
        ↓
Canonical Task Artifact
        ↓
Agent Worker
        ↓
Candidate / Attempt Artifact
        ↓
Deterministic Validator
   ┌───────────────┐
   │               │
failed          passed
   │               │
new attempt     review / done
   │               │
   └──── evidence ─┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LLM no longer acts as the global state machine. It does what probabilistic cognition is good at: understanding open-ended problems, proposing candidate changes, summarizing, and planning. Final state transitions depend on external evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practice Route: Do Not “Migrate to Files”; Migrate to an Open Fact Contract
&lt;/h2&gt;

&lt;p&gt;The practical goal is not to rewrite the Agent stack overnight. It is to gradually reduce the private runtime’s monopoly on factual authority.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Separate Transcript, Working State, and Canonical Fact
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Transcript
    what the Agent said

Working State
    where the runtime currently is

Canonical Facts
    what has been verified and committed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;“Tests pass” in a message is transcript. A validation record containing the commit, exit code, test-suite version, and timestamp can participate in a state transition.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Let Model Output Become a Candidate Before It Becomes State
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LLM proposal
    ↓
Schema validation
    ↓
Deterministic / externally verifiable checks
    ↓
Commit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The core rule is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“The model says the state should change” and “the system permits the state change” must be two different events.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  3. Keep Runtime Checkpoints, but Add a Canonical Artifact Boundary
&lt;/h3&gt;

&lt;p&gt;LangGraph, CrewAI, and workflow-runtime checkpoints remain valuable for execution recovery. There is no need to delete them in the name of externalization.&lt;/p&gt;

&lt;p&gt;Use dual state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;internal checkpoint
    -&amp;gt; serves resume

canonical artifacts
    -&amp;gt; serve collaboration, audit, and interoperability
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During migration, dual-write and reconcile until canonical artifacts can become the cross-system authority.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Treat Storage Semantics, Attempts, and Schemas as Part of the Protocol
&lt;/h3&gt;

&lt;p&gt;Do not pretend POSIX is universal. Local filesystems, general-purpose S3, S3 Express, databases, and Git all have different transition primitives. &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/directory-buckets-objects-rename.html" rel="noopener noreferrer"&gt;[42]&lt;/a&gt; &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html" rel="noopener noreferrer"&gt;[43]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Prefer immutable attempts over repeatedly overwriting one state object:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;task-id
  ├─ attempt-001 [expired]
  ├─ attempt-002 [failed]
  └─ attempt-003 [accepted]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Governance fields should also be structured and validatable, rather than buried only inside natural-language prose.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Test Four System Invariants
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Safety&lt;/strong&gt;: unverified candidates must never become canonical facts.&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Liveness&lt;/strong&gt;: failed workers must not permanently block progress.&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Auditability&lt;/strong&gt;: every committed transition must trace back to its inputs and evidence.&lt;br&gt;&lt;br&gt;
&lt;strong&gt;Portability&lt;/strong&gt;: a third-party implementation must be able to interpret canonical state without the original Agent session.&lt;/p&gt;

&lt;p&gt;If these four properties hold, the underlying store may be LangGraph, PostgreSQL, Git, a POSIX filesystem, S3, or a hybrid.&lt;/p&gt;

&lt;p&gt;If they do not, even a directory tree can become a filesystem-shaped black box.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6u0aa8i927gsugi7joe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu6u0aa8i927gsugi7joe.png" alt="Figure 1 | The State-Externalization Ladder in Agent Engineering" width="800" height="566"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1 | The State-Externalization Ladder: Context → Messages → Checkpointed State → Artifacts → Protocolized Artifacts.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Conclusion and Future Research
&lt;/h2&gt;

&lt;p&gt;What Agent engineering needs to abandon is not one particular framework. It is a deeper myth:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If the model becomes capable enough, it can simultaneously serve as reasoner, state machine, database, scheduler, message bus, authorization arbiter, and auditor.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Recent engineering evolution is steadily undermining that assumption. AutoGPT has moved from a standalone autonomous-agent image toward a workflow platform; CrewAI places Crews beside Flows; LangGraph makes checkpoints, durable execution, and state inspection foundational; Aider returns diff, commit, lint, and test to mature software-engineering tools; EvoGit, tap, PheroPath, and FCoP anchor collaboration facts in Git, files, shared environments, and protocolized artifacts. &lt;a href="https://github.com/Significant-Gravitas/AutoGPT" rel="noopener noreferrer"&gt;[24]&lt;/a&gt; &lt;a href="https://github.com/crewAIInc/crewAI" rel="noopener noreferrer"&gt;[44]&lt;/a&gt; &lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[5]&lt;/a&gt; &lt;a href="https://github.com/Aider-AI/aider" rel="noopener noreferrer"&gt;[30]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2506.02049" rel="noopener noreferrer"&gt;[14]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2606.14445" rel="noopener noreferrer"&gt;[45]&lt;/a&gt; &lt;a href="https://github.com/zi-yue-1129/PheroPath" rel="noopener noreferrer"&gt;[32]&lt;/a&gt; &lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;[37]&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The “rise of physical externalization” should therefore be retained, but redefined.&lt;/p&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; mean:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Filesystem will replace Agent frameworks.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Canonical collaboration state should progressively escape the exclusive custody of model context and proprietary runtime state.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That state may live in files, Git, append-only event logs, databases, content-addressed stores, or object storage. What matters is that it be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Persistent
Explicit
Addressable
Versioned
Inspectable
Validatable
Replayable
Portable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Stigmergy, Physical Externalization, storage media, and concrete protocols occupy different abstraction layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Stigmergy
    │
    │  coordination principle
    ▼
Physical Externalization
    │
    │  systems-design principle
    ▼
Filesystem / Git / DB / Object Store / Artifact Graph
    │
    │  implementation media
    ▼
FCoP / tap / PheroPath / project-specific contracts
       concrete protocols and tools
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Mature next-generation Agent infrastructure is therefore unlikely to be “pure stigmergy” or “pure workflow.” It will more likely be hybrid:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              ┌─────────────────────────┐
              │  Deterministic Workflow │
              │ policy / routing / auth │
              └────────────┬────────────┘
                           │
             candidate work│
                           ▼
┌───────────────┐   ┌──────────────────────┐
│ LLM Workers   │──▶│ Canonical Artifacts  │
│ probabilistic │   │ schema + provenance  │
└───────────────┘   └──────────┬───────────┘
                               │
                  ┌────────────┼─────────────┐
                  ▼            ▼             ▼
               Git/FS       Database     Object Store
                  │            │             │
                  └────────────┼─────────────┘
                               ▼
                       Validators / CI
                               │
                               ▼
                     Committed State
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The most important future research questions are not “which Agent prompt is smartest,” but at least five deeper problems.&lt;/p&gt;

&lt;p&gt;First, &lt;strong&gt;artifact coordination benchmarks&lt;/strong&gt;: compare conversation, supervisor, graph workflow, shared artifacts, stigmergy, and hybrids under the same task, model, and compute budget.&lt;/p&gt;

&lt;p&gt;Second, &lt;strong&gt;portable Agent-state standards&lt;/strong&gt;: durable state still lacks a minimal cross-runtime fact contract comparable to OpenTelemetry or OCI manifests.&lt;/p&gt;

&lt;p&gt;Third, &lt;strong&gt;artifact provenance and security&lt;/strong&gt;: future attack surfaces include forged TASK artifacts, modified front matter, malicious pheromones, replayed stale artifacts, and filename-based routing attacks.&lt;/p&gt;

&lt;p&gt;Fourth, &lt;strong&gt;namespace economics at scale&lt;/strong&gt;: large Agent fleets may create millions of attempts, observations, and validation artifacts per day, forcing serious work on compaction, snapshots, content addressing, hot/cold tiers, indexing, and retention.&lt;/p&gt;

&lt;p&gt;Fifth, &lt;strong&gt;the real boundary of human–Agent isomorphism&lt;/strong&gt;: Markdown is excellent for human readability, while protobuf or database schemas can reintroduce runtime dependence. A durable artifact ecosystem must find a way to combine a human-readable surface with a machine-verifiable core.&lt;/p&gt;

&lt;p&gt;Ultimately, mature Agent infrastructure should not be designed around the assumption that “the model never makes mistakes.”&lt;/p&gt;

&lt;p&gt;That goal is unrealistic.&lt;/p&gt;

&lt;p&gt;The better goal is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Prevent mistakes from silently becoming system facts.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Models may misjudge, hallucinate, crash, upgrade, switch vendors, or forget what they said one session earlier.&lt;/p&gt;

&lt;p&gt;But why a task entered &lt;code&gt;DONE&lt;/code&gt;, which input produced which patch, which test passed, who approved the transition, which worker failed on which version, and which constraint was valid at the time should not depend on whether the model still remembers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model proposes possibilities.&lt;br&gt;&lt;br&gt;
Validators judge evidence.&lt;br&gt;&lt;br&gt;
Workflows constrain transitions.&lt;br&gt;&lt;br&gt;
Persistent media preserve facts.&lt;br&gt;&lt;br&gt;
Open protocols keep facts from belonging to any one framework.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is the Agent engineering worth pursuing after the black-box myth.&lt;/p&gt;

&lt;p&gt;And this is why the most durable idea in “physical externalization” is not “Filesystem forever,” but a more fundamental systems principle:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Agents may be probabilistic; coordination may be emergent; but facts that have happened and been accepted by the system must be addressable, verifiable, recoverable, and auditable.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When an Agent system reaches that point, the large language model can finally move from being imagined as the brain of the entire system back to the role it is best and safest at: &lt;strong&gt;a powerful but replaceable cognitive execution unit, rather than the sole memory and source of truth for the digital world.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/Significant-Gravitas/AutoGPT" rel="noopener noreferrer"&gt;[1]&lt;/a&gt; &lt;a href="https://github.com/Significant-Gravitas/AutoGPT" rel="noopener noreferrer"&gt;[24]&lt;/a&gt; GitHub - Significant-Gravitas/AutoGPT: AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters. · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Significant-Gravitas/AutoGPT" rel="noopener noreferrer"&gt;https://github.com/Significant-Gravitas/AutoGPT&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[2]&lt;/a&gt; &lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[5]&lt;/a&gt; &lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;[39]&lt;/a&gt; GitHub - langchain-ai/langgraph: Build resilient agents. · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/langchain-ai/langgraph" rel="noopener noreferrer"&gt;https://github.com/langchain-ai/langgraph&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2506.02049" rel="noopener noreferrer"&gt;[3]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2506.02049" rel="noopener noreferrer"&gt;[14]&lt;/a&gt; EvoGit: Decentralized Code Evolution via Git-Based Multi-Agent Collaboration&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2506.02049" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2506.02049&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2604.08224" rel="noopener noreferrer"&gt;[4]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2604.08224" rel="noopener noreferrer"&gt;[47]&lt;/a&gt; Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2604.08224" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2604.08224&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html" rel="noopener noreferrer"&gt;[6]&lt;/a&gt; &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html" rel="noopener noreferrer"&gt;[43]&lt;/a&gt; RenameObject - Amazon S3&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;[7]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;[10]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;[38]&lt;/a&gt; Lost in the Middle: How Language Models Use Long Contexts&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2307.03172" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2307.03172&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2308.08155v2" rel="noopener noreferrer"&gt;[8]&lt;/a&gt; &lt;a href="https://arxiv.org/html/2308.08155v2" rel="noopener noreferrer"&gt;[12]&lt;/a&gt; Enabling Next-Gen LLM Applications via Multi-Agent Conversation&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2308.08155v2" rel="noopener noreferrer"&gt;https://arxiv.org/html/2308.08155v2&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Aider-AI/aider" rel="noopener noreferrer"&gt;[9]&lt;/a&gt; &lt;a href="https://github.com/Aider-AI/aider" rel="noopener noreferrer"&gt;[30]&lt;/a&gt; GitHub - Aider-AI/aider: aider is AI pair programming in your terminal · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Aider-AI/aider" rel="noopener noreferrer"&gt;https://github.com/Aider-AI/aider&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2308.00352v7" rel="noopener noreferrer"&gt;[11]&lt;/a&gt; MetaGPT: Meta Programming for a Multi-Agent Collaborative Framework&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2308.00352v7" rel="noopener noreferrer"&gt;https://arxiv.org/html/2308.00352v7&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2307.07924v5" rel="noopener noreferrer"&gt;[13]&lt;/a&gt; ChatDev: Communicative Agents for Software Development&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2307.07924v5" rel="noopener noreferrer"&gt;https://arxiv.org/html/2307.07924v5&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2601.13295v1" rel="noopener noreferrer"&gt;[15]&lt;/a&gt; CooperBench: Why Coding Agents Cannot be Your Teammates Yet&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2601.13295v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2601.13295v1&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2604.07821v2" rel="noopener noreferrer"&gt;[16]&lt;/a&gt; More Capable, Less Cooperative? When LLMs Fail At Zero-Cost ...&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2604.07821v2" rel="noopener noreferrer"&gt;https://arxiv.org/html/2604.07821v2&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2605.03310v1" rel="noopener noreferrer"&gt;[17]&lt;/a&gt; Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2605.03310v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2605.03310v1&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2603.21489v2" rel="noopener noreferrer"&gt;[18]&lt;/a&gt; Effective Strategies for Asynchronous Software Engineering Agents&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2603.21489v2" rel="noopener noreferrer"&gt;https://arxiv.org/html/2603.21489v2&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2606.19911v1" rel="noopener noreferrer"&gt;[19]&lt;/a&gt; Multi-Agent Transactive Memory&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2606.19911v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2606.19911v1&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2606.30668v1" rel="noopener noreferrer"&gt;[20]&lt;/a&gt; Emergent Culture in Minimal LLM Systems&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2606.30668v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2606.30668v1&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2606.14445" rel="noopener noreferrer"&gt;[21]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2606.14445" rel="noopener noreferrer"&gt;[45]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2606.14445" rel="noopener noreferrer"&gt;[48]&lt;/a&gt; tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2606.14445" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2606.14445&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2608.26081" rel="noopener noreferrer"&gt;[22]&lt;/a&gt; &lt;a href="https://arxiv.org/abs/2608.26081" rel="noopener noreferrer"&gt;[36]&lt;/a&gt; SwarmWorld: Stigmergic technological evolution in societies of language-model agents&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/abs/2608.26081" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2608.26081&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2604.05080v1" rel="noopener noreferrer"&gt;[23]&lt;/a&gt; Nidus: Externalized Reasoning for AI-Assisted Engineering&lt;/p&gt;

&lt;p&gt;&lt;a href="https://arxiv.org/html/2604.05080v1" rel="noopener noreferrer"&gt;https://arxiv.org/html/2604.05080v1&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/yoheinakajima/babyagi" rel="noopener noreferrer"&gt;[25]&lt;/a&gt; yoheinakajima/babyagi&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/yoheinakajima/babyagi" rel="noopener noreferrer"&gt;https://github.com/yoheinakajima/babyagi&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/geekan/MetaGPT" rel="noopener noreferrer"&gt;[26]&lt;/a&gt; GitHub - FoundationAgents/MetaGPT: 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/geekan/MetaGPT" rel="noopener noreferrer"&gt;https://github.com/geekan/MetaGPT&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/microsoft/autogen" rel="noopener noreferrer"&gt;[27]&lt;/a&gt; GitHub - microsoft/autogen: A programming framework for agentic AI · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/microsoft/autogen" rel="noopener noreferrer"&gt;https://github.com/microsoft/autogen&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/crewAIInc/crewAI" rel="noopener noreferrer"&gt;[28]&lt;/a&gt; &lt;a href="https://github.com/crewAIInc/crewAI" rel="noopener noreferrer"&gt;[40]&lt;/a&gt; &lt;a href="https://github.com/crewAIInc/crewAI" rel="noopener noreferrer"&gt;[44]&lt;/a&gt; GitHub - crewAIInc/crewAI: Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly, tackling complex tasks. · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/crewAIInc/crewAI" rel="noopener noreferrer"&gt;https://github.com/crewAIInc/crewAI&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/run-llama/workflows-py" rel="noopener noreferrer"&gt;[29]&lt;/a&gt; GitHub - run-llama/workflows-py: event-driven, async-first workflows for AI applications&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/run-llama/workflows-py" rel="noopener noreferrer"&gt;https://github.com/run-llama/workflows-py&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/OpenHands/OpenHands" rel="noopener noreferrer"&gt;[31]&lt;/a&gt; GitHub - OpenHands/OpenHands: AI-Driven Development · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/OpenHands/OpenHands" rel="noopener noreferrer"&gt;https://github.com/OpenHands/OpenHands&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/zi-yue-1129/PheroPath" rel="noopener noreferrer"&gt;[32]&lt;/a&gt; GitHub - zi-yue-1129/PheroPath: stigmergy-based file pheromones for AI Agents · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/zi-yue-1129/PheroPath" rel="noopener noreferrer"&gt;https://github.com/zi-yue-1129/PheroPath&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/FCoP/" rel="noopener noreferrer"&gt;[33]&lt;/a&gt; FCoP — A File-based Coordination Protocol for Multi-Agent AI Systems&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/FCoP/" rel="noopener noreferrer"&gt;https://joinwell52-ai.github.io/FCoP/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI" rel="noopener noreferrer"&gt;[34]&lt;/a&gt; joinwell52-AI&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI" rel="noopener noreferrer"&gt;https://github.com/joinwell52-AI&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Govcraft/pressure-field-experiment" rel="noopener noreferrer"&gt;[35]&lt;/a&gt; &lt;a href="https://github.com/Govcraft/pressure-field-experiment" rel="noopener noreferrer"&gt;[46]&lt;/a&gt; GitHub - Govcraft/pressure-field-experiment: stigmergy-inspired pressure-field coordination for multi-agent LLM systems · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Govcraft/pressure-field-experiment" rel="noopener noreferrer"&gt;https://github.com/Govcraft/pressure-field-experiment&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;[37]&lt;/a&gt; &lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;[41]&lt;/a&gt; GitHub - joinwell52-AI/FCoP: File-based Coordination Protocol for Multi-Agent AI Systems · GitHub&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;https://github.com/joinwell52-AI/FCoP&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/directory-buckets-objects-rename.html" rel="noopener noreferrer"&gt;[42]&lt;/a&gt; Renaming objects in directory buckets - Amazon Simple Storage Service&lt;/p&gt;

&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/directory-buckets-objects-rename.html" rel="noopener noreferrer"&gt;https://docs.aws.amazon.com/AmazonS3/latest/userguide/directory-buckets-objects-rename.html&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://doi.org/10.5281/zenodo.22855630" rel="noopener noreferrer"&gt;[49]&lt;/a&gt; FCoP paper / research record — DOI 10.5281/zenodo.22855630&lt;/p&gt;

&lt;p&gt;&lt;a href="https://doi.org/10.5281/zenodo.22855630" rel="noopener noreferrer"&gt;https://doi.org/10.5281/zenodo.22855630&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://zenodo.org/records/22746175" rel="noopener noreferrer"&gt;[50]&lt;/a&gt; FCoP v4.0.3 code archive — Zenodo Record 22746175&lt;/p&gt;

&lt;p&gt;&lt;a href="https://zenodo.org/records/22746175" rel="noopener noreferrer"&gt;https://zenodo.org/records/22746175&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Chinese Editions:&lt;/em&gt; &lt;a href="https://blog.csdn.net/m0_51507544/article/details/166372784" rel="noopener noreferrer"&gt;CSDN&lt;/a&gt; · &lt;a href="https://juejin.cn/post/7688170325382955008" rel="noopener noreferrer"&gt;Juejin&lt;/a&gt; · &lt;a href="https://zhuanlan.zhihu.com/p/2085851673977677544" rel="noopener noreferrer"&gt;Zhihu&lt;/a&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Research Artifacts:&lt;/em&gt; &lt;a href="https://zenodo.org/records/22855630" rel="noopener noreferrer"&gt;FCoP (Zenodo)&lt;/a&gt; · &lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;FCoP Repository&lt;/a&gt; · &lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Filesystem as Protocol: How FCoP Uses Minimal Markdown for Agent Governance</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Mon, 21 Sep 2026 09:12:17 +0000</pubDate>
      <link>https://dev.to/joinwell52/filesystem-as-protocol-how-fcop-uses-minimal-markdown-for-agent-governance-199i</link>
      <guid>https://dev.to/joinwell52/filesystem-as-protocol-how-fcop-uses-minimal-markdown-for-agent-governance-199i</guid>
      <description>&lt;p&gt;In the previous article, we examined the “communication maze” and governance anxiety emerging as multi-Agent systems move into real-world use in 2026. As more and more agents exchange state, tasks, and tool calls through complex JSON, YAML, and black-box APIs, human operators are finding it increasingly difficult to see clearly who owns a piece of work, how far it has progressed, what justifies the next step, and who is responsible when something goes wrong.&lt;/p&gt;

&lt;p&gt;As major technology companies and frontier engineering research increasingly emphasize &lt;strong&gt;human-readable contracts&lt;/strong&gt; and &lt;strong&gt;process transparency&lt;/strong&gt;, a practical question follows: how can we implement an Agent governance mechanism that is &lt;strong&gt;lightweight, elegant, and capable of industrial-grade enforcement&lt;/strong&gt;?&lt;/p&gt;

&lt;p&gt;This article focuses on an open protocol we designed and incubated ourselves: &lt;strong&gt;FCoP (File-based Coordination Protocol)&lt;/strong&gt;. The aim is to see how it challenges the assumption that governance must begin with complex infrastructure, and how ordinary Markdown, directories, and files can be used to establish real constraints over multi-Agent collaboration.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Why return to the Unix philosophy of files?
&lt;/h2&gt;

&lt;p&gt;One of the main intellectual influences behind FCoP is the Unix philosophy. Unix does not try to place every capability inside one giant central program. Instead, it favors small tools that each do one thing well, exchange information through simple and stable interfaces, and compose into more complex systems.&lt;/p&gt;

&lt;p&gt;FCoP brings the same idea into multi-Agent collaboration. In the protocol, TASK expresses work to be done, REPORT represents a deliverable, ISSUE records unresolved problems, and REVIEW records inspection and decision. Markdown remains directly readable by humans while also being machine-parseable through structured fields.&lt;/p&gt;

&lt;p&gt;So “filesystem as protocol” is not a mechanical restatement of “everything is a file.” What FCoP inherits from Unix is more important: &lt;strong&gt;textual interfaces, small and explicit responsibilities, and composable tools&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;As Agent systems become more complex, a common instinct is to make governance systems equally complex by adding databases, state services, message queues, event buses, and workflow engines. FCoP asks a different question: how many facts are actually indispensable for reliable work handoff across multiple agents?&lt;/p&gt;

&lt;p&gt;The next responsible actor usually needs to know only a limited set of things: what the task is, who currently owns it, which execution attempt is active, what has been delivered, who reviewed that delivery, and whether that review is still applicable to the current work. If those facts can be stored stably, reread later, and verified by software, governance does not necessarily have to begin with a large centralized system.&lt;/p&gt;

&lt;p&gt;FCoP therefore does not try to reinvent a complex workflow engine. Instead, it compresses the most important responsibility relationships in multi-Agent work into a small set of stable facts, then lets the protocol control when those facts are allowed to change.&lt;/p&gt;

&lt;p&gt;At Build 2026, Microsoft introduced ASSERT and the Agent Control Specification (ACS), which similarly separate “finding a problem” from “exerting control along the execution path”: evaluation discovers defects first, while controls are then applied at key points such as input, state, tool execution, and output. &lt;a href="https://devblogs.microsoft.com/foundry/build-2026-open-trust-stack-ai-agents/" rel="noopener noreferrer"&gt;ASSERT and ACS technical overview&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;FCoP is not an implementation of ACS, and the two systems differ in architecture and scope. But they reflect a similar engineering principle: &lt;strong&gt;governance cannot exist only after the fact, and it cannot be reduced to logs; it must be able to constrain work at the moment the work actually moves forward.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;FCoP also does not claim that filesystems should replace all infrastructure. Cross-machine high-throughput messaging, complex queries, large event streams, distributed transactions, and weakly consistent network environments are still better served by databases, message queues, or specialized coordination systems. FCoP addresses a more foundational problem inside a project workspace: how can tasks, deliverables, reviews, and responsibility relationships remain directly understandable to humans while also being re-verifiable by machines?&lt;/p&gt;




&lt;h2&gt;
  
  
  2. FCoP stores work facts, not conversations
&lt;/h2&gt;

&lt;p&gt;FCoP does not try to preserve every Agent conversation, and it does not need to retain every internal reasoning trace. Its concern is narrower: when work moves from one role to another, which confirmed facts does the next responsible actor actually need?&lt;/p&gt;

&lt;p&gt;The core protocol records can be summarized as four types:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;What it stores&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TASK&lt;/td&gt;
&lt;td&gt;Work content, delegation relationships, and lifecycle&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REPORT&lt;/td&gt;
&lt;td&gt;A deliverable produced by one execution attempt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ISSUE&lt;/td&gt;
&lt;td&gt;An unresolved problem discovered during execution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;REVIEW&lt;/td&gt;
&lt;td&gt;A review and decision about a specific deliverable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A simplified multi-Agent software team might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ADMIN ⇄ PM                Requirements and management decisions
          │
          ▼
      FCoP protocol layer
      ├─ PM: creates and manages TASKs, reads deliverables and reviews
      ├─ DEV: claims TASKs and produces REPORT / ISSUE
      └─ QA : reads TASK and REPORT, then produces REVIEW
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The roles above are only an example of one collaboration model; they are not a fixed role set defined by FCoP Core.&lt;/p&gt;

&lt;p&gt;The key change is not simply replacing chat with Markdown. It is separating &lt;strong&gt;a role’s own runtime context&lt;/strong&gt; from &lt;strong&gt;the formal work facts the project depends on collectively&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A PM does not need to enter a DEV conversation and read dozens of chat turns, and DEV does not need to verbally retell the entire implementation process to QA after the work is finished. What is handed off is not the previous conversation but the TASK, REPORT, ISSUE, and REVIEW records that have entered the protocol layer.&lt;/p&gt;

&lt;p&gt;An Agent runtime can end, a model can be replaced, and a conversation can disappear entirely, while the work itself remains. A later responsible actor can reread the protocol files and still determine where the task came from, what stage it reached, what was delivered, and what reviews actually occurred.&lt;/p&gt;

&lt;p&gt;This is one of the most important differences between FCoP and many “multi-Agent chat frameworks”: &lt;strong&gt;collaboration facts do not depend on a conversation continuing to exist.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6zwktudwvd8beqeucmxj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6zwktudwvd8beqeucmxj.png" alt="FCoP 4.0 Protocol Architecture" width="800" height="1000"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Figure 1 — FCoP 4.0 Protocol Architecture.&lt;/strong&gt; FCoP is accessed through Agent Host environments, protocol tools expose standardized capabilities, FCoP Core performs protocol checks and constraints, and confirmed facts are persisted as files and lifecycle state.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  3. The five buckets are protocol state, not directory convention
&lt;/h2&gt;

&lt;p&gt;One of the easiest misunderstandings about FCoP is to treat &lt;code&gt;inbox&lt;/code&gt;, &lt;code&gt;active&lt;/code&gt;, &lt;code&gt;review&lt;/code&gt;, &lt;code&gt;done&lt;/code&gt;, and &lt;code&gt;archive&lt;/code&gt; as ordinary folders used only to organize files.&lt;/p&gt;

&lt;p&gt;In fact, they represent the lifecycle state of a TASK.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;_lifecycle/
├─ inbox/
├─ active/
├─ review/
├─ done/
└─ archive/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;inbox&lt;/code&gt; means that a task has entered the system but has not yet been claimed. &lt;code&gt;active&lt;/code&gt; means that the task is being executed. &lt;code&gt;review&lt;/code&gt; means that the current execution has produced a deliverable and is waiting for review. &lt;code&gt;done&lt;/code&gt; means that the task has satisfied its completion conditions. &lt;code&gt;archive&lt;/code&gt; stores finished tasks that have left the current working view.&lt;/p&gt;

&lt;p&gt;A typical TASK lifecycle can therefore be represented as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;inbox → active → review → done → archive
           ↑         │
           └─────────┘
               rework
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important idea is not the existence of five directories but the fact that &lt;strong&gt;state is materialized by location&lt;/strong&gt;. The lifecycle location of a TASK corresponds to its protocol stage, so the system does not need to rely on an Agent’s own claim that “I am at this step now.”&lt;/p&gt;

&lt;p&gt;But this does &lt;strong&gt;not&lt;/strong&gt; mean that an Agent may move lifecycle files directly. Quite the opposite: if an Agent could move a TASK from &lt;code&gt;active&lt;/code&gt; to &lt;code&gt;done&lt;/code&gt; simply because it believed the work was complete, the five buckets would be only a directory convention rather than a governance protocol.&lt;/p&gt;

&lt;p&gt;FCoP adds a stricter rule: &lt;strong&gt;an Agent may propose an action, but it may not directly change lifecycle facts.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An executor may request to claim a task or submit a deliverable; a reviewer may submit a REVIEW or approve a result. Whether those actions are sufficient to move the TASK to the next stage must be re-validated by FCoP against the current state, the current execution attempt, the deliverable, the review, and the relevant authorization relationships. Only when the required conditions hold does the protocol layer execute the lifecycle transition and materialize the new state in the filesystem.&lt;/p&gt;

&lt;p&gt;The process can be summarized as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent proposes a protocol action
        ↓
FCoP checks current facts and protocol conditions
        ↓
      allow / reject
        ↓
if allowed, execute lifecycle transition
        ↓
new file location becomes the new observable state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So “filesystem as protocol” does not mean that the filesystem itself makes decisions. It means that &lt;strong&gt;the filesystem carries and materializes protocol facts, while FCoP determines which fact changes are allowed to become valid.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is what makes FCoP a protocol rather than merely a Markdown format.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. The five buckets answer “where”; attempt answers “which execution”
&lt;/h2&gt;

&lt;p&gt;The five buckets tell the system which lifecycle stage a TASK is in, but they do not answer another important question: if the same task is sent back and executed again, which execution attempt does the current deliverable belong to?&lt;/p&gt;

&lt;p&gt;That is the purpose of &lt;code&gt;attempt_id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;When a TASK begins a new execution, the system establishes an &lt;code&gt;attempt_id&lt;/code&gt; for that attempt. If the task is rejected during review and returns to execution, a new attempt is created.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;task_id&lt;/code&gt; and &lt;code&gt;attempt_id&lt;/code&gt; therefore answer two different questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;task_id&lt;/code&gt;: Is this the same piece of work?&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;attempt_id&lt;/code&gt;: Is this the same execution attempt?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK-A
│
├─ attempt-1
│    ├─ REPORT-1
│    └─ REVIEW-1
│
└─ attempt-2
     └─ REPORT-2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;TASK-A remains the same task, but &lt;code&gt;attempt-1&lt;/code&gt; and &lt;code&gt;attempt-2&lt;/code&gt; are two different executions.&lt;/p&gt;

&lt;p&gt;This distinction matters because real work rarely passes on the first try. The first implementation may contain problems; the second attempt may produce a new REPORT; a third round of rework may follow. The task identity remains unchanged, but the facts produced by each execution are different.&lt;/p&gt;

&lt;p&gt;If the system knows only &lt;code&gt;task_id&lt;/code&gt;, it can easily apply the first attempt’s REPORT, test results, or REVIEW to the second attempt by mistake. FCoP separates the attempts with &lt;code&gt;attempt_id&lt;/code&gt;, so the first attempt’s REPORT and REVIEW remain valid historical evidence without automatically becoming evidence for the second attempt.&lt;/p&gt;

&lt;p&gt;The division of responsibility is simple: &lt;strong&gt;the five buckets tell us where the task is; &lt;code&gt;attempt&lt;/code&gt; tells us which execution of the task we are looking at.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Separate identity, relationships, and state
&lt;/h2&gt;

&lt;p&gt;File naming is also a useful way to understand how FCoP has evolved.&lt;/p&gt;

&lt;p&gt;Earlier designs placed more human-readable routing information directly in filenames, including sender and recipient:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK-...-ADMIN-to-PM.md
TASK-...-PM-to-DEV.md
REPORT-...-DEV-to-PM-....md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is intuitive because a person opening a directory can often tell who sent a piece of work to whom without reading the file. But it also means that record identity, role routing, and business semantics gradually accumulate inside one filename. If the team’s role model changes, the naming rules have to change with it, and the protocol becomes more tightly coupled to one organizational structure.&lt;/p&gt;

&lt;p&gt;The design direction in FCoP 4.x is to separate those responsibilities. Conceptually, record files can be understood as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK-&amp;lt;stable-id&amp;gt;.md
REPORT-&amp;lt;stable-id&amp;gt;.md
REVIEW-&amp;lt;stable-id&amp;gt;.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here &lt;code&gt;&amp;lt;stable-id&amp;gt;&lt;/code&gt; illustrates the idea of stable identity rather than enumerating every exact filename rule of a specific release. The filename primarily answers “what type of record is this, and which record is it?” Sender, recipient, related task, and execution attempt are expressed through structured fields instead.&lt;/p&gt;

&lt;p&gt;For example, a REPORT may include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;sender&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DEV&lt;/span&gt;
&lt;span class="na"&gt;recipient&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;PM&lt;/span&gt;
&lt;span class="na"&gt;subject_ref&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;TASK-...&lt;/span&gt;
&lt;span class="na"&gt;attempt_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;urn:uuid:...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A PM therefore does not need to infer relevance from a &lt;code&gt;DEV-to-PM&lt;/code&gt; filename string. The protocol relationships can directly express who created the record, who receives it, which task it belongs to, and which execution attempt produced it.&lt;/p&gt;

&lt;p&gt;The design can be summarized in one sentence: &lt;strong&gt;the filename carries identity, structured fields carry relationships and routing, and the lifecycle path carries state.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The same principle extends to operations. Tool calls in Agent systems may be resubmitted because of timeouts. If the system cannot determine whether a request is “the same operation as before,” an ordinary retry may create a second piece of work.&lt;/p&gt;

&lt;p&gt;FCoP therefore uses &lt;code&gt;operation_id&lt;/code&gt; to identify an operation. The same &lt;code&gt;operation_id&lt;/code&gt; with identical content returns the existing result, while reusing the same &lt;code&gt;operation_id&lt;/code&gt; with changed content is rejected.&lt;/p&gt;

&lt;p&gt;In our experiment, the same creation request was submitted 25 times and produced only one TASK. When the body was changed while keeping the same &lt;code&gt;operation_id&lt;/code&gt;, the system returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OPERATION_ID_CONFLICT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The underlying principle remains the same: establish stable identity first, then build relationships and state constraints around that identity. A network retry should not silently become another task.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. YAML expresses responsibility relationships, not merely configuration
&lt;/h2&gt;

&lt;p&gt;FCoP uses Markdown for human-readable content and YAML frontmatter for protocol relationships. The core structure of a REPORT can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;protocol&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fcop&lt;/span&gt;
&lt;span class="na"&gt;version&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;4&lt;/span&gt;
&lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;REPORT&lt;/span&gt;
&lt;span class="na"&gt;report_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;REPORT-...&lt;/span&gt;
&lt;span class="na"&gt;workspace_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;urn:uuid:...&lt;/span&gt;
&lt;span class="na"&gt;sender&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DEV&lt;/span&gt;
&lt;span class="na"&gt;recipient&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;PM&lt;/span&gt;
&lt;span class="na"&gt;subject_ref&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;TASK-...&lt;/span&gt;
&lt;span class="na"&gt;attempt_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;urn:uuid:...&lt;/span&gt;
&lt;span class="na"&gt;report_kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;final&lt;/span&gt;
&lt;span class="na"&gt;result&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;done&lt;/span&gt;
&lt;span class="na"&gt;references&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[]&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;Deliverable after the second round of rework.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At first glance, this looks like ordinary structured Markdown. But the important point is not the number of fields; it is the responsibility relationships those fields establish.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;report_id&lt;/code&gt; identifies the deliverable itself. &lt;code&gt;subject_ref&lt;/code&gt; identifies the TASK it belongs to. &lt;code&gt;attempt_id&lt;/code&gt; identifies which execution produced it. &lt;code&gt;sender&lt;/code&gt; and &lt;code&gt;recipient&lt;/code&gt; express where the record came from and who it is handed to.&lt;/p&gt;

&lt;p&gt;REVIEW follows the same logic. A REVIEW cannot consist only of the word &lt;code&gt;approved&lt;/code&gt;, because the protocol needs to know not merely whether the word “approved” exists, but who approved what, for which task, in which execution attempt, against which REPORT, and whether that approval is still applicable to the current state change.&lt;/p&gt;

&lt;p&gt;A REVIEW must therefore be related explicitly to the appropriate TASK, attempt, and reviewed REPORT, while authorization itself also has a defined scope. In this way, &lt;code&gt;approved&lt;/code&gt; is no longer an isolated natural-language statement but a work fact whose applicability can be checked again.&lt;/p&gt;

&lt;p&gt;That is why, in FCoP, &lt;strong&gt;“this task passed review before” and “the current deliverable has passed review” are not the same statement.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  7. The protocol matters most when it prevents an invalid state from becoming valid
&lt;/h2&gt;

&lt;p&gt;The value of FCoP is not merely that the correct workflow can proceed smoothly. More importantly, when the required conditions do not hold, an incorrect action cannot become a system fact simply because an Agent believes it should.&lt;/p&gt;

&lt;p&gt;An Agent may say “I am done,” and it may request that a task move into review, but the protocol does not accept the transition merely because the natural-language statement exists. FCoP rechecks the current TASK, attempt, REPORT, REVIEW, and authorization relationships. The lifecycle moves forward only when the current facts support the requested transition.&lt;/p&gt;

&lt;p&gt;For example, an executor may claim a task through an operation such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;project&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transition&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;from_stage&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inbox&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;to_stage&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;active&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claim_task&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;actor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;actor&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is not these few lines of code but whether the preconditions are still true at the moment of execution.&lt;/p&gt;

&lt;p&gt;If one process has already claimed the TASK and moved it from &lt;code&gt;inbox&lt;/code&gt; to &lt;code&gt;active&lt;/code&gt;, another concurrent process that attempts the same claim will cause FCoP to reread the current facts, discover that the original &lt;code&gt;inbox → active&lt;/code&gt; precondition no longer exists, and reject the second transition.&lt;/p&gt;

&lt;p&gt;Likewise, if the current attempt has not yet produced a REPORT, the task cannot move into review out of thin air. If a REVIEW belongs to a previous attempt’s REPORT, it cannot approve the current attempt’s new deliverable.&lt;/p&gt;

&lt;p&gt;This is the purpose of the protocol: &lt;strong&gt;an Agent may propose a judgment or request, but it may not turn its own judgment directly into a system fact.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Experiment: what can these Markdown and directory constraints actually stop?
&lt;/h2&gt;

&lt;p&gt;To verify that these constraints are executable rather than merely descriptive, we ran several local experiments against FCoP 4.0.3. The test environment used a Windows local filesystem, Python 3.10.11, and &lt;code&gt;fcop==4.0.3&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The main results were:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10 rounds, 4 processes claiming concurrently in each round&lt;/td&gt;
&lt;td&gt;10 successes, 30 &lt;code&gt;INVALID_TRANSITION&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Submit for review without a current REPORT&lt;/td&gt;
&lt;td&gt;&lt;code&gt;REPORT_REQUIRED&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use a REPORT from another execution attempt&lt;/td&gt;
&lt;td&gt;&lt;code&gt;ATTEMPT_MISMATCH&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reuse an old REVIEW to approve a new deliverable after rework&lt;/td&gt;
&lt;td&gt;&lt;code&gt;AUTHORIZATION_INVALID&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first test focused on concurrent claiming. In each round, four processes competed for the same TASK. Across ten rounds, this produced forty claim attempts. Exactly one process succeeded in each round; the other three received &lt;code&gt;INVALID_TRANSITION&lt;/code&gt;. The final result was ten successes and thirty rejections, with no case in which the same TASK was successfully claimed by multiple executors.&lt;/p&gt;

&lt;p&gt;The second and third tests focused on evidence boundaries. If the current attempt has not produced a REPORT, the system does not allow submission for review and instead returns &lt;code&gt;REPORT_REQUIRED&lt;/code&gt;. If a REPORT from another attempt is used as the current deliverable, the system returns &lt;code&gt;ATTEMPT_MISMATCH&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The fourth test shows the governance meaning of rework most clearly.&lt;/p&gt;

&lt;p&gt;Suppose the first execution produces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK-A
  └─ attempt-1
       └─ REPORT-1
            └─ REVIEW-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After QA reviews &lt;code&gt;REPORT-1&lt;/code&gt;, a problem is found and the task is returned for rework. A new execution attempt is created:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TASK-A
  └─ attempt-2
       └─ REPORT-2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;TASK-A&lt;/code&gt; is still the same task, but the execution has changed from &lt;code&gt;attempt-1&lt;/code&gt; to &lt;code&gt;attempt-2&lt;/code&gt;, and the deliverable has changed from &lt;code&gt;REPORT-1&lt;/code&gt; to &lt;code&gt;REPORT-2&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The old &lt;code&gt;REVIEW-1&lt;/code&gt; remains a real historical record, but it proves only that &lt;code&gt;REPORT-1&lt;/code&gt; from the first attempt was reviewed. It does not prove that &lt;code&gt;REPORT-2&lt;/code&gt; from the second attempt has been reviewed.&lt;/p&gt;

&lt;p&gt;If the system tries to use the old REVIEW to advance the new deliverable, FCoP returns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AUTHORIZATION_INVALID
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rule protects a very simple responsibility boundary: &lt;strong&gt;an old review is a real historical fact, but it is not authorization for the current deliverable.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Files can of course be read, and a body may contain the words “done” or “approved,” but those words do not automatically advance the lifecycle. What determines whether the work may continue is whether the current state, current execution, current deliverable, and current authorization are consistent.&lt;/p&gt;




&lt;h2&gt;
  
  
  9. What does “filesystem as protocol” actually mean?
&lt;/h2&gt;

&lt;p&gt;At this point, the phrase “filesystem as protocol” can be understood more precisely.&lt;/p&gt;

&lt;p&gt;It does not mean saving every conversation as Markdown. It does not mean allowing Agents to move files between directories on their own. And it does not mean that the filesystem itself makes business decisions.&lt;/p&gt;

&lt;p&gt;FCoP becomes a protocol because several constraints coexist: records have stable identity, structured fields preserve explicit relationships, the five buckets materialize lifecycle state, attempts separate repeated executions, and every state change must pass protocol actions and condition checks.&lt;/p&gt;

&lt;p&gt;In other words, &lt;strong&gt;Markdown stores work content, structured fields store responsibility relationships, file paths materialize lifecycle state, and the FCoP protocol determines when those state changes are allowed to occur.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The “minimalism” of FCoP does not mean the absence of rules, nor does it weaken governance requirements. Instead, it places complexity where it actually belongs: identity, relationships, execution attempts, and transition boundaries, rather than concentrating every concern inside one centralized workflow engine.&lt;/p&gt;

&lt;p&gt;That is why FCoP can rely on ordinary files and Markdown while still establishing real governance constraints.&lt;/p&gt;




&lt;h2&gt;
  
  
  10. What FCoP solves — and what it does not
&lt;/h2&gt;

&lt;p&gt;FCoP does not replace testing. Nor can it infer that tests were actually executed simply because a REPORT says “all tests passed.” It does not replace QA, human review, or independent evaluation, and it does not automatically determine whether a technical claim is true.&lt;/p&gt;

&lt;p&gt;Content-level truth still requires testing systems, code review, QA, independent evaluation, or other evidence mechanisms.&lt;/p&gt;

&lt;p&gt;FCoP addresses a more basic layer that must be correct before those judgments can be trusted: which task are we dealing with, which execution produced this deliverable, which REPORT did the current REVIEW actually examine, and does this authorization still apply to the current work?&lt;/p&gt;

&lt;p&gt;These questions may not sound “intelligent.” In fact, they are deliberately conventional. But if these basic relationships are uncertain, even the strongest model, Agent, or workflow system can continue operating on the wrong object.&lt;/p&gt;




&lt;h2&gt;
  
  
  11. Conclusion: the value of a protocol is that an Agent cannot define system facts by itself
&lt;/h2&gt;

&lt;p&gt;Agent systems are becoming more complex very quickly. There are more models, more tools, more roles, more autonomous execution, and deeper cross-system integration. It is easy to assume that governing such systems therefore requires an equally complex centralized infrastructure.&lt;/p&gt;

&lt;p&gt;FCoP proposes another answer.&lt;/p&gt;

&lt;p&gt;It does not introduce a new database or a large orchestration language. Instead, it builds on things developers have used for decades: files, directories, Markdown, and a small number of stable structural relationships.&lt;/p&gt;

&lt;p&gt;Filenames carry identity, structured fields carry relationships and routing, the five buckets materialize lifecycle state, attempts separate repeated executions, REPORT and REVIEW preserve delivery and review evidence, and protocol gates determine when those facts are sufficient to support the next transition.&lt;/p&gt;

&lt;p&gt;Our experiments show that, under the tested conditions of a Windows local filesystem, Python 3.10.11, and &lt;code&gt;fcop==4.0.3&lt;/code&gt;, constraints around identity, concurrent claiming, execution attempts, and authorization scope can be reproduced consistently. These results are not performance claims about cross-machine coordination, network filesystems, or large-scale throughput. But they demonstrate one important point: &lt;strong&gt;industrial-grade enforceability does not necessarily require industrial-grade complexity.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For multi-Agent teams, reliable collaboration may begin not with sending more messages between Agents, but with ensuring that no Agent can turn “I think this is finished” directly into “the system says this is finished.”&lt;/p&gt;

&lt;p&gt;An Agent may propose an action; FCoP validates it. Once the protocol accepts the transition, the filesystem preserves the new fact.&lt;/p&gt;

&lt;p&gt;That may be the most important meaning of “filesystem as protocol”: &lt;strong&gt;the work may be executed by Agents, but the facts cannot be declared valid by the Agents themselves.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;a href="https://zenodo.org/records/22855630" rel="noopener noreferrer"&gt;FCoP Paper (Zenodo)&lt;/a&gt; · &lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;FCoP Protocol and Examples&lt;/a&gt; · &lt;a href="https://pypi.org/project/fcop/" rel="noopener noreferrer"&gt;FCoP on PyPI&lt;/a&gt; · &lt;a href="https://pypi.org/project/fcop-mcp/" rel="noopener noreferrer"&gt;FCoP MCP on PyPI&lt;/a&gt; · &lt;a href="https://blog.csdn.net/m0_51507544/article/details/166249570" rel="noopener noreferrer"&gt;Chinese Edition (CSDN)&lt;/a&gt; · &lt;a href="https://juejin.cn/post/7687627812326457353" rel="noopener noreferrer"&gt;Chinese Edition (Juejin)&lt;/a&gt; · &lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
      <category>architecture</category>
    </item>
    <item>
      <title>When Multi-Agent Systems Go Out of Control: What Agent Governance Do We Need in 2026?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Fri, 18 Sep 2026 09:10:14 +0000</pubDate>
      <link>https://dev.to/joinwell52/when-multi-agent-systems-go-out-of-control-what-agent-governance-do-we-need-in-2026-227c</link>
      <guid>https://dev.to/joinwell52/when-multi-agent-systems-go-out-of-control-what-agent-governance-do-we-need-in-2026-227c</guid>
      <description>&lt;p&gt;In 2026, developers are giving AI more than questions. They are giving it work.&lt;/p&gt;

&lt;p&gt;Breaking down requirements, writing code, running tests and reviewing results—steps once connected by people are increasingly assigned to different agents. A Planner develops a plan, a Coder implements it, and a Reviewer examines the delivery. That is the attraction of a digital team: a person sets the objective, multiple agents divide the work, and progress continues.&lt;/p&gt;

&lt;p&gt;But delegating the tasks does not make management disappear. Scope, progress, disagreements and acceptance conditions once held in people's heads now need to remain clear across independently acting roles.&lt;/p&gt;

&lt;p&gt;Why did an agent suddenly change the scope? Which of two conflicting reports should we trust? When an executor says “done,” has anyone actually checked? These questions lead to one concern: &lt;strong&gt;once work is delegated to AI, who retains control of its boundaries and progress?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Tool calls can become more frequent while collaboration becomes harder to explain. That gap is where anxiety about multi-agent governance begins.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Beneath the Momentum: Communication Confusion and Governance Anxiety in 2026
&lt;/h3&gt;

&lt;p&gt;Consider a collaboration chain: a planner turns an ambiguous objective into tasks, an executor modifies code according to its interpretation, and a reviewer reads only the execution summary before declaring a pass. Every role has produced a response, yet the work may have drifted away from the original assignment. Finding the divergence requires tracing the instructions each participant received, the actions taken and the material behind each judgment.&lt;/p&gt;

&lt;p&gt;This kind of loss of control rarely starts dramatically. It hides in seemingly smooth handoffs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The message arrived, but the evidence did not.&lt;/strong&gt; The next participant knows that someone said the work was finished, but not what was checked or missed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tools have permissions, but the assignment lacks boundaries.&lt;/strong&gt; Permission to inspect a project is interpreted as permission to modify it. A limited authorization expands along the task chain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The indicator turned green, but responsibility remains unresolved.&lt;/strong&gt; Code written, report submitted, tests passed and formal acceptance collapse into a single “done.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe05ynxmliij0baa471xh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe05ynxmliij0baa471xh.png" alt="Scope, checks and review targets across a handoff" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Receiving a message does not mean the supporting evidence is complete. The roles illustrate one collaboration pattern.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Governance therefore starts with a straightforward requirement: &lt;strong&gt;make both “what happened” and “what justifies continuing” possible to establish.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In its September 2026 responsible AI statement, Microsoft moves governance further into system operation: organizations need to see what agents are doing, test their behavior and intervene when necessary. &lt;a href="https://blogs.microsoft.com/on-the-issues/2026/09/01/responsible-ai-in-2026-how-we-are-adapting-for-whats-ahead/" rel="noopener noreferrer"&gt;Microsoft responsible AI statement&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That direction has concrete engineering expressions. ASSERT translates organizational policies into evaluation scenarios. The Agent Control Specification (ACS) defines controls at key points including input, model, state, tool execution and output. Microsoft describes a cycle of identifying problems through evaluation, configuring controls and evaluating the improvement. &lt;a href="https://devblogs.microsoft.com/foundry/build-2026-open-trust-stack-ai-agents/" rel="noopener noreferrer"&gt;ASSERT and ACS technical explanation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A fix in the OpenAI Agents SDK makes the issue concrete. In September 2026, contributor jbeckwith-oai submitted a fix describing how, when an output-checking program threw an exception without reaching a verdict, the run could fail while the unvetted answer was still saved to the session and sent back to the model in the next turn. The fix routes that output through the path that prevents persistence, keeping it out of subsequent context. &lt;a href="https://github.com/openai/openai-agents-js/pull/1938" rel="noopener noreferrer"&gt;OpenAI Agents JS #1938&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A run ending does not mean everything it leaves behind is qualified to become evidence for subsequent work.&lt;/strong&gt; In multi-agent collaboration, we must ask whether handed-off content carries its checking status, and whether the recipient can distinguish a verified conclusion from output whose checks never completed.&lt;/p&gt;

&lt;p&gt;Further questions follow: whose delivery was checked, and which attempt produced it? Which result does an approval cover? After someone else takes over, can the original basis still be inspected?&lt;/p&gt;

&lt;p&gt;The problem has grown beyond getting agents to communicate. It is now &lt;strong&gt;how to leave a basis for the work that later participants can inspect and use.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Why Do Traditional Enterprise Governance Frameworks Still Need an Agent Layer?
&lt;/h3&gt;

&lt;p&gt;Enterprise software already has access control, audit logs, database transactions and service gateways. Why does organizing agent work require additional governance rules?&lt;/p&gt;

&lt;p&gt;Because technical operations and business assignments are connected by relationships that must be explicitly represented. A gateway can check whether a request carries credentials without knowing whether the action belongs to the current assignment. A database can store task state, but business rules must specify who may change it and what evidence must exist first.&lt;/p&gt;

&lt;p&gt;When an agent adapts its execution path to intermediate results, those questions recur. A successful tool call establishes that an action ran. Determining whether the overall process is valid requires placing that action back into its task, role and authorization relationships.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first gap is the distance between technical permission and responsibility for the work.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Being able to write a file is a capability. Being entitled to accept a delivery is a responsibility. Executors, reviewers and approvers should not acquire identical decision rights merely because they use the same tools.&lt;/p&gt;

&lt;p&gt;Without explicit task relationships, teams must reconstruct events from logs: who requested this step, whether a modification exceeded the assignment, and why work advanced without review. The more autonomous the system and the longer the chain, the harder that reconstruction becomes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The second gap is how a human decision takes effect after intervention.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A stop button addresses only part of the problem. The system must also know which work has been submitted, which actions have not yet run, which results should be preserved and where execution should resume.&lt;/p&gt;

&lt;p&gt;Human approval raises similar questions. For delivery acceptance, approval needs to refer to a specific delivery and its scope. If the executor subsequently changes the content, does the original approval still cover the current result?&lt;/p&gt;

&lt;p&gt;Without that relationship, “someone clicked approve” establishes that an event occurred, but provides a weak basis for later work.&lt;/p&gt;

&lt;p&gt;Both gaps point to the same issue: a system must understand what an action means within an assignment, not merely that it happened. Generating a report does not establish formal delivery. An approval cannot remain valid independently of its object.&lt;/p&gt;

&lt;p&gt;For long-running, continuous, multi-role work, these relationships cannot depend indefinitely on chat history and retrospective explanation. Tasks need identities, deliveries need supporting material, reviews need accountable roles, and key transitions need explicit conditions.&lt;/p&gt;

&lt;p&gt;The next question is therefore: &lt;strong&gt;how should these working relationships be preserved so that later participants can read, check and use them?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Industry and Frontier Engineering: Why Are Text and File Contracts Worth Attention?
&lt;/h3&gt;

&lt;p&gt;Following that line of thought, human-readable work contracts become valuable.&lt;/p&gt;

&lt;p&gt;GitHub Agentic Workflows lets developers describe work in Markdown with YAML frontmatter, compile it into an execution environment, and apply permissions and controls over permitted write operations. The work definition can be read directly while execution mechanisms enforce constraints. &lt;a href="https://docs.github.com/en/copilot/concepts/agents/about-github-agentic-workflows" rel="noopener noreferrer"&gt;GitHub documentation&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anthropic's engineering article on its multi-agent research system describes a different handoff practice: saving some artifacts externally and returning references to the coordinating agent, reducing information loss from repeatedly passing results through intermediaries. &lt;a href="https://www.anthropic.com/engineering/multi-agent-research-system" rel="noopener noreferrer"&gt;Anthropic engineering article&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The two approaches operate at different levels but touch a shared need: &lt;strong&gt;work must leave material that the next participant can inspect again.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Choosing Markdown does not require rejecting other structures. In its work on long-running agents, Anthropic used JSON for a feature list, together with progress material and tests supporting subsequent sessions. &lt;a href="https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents" rel="noopener noreferrer"&gt;Long-running agent engineering&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The important question is whether work records can outlive a session and be read, referenced and checked again—not whether the medium is Markdown, YAML, JSON or a database. Sessions end, models change and executors are replaced. The reason a task exists, its current delivery, who checked it and on what basis should survive those changes.&lt;/p&gt;

&lt;p&gt;That addresses preservation. Governance requires a further step.&lt;/p&gt;

&lt;p&gt;Material used only for retrospective review supports recording and traceability. Material that also determines whether work may proceed starts to participate in governance. Whether a task may advance, whether an approval remains applicable and whether an old check covers a new delivery cannot depend solely on remembering what someone previously said.&lt;/p&gt;

&lt;p&gt;A more important question follows: &lt;strong&gt;should governance rules depend on agents remembering them, or become explicit rules shared by all participants?&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Why Does Governance Ultimately Lead Toward a Protocol?
&lt;/h3&gt;

&lt;p&gt;Context reminders and prompt instructions alone struggle to preserve governance rules across long tasks and repeated handoffs. Context grows, is compressed and is truncated. Sessions end, models change and different roles interpret the same requirement differently. Prompts remain important, but they primarily tell an individual executor what it should do.&lt;/p&gt;

&lt;p&gt;Governance must answer another question: &lt;strong&gt;regardless of which model executes today or which agent takes over tomorrow, under what conditions may the work continue?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A rule that depends solely on model recall resembles an instruction. When it is explicit, shared by participants and checked by the system at critical operations, it starts to become part of governance.&lt;/p&gt;

&lt;p&gt;That is why governance needs a protocol.&lt;/p&gt;

&lt;p&gt;A protocol is more than a longer prompt, and it need not encode every business judgment. It defines the working rules: what constitutes a task and a formal delivery, which result a review covers, who may make which decisions, and which missing conditions require the system to refuse further progress.&lt;/p&gt;

&lt;p&gt;Organizations use procedures, rules and contracts for a related reason: important obligations cannot depend solely on each participant remembering them. Extending execution from people to models makes that concern more pronounced.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A workflow tells an agent where to go next; a protocol establishes what entitles it to take that step.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;These responsibilities can coexist in one system: the process organizes execution, while the protocol specifies the conditions participants must satisfy to advance.&lt;/p&gt;

&lt;p&gt;FCoP—the File-based Coordination Protocol—starts with this question.&lt;/p&gt;

&lt;p&gt;It does not prescribe how an agent thinks or replace a reviewer's judgment about code quality. It specifies how work is established: tasks need persistent identities, deliveries must relate to particular tasks and execution attempts, reviews must identify their subjects, and transitions requiring authorization need the corresponding basis. A statement that the work is “done” cannot substitute for missing conditions.&lt;/p&gt;

&lt;p&gt;FCoP uses ordinary files to hold these records. TASK, REPORT, ISSUE and REVIEW represent tasks, delivery reports, problems and reviews. People can open them; agents can read and inspect them. Files preserve records and evidence, while the protocol defines their relationships and the conditions under which they support further action.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe3vfdw5z45bx5v0l4zzx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe3vfdw5z45bx5v0l4zzx.png" alt="Linked records and checks at the operation boundary" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The four record types are related records, not sequential lifecycle stages. Allowing an operation does not establish that its deliverable is correct.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This apparently simple design immediately faces difficult questions.&lt;/p&gt;

&lt;p&gt;If network retries send the same creation request 25 times, should there be 25 tasks or just one? If four processes claim the same task simultaneously, can all four believe they succeeded? If a reviewer checks one delivery and the executor subsequently produces a new result, does the old REVIEW still apply? Can work advance without a REPORT, a required review or authorization?&lt;/p&gt;

&lt;p&gt;There is an even less intuitive question: &lt;strong&gt;if agents never chat with one another and exchange work records only through TASK, REPORT, ISSUE and REVIEW, can they still form a team capable of sustained collaboration?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is what makes FCoP worth investigating. The interesting question is whether &lt;strong&gt;rules expressed through ordinary files can actually constrain multi-agent work.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  5. What Matters Is Whether It Can Hold These Boundaries
&lt;/h3&gt;

&lt;p&gt;FCoP proposes a deceptively simple answer: preserve multi-agent work records and collaboration rules in ordinary files, and check those rules at critical operations.&lt;/p&gt;

&lt;p&gt;Retries, competing claims, changed deliveries and missing evidence are more revealing tests than a smooth demonstration. Does the system reuse an existing result when it should? Does it reject an invalid request? Can it stop when conditions are missing? Those behaviors establish what the protocol actually constrains.&lt;/p&gt;

&lt;p&gt;If these boundaries hold, “files as protocol” becomes more than a design slogan. &lt;strong&gt;It may mean that the order governing multi-agent work need not remain hidden inside a model, a prompt or a large central system.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/FCoP" rel="noopener noreferrer"&gt;FCoP protocol and examples&lt;/a&gt; · &lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
      <category>programming</category>
    </item>
    <item>
      <title>The quota is exhausted. Why is the agent still retrying?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:56:43 +0000</pubDate>
      <link>https://dev.to/joinwell52/the-quota-is-exhausted-why-is-the-agent-still-retrying-5d2o</link>
      <guid>https://dev.to/joinwell52/the-quota-is-exhausted-why-is-the-agent-still-retrying-5d2o</guid>
      <description>&lt;h1&gt;
  
  
  The quota is exhausted. Why is the agent still retrying?
&lt;/h1&gt;

&lt;p&gt;The agent stops. Its provider says the quota is exhausted and gives a reset time for later that evening.&lt;/p&gt;

&lt;p&gt;Then the task retries. The counter rises; the work goes nowhere. Which part of the message did the system fail to understand?&lt;/p&gt;

&lt;p&gt;One possibility is that the provider supplied a cause, while the program recorded only “this step failed.” “The warehouse is closed today” became “delivery failed,” so another delivery was scheduled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where did the cause get lost?
&lt;/h2&gt;

&lt;p&gt;The source is &lt;a href="https://github.com/paperclipai/paperclip/pull/13549" rel="noopener noreferrer"&gt;Paperclip #13549&lt;/a&gt;, by electrumnz. Paperclip coordinates ongoing AI work, including recovery after failures.&lt;/p&gt;

&lt;p&gt;The author reported 320 quota-related terminal runs in a production review, 186 carrying retry links, while quota reasons were missing. That is upstream evidence; we did not access the production data. It caught our attention because trying again may help a temporary network failure, while repeatedly trying before quota resets may achieve nothing.&lt;/p&gt;

&lt;p&gt;To choose a next step, the program must first identify the cause. The patch adds a previously missed entry point: when a message arrives under a failed-turn error code, the classifier also checks it for quota information.&lt;/p&gt;

&lt;p&gt;The question is therefore not just whether the provider explained the failure. &lt;strong&gt;Did that explanation reach the code responsible for producing a waiting time?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Send the same sentence through again
&lt;/h2&gt;

&lt;p&gt;We extracted the original classification code and parsing helpers from both revisions and supplied the same nine inputs. One described exhausted quota and specified when it would reset.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;Candidate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Quota message with a reset time, through the failed-turn entry point&lt;/td&gt;
&lt;td&gt;No classification&lt;/td&gt;
&lt;td&gt;Quota recognized; reset time calculated from the message&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quota message without a reset time&lt;/td&gt;
&lt;td&gt;No classification&lt;/td&gt;
&lt;td&gt;Quota recognized; default one-hour wait&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ordinary socket failure&lt;/td&gt;
&lt;td&gt;No classification&lt;/td&gt;
&lt;td&gt;Still no classification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Missing key through the existing adapter-failure entry point&lt;/td&gt;
&lt;td&gt;Configuration incomplete&lt;/td&gt;
&lt;td&gt;Still configuration incomplete&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both of the first two rows produce a time, but for different reasons. One comes from the provider's message; the other is a fallback used when information is missing.&lt;/p&gt;

&lt;p&gt;That distinction matters to someone waiting. “Try again at this time” can sound like a promise that service will be available. &lt;strong&gt;A default one-hour delay says when the program proposes to try again; it makes no such promise about the provider.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F49tfi87ijopmrkbmgb8r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F49tfi87ijopmrkbmgb8r.png" alt="Similar-looking retry times can have different foundations" width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. Source: author illustration from the two original classifiers. Exact timestamps and timezone conversion remain in the experimental supplement.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This verifies what the classifier returns. Whether the program arranging the next execution actually waits accordingly needs a further test. The table alone cannot establish that the complete retry problem is solved.&lt;/p&gt;

&lt;h2&gt;
  
  
  A keyword can appear in a denial
&lt;/h2&gt;

&lt;p&gt;We deliberately supplied this sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;No quota exceeded. Socket failed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The candidate still classified it as a quota problem and returned the default one-hour wait.&lt;/p&gt;

&lt;p&gt;This was a constructed counterexample; we have no evidence that a real provider sends that wording. It demonstrates a limit of keyword matching: recognizing a word does not necessarily establish what the sentence means.&lt;/p&gt;

&lt;p&gt;Another input supplied an invalid timezone and also triggered the fallback. An equally precise-looking time can therefore come from successful parsing or a backup rule. Preserving that distinction helps explain why the program waits as it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  A timestamp needs an explanation
&lt;/h2&gt;

&lt;p&gt;A useful next direction is for provider adapters to pass an explicit cause and reset time, reducing reliance on prose. When inference or fallback is needed, its basis should remain available. That is a design direction; this experiment tested only the classifiers.&lt;/p&gt;

&lt;p&gt;Users can help examine the next part of the chain. Save the original message and time, then record whether retries keep accumulating and whether work resumes when expected. Those observations help connect “cause recognized” to “execution actually waited.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When a tool says “wait until this time,” can it also tell you whether the provider supplied that time or the program estimated it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Revisions and experimental boundary&lt;/p&gt;
&lt;h3&gt;
  
  
  Original detailed comparison
&lt;/h3&gt;

&lt;p&gt;The experiment clock was September 16, 2026, 04:50 UTC. The message specified 21:20 Pacific/Auckland, parsed as 09:20 UTC. Without a reset time, the one-hour fallback yielded 05:50 UTC. The PR was open at our source read; that timestamped status does not claim it remains unchanged.&lt;/p&gt;

&lt;p&gt;The newly included entry point is &lt;code&gt;acpx_turn_failed&lt;/code&gt;. The output field &lt;code&gt;parsedResetTime&lt;/code&gt; distinguishes a parsed reset from a fallback. “No classification” means this function returned null; it does not establish that all downstream handling immediately retries.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Before&lt;/th&gt;
&lt;th&gt;Candidate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Quota message through the failed-turn entry point&lt;/td&gt;
&lt;td&gt;No classification&lt;/td&gt;
&lt;td&gt;Provider quota; retry at 09:20 UTC&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quota message without a reset time&lt;/td&gt;
&lt;td&gt;No classification&lt;/td&gt;
&lt;td&gt;Provider quota; retry at 05:50 UTC&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ordinary socket failure&lt;/td&gt;
&lt;td&gt;No classification&lt;/td&gt;
&lt;td&gt;Still no classification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Missing key through the existing adapter-failure entry point&lt;/td&gt;
&lt;td&gt;Configuration incomplete&lt;/td&gt;
&lt;td&gt;Still configuration incomplete&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Base: &lt;code&gt;d0b67bfe71277f5a5ad21c7a7b6cb7af09e2ccca&lt;/code&gt;. Candidate: &lt;code&gt;688c994a95d6c447816c50b14e30745d8ea9e855&lt;/code&gt;. Checked source anchors extract the original classifier block and its original parsing helpers. Nine inputs across two versions produce 18 observations, not 18 live provider calls.&lt;/p&gt;

&lt;p&gt;Additional inputs cover a structured retryNotBefore value yielding 12:00 UTC, an unknown error code, an invalid timezone, and missing-key messages entering through different error codes. The negated sentence is deliberately synthetic.&lt;/p&gt;

&lt;p&gt;We did not run the recovery scheduler, database scan, production account or model. &lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-17-control-intent" rel="noopener noreferrer"&gt;Results and reproduction scripts&lt;/a&gt;. The experiment does not establish the same error path or quota-retry defect in CodeFlowMu.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-17-quota-behind-error-code" rel="noopener noreferrer"&gt;Original article&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-17-quota-behind-error-code" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>No reply. Will retrying launch a second AI agent?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:53:16 +0000</pubDate>
      <link>https://dev.to/joinwell52/no-reply-will-retrying-launch-a-second-ai-agent-3dc9</link>
      <guid>https://dev.to/joinwell52/no-reply-will-retrying-launch-a-second-ai-agent-3dc9</guid>
      <description>&lt;h1&gt;
  
  
  No reply. Will retrying launch a second AI agent?
&lt;/h1&gt;

&lt;p&gt;You tap “Start agent” on your phone. A few seconds pass. Nothing answers.&lt;/p&gt;

&lt;p&gt;Should you tap again? If the first request never arrived, that seems sensible. If the computer already began acting but its reply disappeared, another tap could create a second job.&lt;/p&gt;

&lt;p&gt;Imagine two agents modifying separate copies of the files, leaving you to choose which work to keep. Duplicate work may also consume additional resources. These are possible consequences, not incidents observed in our experiment. They explain why the silent click matters: &lt;strong&gt;no reply is not evidence that nothing happened.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Look up the original operation
&lt;/h2&gt;

&lt;p&gt;The inspiration is &lt;a href="https://github.com/stablyai/orca/pull/21106" rel="noopener noreferrer"&gt;Orca #21106&lt;/a&gt;, by brennanb2025. Orca connects remote devices, such as phones, with computers running AI agents. This change saves launch records to a file so a retry can recover an existing result.&lt;/p&gt;

&lt;p&gt;We were interested because starting an agent may also create a workspace and a session. Repeated words can be ignored; two jobs may keep producing consequences.&lt;/p&gt;

&lt;p&gt;The mechanism gives each launch an operation number. The first request and its retry use the same number, allowing the computer to look up the original record. In code, this is the operation ID. The check also considers the caller and task contents, so an old number cannot silently stand for a different request.&lt;/p&gt;

&lt;p&gt;One distinction carries the design: &lt;strong&gt;“I accepted this operation” and “I completed this launch” need separate records.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What does the record contain when a retry arrives?
&lt;/h2&gt;

&lt;p&gt;We ran the original code that stores these records and decides whether a launch may proceed, using real temporary files. We saved a synthetic successful result, discarded the initial response, then returned with the same operation number.&lt;/p&gt;

&lt;p&gt;“Successful result” here means the result of launching, such as which workspace was returned. It does not mean the agent finished its subsequent assignment. No real agent ran in our probe.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Controlled situation&lt;/th&gt;
&lt;th&gt;What the retry returned&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Launch success saved, reply discarded&lt;/td&gt;
&lt;td&gt;The same saved result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operation accepted, final outcome missing&lt;/td&gt;
&lt;td&gt;Unknown; another execution was refused&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same number, different task contents&lt;/td&gt;
&lt;td&gt;Conflict: this is not the original request&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first row shows why keeping records helps: a lost reply does not erase a saved result. The third shows why numbers cannot be reused indiscriminately. The second is harder to accept—if the operation was recorded, why can it not simply continue?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F58df5235k1ob3mkdkjrr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F58df5235k1ob3mkdkjrr.png" alt="The same number returns: is the record complete?" width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. Source: author illustration from the real-file experiment. Completion here concerns the launch operation; successful outcomes were synthetic.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened after acceptance?
&lt;/h2&gt;

&lt;p&gt;Suppose the computer records “I will execute this,” then loses contact before saving an outcome. That record alone allows at least three explanations: work never began, partly happened, or the launch finished without its final result being written.&lt;/p&gt;

&lt;p&gt;An accepted operation therefore cannot be translated directly into either “not done” or “successful.” Executing again might duplicate work; declaring success might leave someone waiting for a job that does not exist.&lt;/p&gt;

&lt;p&gt;The tested code preserves that uncertainty. Receiving the same number again does not issue another execution permission. We also reopened the store in the same process and could still recover saved results. That is not a test of complete recovery after power loss or a process crash.&lt;/p&gt;

&lt;p&gt;At our read, the change had merged on September 17 UTC, but shipped clients did not yet supply the required outer operation ID. The source also explicitly leaves broader recovery outside this change. Our experiment explains how this record layer handles retries; it cannot establish protection for every tap in the mobile product.&lt;/p&gt;

&lt;h2&gt;
  
  
  An uncertain answer still needs a next step
&lt;/h2&gt;

&lt;p&gt;Preventing a duplicate action matters. The person waiting still needs to decide whether to wait, inspect the work, or seek help.&lt;/p&gt;

&lt;p&gt;Our experiment points toward evidence gathering: the original operation number, any still-running task and any resources already created could help reduce guesswork. Those are directions for future design and validation. We did not implement inspection or recovery.&lt;/p&gt;

&lt;p&gt;Users can contribute a concrete timeline: what the interface said when the reply was missing, what they did next, and whether they later discovered work already running. That helps distinguish an action that never happened from a reply that never arrived.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When a system can only say “unable to confirm,” what evidence should it provide so a person can confidently choose the next step?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Fixed revision, scope and complete observations&lt;/p&gt;
&lt;h3&gt;
  
  
  Original detailed comparison
&lt;/h3&gt;

&lt;p&gt;Admission is the check that permits execution to begin. The eight-request result below concerns this layer only. The higher-level handler joins live work; it was not exercised, so seven refusals cannot be read as seven failed user retries.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Controlled situation&lt;/th&gt;
&lt;th&gt;Observed result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Success was saved but the reply was discarded&lt;/td&gt;
&lt;td&gt;Replayed the same saved result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Store reopened in the same process&lt;/td&gt;
&lt;td&gt;Replayed the saved result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same operation ID, changed task&lt;/td&gt;
&lt;td&gt;Rejected as a conflict&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same paired device, rotated bearer token&lt;/td&gt;
&lt;td&gt;Replayed the existing result&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Operation claimed, final result missing&lt;/td&gt;
&lt;td&gt;Refused another execution; outcome stayed unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eight concurrent admissions for one operation&lt;/td&gt;
&lt;td&gt;One execution admission; seven unknown refusals&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Revision: &lt;code&gt;d60ff2db02bf2b6c6ac4fd39c179c9854409d2a4&lt;/code&gt;. We executed the complete original durableAgentSessionRecordStore plus admitAgentLaunchOperation, fingerprint and registry logic with real temporary files and proper-lockfile.&lt;/p&gt;

&lt;p&gt;The 11 observations also cover caller partitioning between paired devices, replay of recorded failures, missing paired identity for remote callers, and an expired previously unseen operation ID. Bearer rotation retained the same paired-device identity.&lt;/p&gt;

&lt;p&gt;The concurrent test stops at admission; the higher-level live-join handler was not exercised. Reopening occurred within one process. Saved launch outcomes were synthetic: no real workspace, model or mobile call was created.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-17-control-intent" rel="noopener noreferrer"&gt;Results and reproduction scripts&lt;/a&gt;. These observations do not establish a duplicate-launch defect or missing idempotency protection in CodeFlowMu.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-17-retry-without-second-launch" rel="noopener noreferrer"&gt;Original article&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-17-retry-without-second-launch" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>More backups. Less history. How does that happen?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Thu, 17 Sep 2026 03:53:12 +0000</pubDate>
      <link>https://dev.to/joinwell52/more-backups-less-history-how-does-that-happen-33d1</link>
      <guid>https://dev.to/joinwell52/more-backups-less-history-how-does-that-happen-33d1</guid>
      <description>&lt;h1&gt;
  
  
  More backups. Less history. How does that happen?
&lt;/h1&gt;

&lt;p&gt;You open the backup folder and find ten neatly named files. Only when you need yesterday’s settings do you discover that all ten contain today’s version.&lt;/p&gt;

&lt;p&gt;Nothing crashed. Every launch faithfully made a backup. That repetition pushed the useful old version out of the retention window.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A full backup list does not necessarily contain a useful history.&lt;/strong&gt; We tested the original code from an open-source project, asking whether the content from before a change survived.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this small change caught our attention
&lt;/h2&gt;

&lt;p&gt;The source is &lt;a href="https://github.com/yzhao062/anywhere-agents/commit/ce6457a830dead5e05c3a17f685f7ff32575ce20" rel="noopener noreferrer"&gt;commit ce6457a in anywhere-agents&lt;/a&gt;, by Yue Zhao, or yzhao062 on GitHub. The project helps AI coding tools share settings. One Python program merges shared configuration into a local file.&lt;/p&gt;

&lt;p&gt;The upstream investigation found that repeatedly merging unchanged settings created identical backups and displaced useful older contents.&lt;/p&gt;

&lt;p&gt;This matters to us because automated tools often synchronize on every launch. If each synchronization consumes a backup slot, opening more terminals can shorten recoverable history without changing a single setting.&lt;/p&gt;

&lt;p&gt;The patch compares the proposed output with the existing bytes while holding the original file lock. Identical output means no write and no new backup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Change once, then repeat
&lt;/h2&gt;

&lt;p&gt;In a temporary directory, we started with old settings, merged a new value, then performed 27 identical merges. Both versions ran their original file operations and locking code.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Observation&lt;/th&gt;
&lt;th&gt;Before the patch&lt;/th&gt;
&lt;th&gt;After the patch&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Retained backups&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Content from before the change survives&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backups identical to the current file&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identical merges change the target modification time&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first merge really did save the old settings in both versions. Before the patch, later merges kept saving the new settings again until the first backup fell outside the ten-file window.&lt;/p&gt;

&lt;p&gt;After the patch, unchanged merges stopped consuming that window. Fewer backup files preserved more useful history.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fadesvxn6b30s7rcflzv0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fadesvxn6b30s7rcflzv0.png" alt="Repeated copies versus recoverable history" width="799" height="433"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. One real change followed by 27 identical merges. Source: author illustration from our local original-code experiment. Each block is a retained backup, not an upstream consumer.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;We also ran eight concurrent processes after the initial change. The old version retained nine backups; the patched version retained only the original old content. That supports the lock-protected check under these conditions, not a claim about every filesystem failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does “unchanged” mean?
&lt;/h2&gt;

&lt;p&gt;Next, we deliberately reformatted the target before each merge. Keys and values stayed the same; whitespace changed.&lt;/p&gt;

&lt;p&gt;After 12 such rewrites, both versions retained ten backups and had evicted the original old settings. This does not contradict the patch: different formatting means different bytes. The implementation skips byte-identical writes; it does not promise ten semantically distinct revisions.&lt;/p&gt;

&lt;p&gt;This was a constructed counterexample, not a reported production incident. It exposes a useful design question: should history preserve every write, or every meaningful change?&lt;/p&gt;

&lt;p&gt;Ignoring formatting is not automatically the right answer. Comments, ordering and presentation can themselves be worth restoring. A comparison that is too coarse can discard changes that matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try recovering, not just counting
&lt;/h2&gt;

&lt;p&gt;A practical check is small: change a harmless setting, restart or synchronize repeatedly, then look for the pre-change version. The target is specific: can you recover what existed before that change?&lt;/p&gt;

&lt;p&gt;If you have found plenty of backups but not the version you needed, record when the setting changed and how often the tool restarted or synchronized afterward. That helps locate when repeated copies displaced useful history.&lt;/p&gt;

&lt;p&gt;For developers, the formatting counterexample suggests another experiment: let two tools alternate writes to the same configuration, then compare retention by write count, distinct content and elapsed time. Each policy may lose something different. Formatting or comments can themselves be worth restoring, so deduplication should not be assumed to be the answer.&lt;/p&gt;

&lt;p&gt;The tested improvement is clear: unchanged bytes no longer create backups, preserving the original revision in that scenario. It leaves one larger question: &lt;strong&gt;does a backup system promise the last few saves, or the changes a person needs to return to?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Versions, scope and reproduction&lt;/p&gt;

&lt;p&gt;The upstream investigation reported 27 consumers and 198 byte-identical backups. Those are &lt;strong&gt;upstream observations, not the scale of our experiment&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We ran the complete original Python merge module through its main entry point, using real Windows temporary files and its original locks. Six scenarios ran against both versions, producing 12 observations: identical merges, deliberate formatting rewrites, another real change, first installation, empty-merge refusal, and eight concurrent identical merges.&lt;/p&gt;

&lt;p&gt;Base: &lt;code&gt;91c64c153f2034bf4aa88121cbe546cbf6a22aeb&lt;/code&gt;. Patched revision: &lt;code&gt;ce6457a830dead5e05c3a17f685f7ff32575ce20&lt;/code&gt;. Except for first installation, scenarios begin with an initial merge. The 27 repeats are not 27 real consumers. We used synthetic settings and did not test power loss or disk corruption.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-17-control-intent" rel="noopener noreferrer"&gt;Results and reproduction scripts&lt;/a&gt;. This experiment does not establish the same problem in CodeFlowMu.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-17-backup-without-history" rel="noopener noreferrer"&gt;Original article&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-17-backup-without-history" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>MIIT Document No. 209 Explained: Agent Software and the Transformation of the Software Industry</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Wed, 16 Sep 2026 05:19:05 +0000</pubDate>
      <link>https://dev.to/joinwell52/miit-document-no-209-explained-agent-software-and-the-transformation-of-the-software-industry-1i9a</link>
      <guid>https://dev.to/joinwell52/miit-document-no-209-explained-agent-software-and-the-transformation-of-the-software-industry-1i9a</guid>
      <description>&lt;p&gt;On September 11, 2026, China's Ministry of Industry and Information Technology (MIIT) published the &lt;a href="https://www.miit.gov.cn/jgsj/xxjsfzs/wjfb/art/2026/art_91a26793271e4c77aca49ecde4472999.html" rel="noopener noreferrer"&gt;Implementation Plan for the “AI Plus Software” Special Action&lt;/a&gt;. Numbered Gong Xin Bu Xin Fa [2026] No. 209 and dated September 2, the document was issued to industry and information technology authorities in all provinces, autonomous regions, municipalities and the Xinjiang Production and Construction Corps, asking them to implement it in light of local conditions.&lt;/p&gt;

&lt;p&gt;The title might suggest a simple call for software companies to use more AI. But the eight-page plan and its 19 tasks address much more than AI-assisted coding.&lt;/p&gt;

&lt;p&gt;The document tackles three fundamental questions for the software industry.&lt;/p&gt;

&lt;p&gt;First, how will software be produced?&lt;/p&gt;

&lt;p&gt;Second, what will software itself become?&lt;/p&gt;

&lt;p&gt;Third, when software starts acting continuously as an agent, who will provide the execution frameworks, skills, markets, standards, security and boundaries of responsibility?&lt;/p&gt;

&lt;p&gt;Document No. 209 therefore matters for more than adding another “AI Plus” policy. China is beginning to organize &lt;strong&gt;intelligent programming, agent software and intelligent services&lt;/strong&gt; as new growth drivers for the software industry. Agents, previously often grouped together with large models, chatbots and automation tools, now sit within a complete software value chain: development tools, execution frameworks, software products, skill packages, application marketplaces, delivery services, and standards for identity, interfaces, behavior and security.&lt;/p&gt;

&lt;p&gt;Agents are moving from a model feature toward a distinct form of software.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Understanding Document No. 209 in Six Terms&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Intelligent programming&lt;/strong&gt; — Transform software production&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intelligent companions&lt;/strong&gt; — Upgrade existing software&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent software&lt;/strong&gt; — Develop new product forms&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills&lt;/strong&gt; — Package specialist capabilities&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intelligent services&lt;/strong&gt; — Deliver value continuously&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verifiable behavior&lt;/strong&gt; — Inspect execution and outcomes&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  I. Beyond Document No. 414: Adding Products and Infrastructure
&lt;/h2&gt;

&lt;p&gt;To understand Document No. 209, it helps to place it in the sequence of policies released in 2026, without treating the documents as interchangeable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.miit.gov.cn/jgsj/kjs/wjfb/art/2026/art_331b6f1d28ad410aa9df711acf42dbfe.html" rel="noopener noreferrer"&gt;Document No. 414&lt;/a&gt;, published on August 31, organizes a resource pool, service teams, real use cases, delivery arrangements and ongoing evaluation around AI application service providers. Its question is: &lt;strong&gt;who brings AI into enterprises, and who takes responsibility for consulting, implementation, operations and security governance?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://www.miit.gov.cn/jgsj/qyj/wjfb/art/2026/art_86c400b4473849818629663a94a6d44b.html" rel="noopener noreferrer"&gt;Artificial Intelligence SME Entrepreneurship Support Plan (2026–2028)&lt;/a&gt;, published on September 4, brings AI-native firms, agent developers, open-source contributors, “one-person companies” and highly capable individual entrepreneurs into the scope of startup support. Its question is: &lt;strong&gt;who starts new AI businesses, and how will the computing, data, use cases, capital and services they need become available?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We analyzed those documents in &lt;a href="https://joinwell52-ai.github.io/joinwell52/en/industry/2026-09-04-miit-414-ai-application-delivery" rel="noopener noreferrer"&gt;MIIT Document No. 414 Explained: From Model Supply to Application Delivery&lt;/a&gt; and &lt;a href="https://joinwell52-ai.github.io/joinwell52/en/industry/2026-09-08-miit-ai-sme-support-plan" rel="noopener noreferrer"&gt;MIIT's AI SME Entrepreneurship Support Plan Explained: An AI Startup Ecosystem Is Taking Shape&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Document No. 209 addresses a third question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What software will these entrepreneurs and service providers actually produce, run and deliver?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It takes policy further inside the software industry. From production tools, technical foundations, product forms and service models to open source, computing, data, standards, talent and security, it offers a more complete plan for building the industry.&lt;/p&gt;

&lt;p&gt;The three documents have distinct roles:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Document&lt;/th&gt;
&lt;th&gt;Main subject&lt;/th&gt;
&lt;th&gt;Central question&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Document No. 414&lt;/td&gt;
&lt;td&gt;AI application service providers&lt;/td&gt;
&lt;td&gt;Who handles consulting, integration, delivery, operations and governance?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Entrepreneurship Support Plan&lt;/td&gt;
&lt;td&gt;AI SMEs and entrepreneurs&lt;/td&gt;
&lt;td&gt;Who starts businesses, and how do they obtain computing, data, use cases and capital?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Document No. 209&lt;/td&gt;
&lt;td&gt;Software and information technology services&lt;/td&gt;
&lt;td&gt;How is next-generation software produced, how does it run, and what industrial foundations does it need?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If Document No. 414 organizes the delivery teams and the entrepreneurship plan cultivates new companies, Document No. 209 defines the software industry system they will help build.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmzon4ycchfbmajksuns.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcmzon4ycchfbmajksuns.png" alt="The three policies address delivery providers, startup support and the software industry: distinct but complementary roles." width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. The three policies address delivery providers, startup support and the software industry: distinct but complementary roles. Source: Relevant MIIT policy documents; compiled by the author.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  II. Why “AI Plus Software” Means More Than Adding AI to Software
&lt;/h2&gt;

&lt;p&gt;The opening of the document identifies three things AI is changing: &lt;strong&gt;development methods, product forms and service models&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Each represents a different transformation.&lt;/p&gt;

&lt;h3&gt;
  
  
  The first transformation: how software is produced
&lt;/h3&gt;

&lt;p&gt;AI will no longer be merely a completion tool used while a programmer writes a piece of code. It will participate throughout a project: understanding requirements, designing architecture, coding, testing, checking security, deployment, operations and ongoing changes.&lt;/p&gt;

&lt;p&gt;The plan therefore calls for agent-driven intelligent programming tools, stronger project-level autonomous development across the full process, and deeper integration with code hosting platforms, cloud services and open-source communities. The relevant unit is no longer how quickly a function can be written, but whether an entire software project can be completed more efficiently and reliably.&lt;/p&gt;

&lt;p&gt;This also explains the call for intelligent upgrading within software companies. Genuine upgrading means more than installing a chat plugin for every programmer. It requires redesigning development processes, quality control, security checks and collaboration, with results assessed through code quality, development effectiveness and other practical outcomes.&lt;/p&gt;

&lt;h3&gt;
  
  
  The second transformation: what a software product is
&lt;/h3&gt;

&lt;p&gt;Traditional software mainly waits for users to click menus, fill in forms and issue commands. Agent software can receive goals, understand its environment, select tools, execute successive steps and adjust subsequent actions based on results.&lt;/p&gt;

&lt;p&gt;The document therefore treats agent software as a new product form in its own right, rather than discussing only AI features added to existing software. It covers capabilities embedded in operating systems, databases, industrial software and devices, as well as general-purpose agents, specialist industry agents and industrial agents.&lt;/p&gt;

&lt;p&gt;This raises the agent from an interface feature to a software product category.&lt;/p&gt;

&lt;h3&gt;
  
  
  The third transformation: how software services work
&lt;/h3&gt;

&lt;p&gt;Traditional software services tend to revolve around licenses, implementation projects and maintenance. Intelligent software services may instead become a continuously operating service chain: models keep changing, knowledge is updated, agents keep executing, permissions and costs evolve, and results need ongoing verification.&lt;/p&gt;

&lt;p&gt;The plan calls for an end-to-end intelligent application service chain, encourages software firms to become AI application service providers, and promotes a shift from one-off delivery to ongoing service and value delivery.&lt;/p&gt;

&lt;p&gt;Together, these three changes express the full meaning of “AI Plus Software”: AI is reorganizing software production, products and business models.&lt;/p&gt;

&lt;h2&gt;
  
  
  III. 2028 and 2030: From Adoption at Scale to Industry Upgrading
&lt;/h2&gt;

&lt;p&gt;Document No. 209 sets two stages of objectives.&lt;/p&gt;

&lt;p&gt;By 2028:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Intelligent programming tools and intelligent development platforms should reach 20,000 software enterprises above the designated size.&lt;/li&gt;
&lt;li&gt;A total of 100 intelligent upgrading projects should be implemented in software companies.&lt;/li&gt;
&lt;li&gt;Priority industries should develop 100 benchmark agent software applications.&lt;/li&gt;
&lt;li&gt;At least 15 hardware–software adaptation centers should be established.&lt;/li&gt;
&lt;li&gt;At least five high-quality open-source projects should be incubated.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By 2030:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Key software should have undergone comprehensive intelligent upgrading.&lt;/li&gt;
&lt;li&gt;Intelligent programming, agent software and intelligent services should become new industry growth drivers.&lt;/li&gt;
&lt;li&gt;A number of internationally influential open-source communities and industry clusters should be established.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These figures should not simply be added together: each serves a different policy purpose.&lt;/p&gt;

&lt;p&gt;Reaching 20,000 companies measures diffusion of production tools. The 100 upgrading projects are meant to produce reusable examples of company transformation. The 100 benchmark agent applications test whether new products can enter real industries. Adaptation centers support hardware–software coordination and industrial validation. Open-source projects contribute to a shared technical ecosystem.&lt;/p&gt;

&lt;p&gt;The plan is therefore building adoption at scale, demonstration projects, validation facilities and open-source foundations at the same time.&lt;/p&gt;

&lt;p&gt;Read together, the targets emphasize adoption and validation at scale by 2028, followed by upgrading key software, growing new business forms, and developing internationally influential open-source communities and industry clusters by 2030. They describe a policy path from broader adoption toward industry upgrading, rather than suggesting that the industry's structure will be settled by then.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh51nw0mt87w974vpsf1e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh51nw0mt87w974vpsf1e.png" alt="The 2028 targets cover firms, projects, applications, compatibility centers and open source; 2030 points toward industry transformation." width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 2. The 2028 targets cover firms, projects, applications, compatibility centers and open source; 2030 points toward industry transformation. Source: Relevant MIIT policy documents; compiled by the author.&lt;/em&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Agents Are a New Growth Area, but Document No. 209 Is Not an Agent-Only Policy
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://www.miit.gov.cn/jgsj/xxjsfzs/wjfb/art/2026/art_91a26793271e4c77aca49ecde4472999.html" rel="noopener noreferrer"&gt;implementation plan&lt;/a&gt; also addresses intelligent upgrading of foundational and industrial software itself. Operating systems and databases are to improve agent scheduling, performance tuning, operations and security, with chip and server manufacturers collaborating on hardware–software compatibility and adaptation centers. Industrial software such as CAD, CAE and EDA is to apply AI to drawing, design and simulation, while industrial internet platforms strengthen data analysis, production scheduling and multi-agent decision-making.&lt;/p&gt;

&lt;p&gt;Document No. 209 should therefore be read as combining upgrades to existing software with the development of new software forms. Foundational software, industrial software and development methods all need intelligent upgrading. Agents are an important emerging form within that broader software-industry transformation.&lt;/p&gt;

&lt;h2&gt;
  
  
  IV. The Most Consequential New Concept Is Agent Software
&lt;/h2&gt;

&lt;p&gt;Intelligent programming matters: it directly affects software production efficiency and can scale relatively quickly. From an industry-structure perspective, however, the most notable part of Document No. 209 is its fourth section, on &lt;strong&gt;cultivating new forms of agent software and accelerating their continuous evolution&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It introduces at least four layers.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Agent execution frameworks
&lt;/h3&gt;

&lt;p&gt;The plan calls for engineering research into agent execution frameworks, improving reliable execution of complex tasks and collaboration among multiple agents.&lt;/p&gt;

&lt;p&gt;The emphasis on execution frameworks shows attention to engineering foundations beyond models. &lt;strong&gt;From an enterprise engineering perspective,&lt;/strong&gt; a large model can generate an answer, but it does not automatically solve persistent state, tool use, permission constraints, failure recovery, concurrency conflicts, handoffs and acceptance of results over a long-running task. These are engineering questions the author derives from reliable execution and multi-agent collaboration, not a feature checklist prescribed item by item in the document.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Agent software development platforms
&lt;/h3&gt;

&lt;p&gt;The document proposes agent software development platforms and better toolchains for development, testing, deployment and operations.&lt;/p&gt;

&lt;p&gt;Agent development cannot remain a collection of prompts, demonstration pages and temporary scripts. Like mature software, it needs design, testing, releases, monitoring, upgrades and maintenance, supported by repeatable engineering processes.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Agent software products
&lt;/h3&gt;

&lt;p&gt;The plan separately addresses general-purpose agents, device agents, specialist agents for vertical domains and industrial agents. Each category brings different reliability and accountability requirements.&lt;/p&gt;

&lt;p&gt;An office information-processing agent can work under human review. Device agents need lightweight inference and privacy protection. Vertical-domain agents need knowledge graphs and industry data. Industrial agents may enter core production processes and interact with industrial control systems and critical business systems. They therefore require greater safety, reliability and verifiable behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Agent software application markets
&lt;/h3&gt;

&lt;p&gt;The document also proposes agent software application stores and skill-package repositories, with rules for listing reviews and operational management.&lt;/p&gt;

&lt;p&gt;This is a consequential step. Agents, skills and components must be discoverable, installable, composable, updatable and tradable if an agent software ecosystem is to develop, instead of every project starting from scratch.&lt;/p&gt;

&lt;p&gt;Together, these four layers form a complete chain:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Execution frameworks provide the foundation; development platforms produce software; agent products enter real use cases; marketplaces distribute them; service providers handle delivery and operations.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz5rq0gqd0lc8izvl7byr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz5rq0gqd0lc8izvl7byr.png" alt="Execution frameworks, development platforms, agent products and distribution form an industry, with delivery and governance spanning the chain." width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 3. Execution frameworks, development platforms, agent products and distribution form an industry, with delivery and governance spanning the chain. Source: Relevant MIIT policy documents; compiled by the author.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  V. Skills Are Formally Included in Agent Software Markets: More Than a Prompt Marketplace
&lt;/h2&gt;

&lt;p&gt;Document No. 209 explicitly proposes repositories for skill packages, or Skills, and encourages developers to build high-quality specialist skills, knowledge bases and functional components for agents.&lt;/p&gt;

&lt;p&gt;A skill may appear to be a packaged capability: organizing contracts, analyzing equipment faults, generating quotations or operating a business system. Once it enters an enterprise or an application marketplace, however, it cannot remain a few paragraphs of prompts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;From an enterprise engineering perspective,&lt;/strong&gt; a commercially usable skill package raises the following questions. They are the author's deductions from the proposed skill repositories, listing reviews and operational management, rather than requirements individually prescribed by the document:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Who publishes it, and can that identity be trusted?&lt;/li&gt;
&lt;li&gt;Which agents and execution environments can use it?&lt;/li&gt;
&lt;li&gt;Which models, data and tools does it depend on?&lt;/li&gt;
&lt;li&gt;What permissions can it receive?&lt;/li&gt;
&lt;li&gt;How are versions upgraded and rolled back?&lt;/li&gt;
&lt;li&gt;What external effects does execution produce?&lt;/li&gt;
&lt;li&gt;Can it stop, recover or undo work after an error?&lt;/li&gt;
&lt;li&gt;Who approves its listing, and who handles ongoing operations?&lt;/li&gt;
&lt;li&gt;Are its intellectual property, data sources and security risks clear?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A mature Skills market will therefore resemble a software-package ecosystem for agents. Capabilities are modular, dependencies and permissions are explicit, effects can be inspected, and versions and responsibility can be traced.&lt;/p&gt;

&lt;p&gt;This also creates new specializations. Some firms will build general-purpose skills; others will provide industry knowledge packages, maintain internal enterprise repositories, review security and compatibility, or operate agent marketplaces for particular industries.&lt;/p&gt;

&lt;p&gt;Many small AI firms may not need to train their own foundation model. They can instead package knowledge, tools and workflows around a valuable business process. This connects directly with the entrepreneurship plan's emphasis on agent developers, open-source contributors and small specialist businesses.&lt;/p&gt;

&lt;h2&gt;
  
  
  VI. Why Verifiable Behavior Could Be the Dividing Line for Industrial Agents
&lt;/h2&gt;

&lt;p&gt;In its provisions for industrial agents, the plan calls for safety, reliability and verifiable behavior.&lt;/p&gt;

&lt;p&gt;The final requirement is easy to overlook, yet it may determine whether agents can enter actual production and business processes.&lt;/p&gt;

&lt;p&gt;When a chatbot says something wrong, a user can often ignore or correct it. If an industrial agent changes production schedules, operates equipment, adjusts parameters, writes to a business system or triggers procurement, an error can become a real loss. A chat history alone cannot prove what happened or whether the result stayed within authorization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;From an enterprise engineering perspective,&lt;/strong&gt; the author breaks verifiable behavior into the following questions. This is an analytical framework, not an acceptance standard itemized in the document:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What task did the agent receive?&lt;/li&gt;
&lt;li&gt;Who authorized execution?&lt;/li&gt;
&lt;li&gt;Which tools did it call?&lt;/li&gt;
&lt;li&gt;What effects did those tools actually produce?&lt;/li&gt;
&lt;li&gt;Did the results meet the business objective?&lt;/li&gt;
&lt;li&gt;Who has authority to accept results and resolve disputes?&lt;/li&gt;
&lt;li&gt;Did recovery repeat external effects whose status was unknown?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A model's account of its own behavior cannot answer all these questions. Model output can contribute evidence, but does not automatically establish facts. Software can inspect files, states, hashes and tool results, but cannot assume the business owner's authority to judge whether the result is acceptable.&lt;/p&gt;

&lt;p&gt;Industrial agents therefore need more than higher model accuracy. They need task agreements, identity and authorization, execution records, evidence of effects, independent verification and recovery mechanisms. Providers that turn these engineering capabilities into affordable, integrable infrastructure can help agents move from convincing demonstrations to systems that enterprises can actually trust to use.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fltveyk3oj0jezxeftvwr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fltveyk3oj0jezxeftvwr.png" alt="Verifiable behavior separates task, authority, execution, effects and acceptance; recovery also requires state and side-effect checks." width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 4. Verifiable behavior separates task, authority, execution, effects and acceptance; recovery also requires state and side-effect checks. Source: Relevant MIIT policy documents; compiled by the author.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  VII. Four Main Directions for Agent Standards
&lt;/h2&gt;

&lt;p&gt;Document No. 209 does not announce a completed set of national agent standards. It commissions and advances subsequent standards development. That distinction matters.&lt;/p&gt;

&lt;p&gt;The document proposes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Graded maturity assessment standards for intelligent programming capabilities.&lt;/li&gt;
&lt;li&gt;Agent interface standards supporting hardware–software coordination.&lt;/li&gt;
&lt;li&gt;Research into intelligent service standards, including Model as a Service and Agent as a Service.&lt;/li&gt;
&lt;li&gt;Research into agent identity, trusted interconnection, data security and behavior control.&lt;/li&gt;
&lt;li&gt;Security management requirements covering agent development, deployment and application.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These provisions suggest four main directions.&lt;/p&gt;

&lt;p&gt;The first is &lt;strong&gt;capability standards&lt;/strong&gt;: how broad a task can a coding tool or development platform complete, and how should maturity be classified? Vendor claims are insufficient.&lt;/p&gt;

&lt;p&gt;The second is &lt;strong&gt;interface standards&lt;/strong&gt;: how should different models, devices, hardware, software systems and agents connect, without locking every capability into a single platform?&lt;/p&gt;

&lt;p&gt;The third is &lt;strong&gt;service standards&lt;/strong&gt;: how should MaaS and AaaS define scope, quality, measurement, responsibility and requirements for ongoing operations?&lt;/p&gt;

&lt;p&gt;The fourth is &lt;strong&gt;governance standards&lt;/strong&gt;: under what identity does an agent enter a system, how does it obtain permission, how can it connect trustworthily, and how are data and behavior protected, recorded and controlled?&lt;/p&gt;

&lt;p&gt;The direction is now clear, but standards are still taking shape. The most useful work companies can do is accumulate real execution evidence, interface experience, security mechanisms and failure cases that can support standards development and pilot validation. Claiming compliance with unpublished national standards is premature.&lt;/p&gt;

&lt;h2&gt;
  
  
  VIII. Open Source as a Mechanism for Discovery, Validation and Diffusion
&lt;/h2&gt;

&lt;p&gt;The plan makes collaborative open-source innovation part of the foundation for intelligent software development. It calls for nationwide open-source foundations and AI communities to encourage firms to release work on intelligent programming, intelligent software products and agents. It supports first releases of key projects through communities, greater use of domestically developed open-source licenses and compliance guidance, and mechanisms for selecting and cultivating high-quality projects.&lt;/p&gt;

&lt;p&gt;One particularly notable proposal is to explore open-source contributions as a reference for identifying promising software SMEs.&lt;/p&gt;

&lt;p&gt;This echoes the entrepreneurship plan's use of public indicators such as project stars to identify high-growth potential. Document No. 209 goes further: open-source activity should help identify firms, while strong projects should also enter actual manufacturing, finance and other industry applications.&lt;/p&gt;

&lt;p&gt;That could change how agent infrastructure projects develop.&lt;/p&gt;

&lt;p&gt;A project once viewed primarily as a technical community contribution may become an entry point for policy support, industrial validation, service-provider partnerships and deployment. But visible code must become usable capability: clear installation instructions, stable interfaces, traceable versions, compliant licensing, reproducible tests and evidence of real use.&lt;/p&gt;

&lt;p&gt;Stars increase visibility, but do not establish reliability. Publishing code alone does not establish industrial capability. The policy seeks projects that can participate in real software production and industry applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  IX. Funding, Projects and Use Cases: Direction Is Clear, Applications Are Still to Come
&lt;/h2&gt;

&lt;p&gt;In its implementation provisions, Document No. 209 proposes mechanisms such as challenge-based project selection, coordination of funding channels and first-version software product recognition policies, support for eligible projects, and dedicated fiscal funding for key technology research, infrastructure and practical applications.&lt;/p&gt;

&lt;p&gt;This creates the prospect of projects, pilots, standards work, adaptation centers, benchmark applications and local support measures.&lt;/p&gt;

&lt;p&gt;The plan's computing-service provisions already call for &lt;strong&gt;compute vouchers and other broadly accessible support policies&lt;/strong&gt; to reduce software companies' computing costs. &lt;a href="https://www.miit.gov.cn/jgsj/xxjsfzs/gzdt/art/2026/art_45df9463e105446c996568f299e07bca.html" rel="noopener noreferrer"&gt;MIIT's official explanation&lt;/a&gt; goes further: local authorities are encouraged to use compute vouchers and related policies to cover a proportion of the costs of purchasing domestically developed programming tools and using programming services from domestically developed large models. Support is therefore not confined to major projects; the policy also points toward assistance with companies' purchases of intelligent programming tools and model services.&lt;/p&gt;

&lt;p&gt;Policy direction must nevertheless be distinguished from an entitlement already granted.&lt;/p&gt;

&lt;p&gt;Neither the plan nor this explanation sets a nationwide subsidy percentage, claim requirements or implementation timetable. They also do not publish unified project application dates, company eligibility rules, scoring criteria, funding pools or award amounts per project. Words such as support, encourage and explore do not mean every relevant company automatically receives a subsidy. Eligibility, covered expenses and payment timing remain subject to local policies and application notices.&lt;/p&gt;

&lt;p&gt;Companies should prepare five kinds of material:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reproducible product capabilities and clear technical limits.&lt;/li&gt;
&lt;li&gt;Real customer use cases and evidence of outcomes.&lt;/li&gt;
&lt;li&gt;Complete development, testing, deployment and operating processes.&lt;/li&gt;
&lt;li&gt;Explanations of data, permissions, security and intellectual property.&lt;/li&gt;
&lt;li&gt;Demonstration cases that can be replicated, extended and inspected by third parties.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When applications open, the strongest teams will be those that can demonstrate solutions to the engineering problems identified by the plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  X. The New Software Value Chain Taking Shape
&lt;/h2&gt;

&lt;p&gt;Document No. 209 covers far more participants than foundation-model companies. Its provisions point toward a new value chain:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg7sr2vo7qgcww58ketvb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg7sr2vo7qgcww58ketvb.png" alt="Eight specializations in the agent software value chain and their core capabilities." width="800" height="658"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 5. Eight specializations in the agent software value chain and their core capabilities. Source: Author's analysis of MIIT Document No. 209.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Not every company needs to build an all-encompassing agent platform. Specialist strengths in execution, skills, evaluation, security, integration or a particular industry can support independent businesses.&lt;/p&gt;

&lt;p&gt;This is especially relevant to SMEs. They may struggle to compete with large firms in general-purpose foundation models, but can build advantages in data, processes, knowledge and delivery around a narrow, demanding use case. An agent that reliably handles procurement inquiries, equipment-fault analysis, research-document organization, quality tracking or export-customer management may have more commercial value than a feature-rich general assistant that cannot enter real operations.&lt;/p&gt;

&lt;h2&gt;
  
  
  XI. Competition Will Turn to What AI Actually Achieves
&lt;/h2&gt;

&lt;p&gt;The provisions on intelligent upgrading in software companies identify code quality and development effectiveness as important measures of intelligent programming adoption. This directly affects how companies should approach AI.&lt;/p&gt;

&lt;p&gt;Connecting a foundation model will increasingly be insufficient as evidence of progress. Companies will need to answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How much shorter is the delivery cycle?&lt;/li&gt;
&lt;li&gt;Have defect rates fallen?&lt;/li&gt;
&lt;li&gt;Are security issues detected earlier?&lt;/li&gt;
&lt;li&gt;Have test coverage and maintenance efficiency improved?&lt;/li&gt;
&lt;li&gt;Can complex projects progress through completion?&lt;/li&gt;
&lt;li&gt;Do workflows remain stable after models or tools change?&lt;/li&gt;
&lt;li&gt;How are jobs redesigned, rather than simply eliminated?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  An Often-Overlooked Policy Signal: Stabilize Jobs, Expand Opportunities and Support Transitions
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.miit.gov.cn/jgsj/xxjsfzs/gzdt/art/2026/art_45df9463e105446c996568f299e07bca.html" rel="noopener noreferrer"&gt;MIIT's official explanation&lt;/a&gt; addresses employment directly through three priorities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Stabilize jobs:&lt;/strong&gt; protect existing employment in software and prevent sudden, large-scale unemployment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expand opportunities:&lt;/strong&gt; use agent software and intelligent services to create employment opportunities, including new roles in code review and AI security.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support transitions:&lt;/strong&gt; strengthen job training and help software workers move into higher-value roles.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Document No. 209 consequently addresses technology and products alongside a practical question: &lt;strong&gt;what happens to people as AI transforms the software industry?&lt;/strong&gt; It does not simply encourage replacing programmers with AI. As production is reorganized, the policy brings job redesign, skills training, new roles and assessments of employment effects into the same framework.&lt;/p&gt;

&lt;p&gt;For companies implementing the changes, a transformation plan should address more than development efficiency and tool costs. It should explain which roles will change, how people will be trained and redeployed, who will take responsibility for code review and AI security, and how employment risks will be identified and managed. A reduction in manual steps alone is not a complete measure of successful transformation.&lt;/p&gt;

&lt;p&gt;Mature transformation requires a renewed division of work: machines perform executable tasks at scale, while people retain responsibility for goals, boundaries, professional judgment, exceptions and final accountability.&lt;/p&gt;

&lt;h2&gt;
  
  
  XII. How Different Software Firms Can Transform
&lt;/h2&gt;

&lt;p&gt;Document No. 209 does not give every software company the same task of connecting a large language model. Firms have different customers, products, technologies and delivery models, so their starting points should differ. The right sequence is to &lt;strong&gt;identify what is worth preserving in the existing business, determine which work can be delegated to agents, and then decide which capabilities are missing.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Established software vendors: retain the business foundation and add task-performing capabilities
&lt;/h3&gt;

&lt;p&gt;For companies with ERP, CRM, finance, collaboration or industry software, the most valuable assets often include business objects, permission systems, historical data and customer trust. A natural-language interface is no reason to discard those assets.&lt;/p&gt;

&lt;p&gt;A more practical path starts with bounded tasks: detecting unusual orders, preparing customer follow-up recommendations, explaining business data, or drafting transactions for review. Agents should access existing systems through controlled interfaces, rather than bypassing permissions to modify databases directly.&lt;/p&gt;

&lt;p&gt;The transformation is not simply replacing menus with a chat box. It is moving from users performing every step to users setting a goal and the system completing a task within authorized limits. Existing approval and accountability mechanisms should remain for consequential actions involving payments, contracts, inventory or commitments to customers.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Project providers and integrators: move toward reusable products and ongoing operations
&lt;/h3&gt;

&lt;p&gt;These firms understand customer environments, business processes and system integration. Their advantage is knowing which problems matter, not necessarily training another foundation model.&lt;/p&gt;

&lt;p&gt;They should look for recurring needs across existing projects and turn reusable knowledge, interfaces, skills and acceptance methods into product packages. They must distinguish standard delivery from customization, while including model updates, knowledge maintenance, incident handling, security checks and outcome reviews in ongoing service.&lt;/p&gt;

&lt;p&gt;The business model also needs clear boundaries: what implementation fees, subscriptions or operating fees cover; which data the customer supplies; who handles exceptions; and how service is handed over when a contract ends. Outcome-based pricing may be possible where outcomes can be measured and responsibility attributed. Promising it before defining acceptance is premature.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Tool and platform companies: move from isolated generation to verifiable engineering
&lt;/h3&gt;

&lt;p&gt;For coding tools, development platforms, cloud services and execution-framework providers, competition shifts from generating a piece of code to advancing a project reliably and producing maintainable results.&lt;/p&gt;

&lt;p&gt;Within their product boundaries, firms need to build or integrate version control, testing, security checks, deployment, persistent state, tool permissions, cost control and recovery. No platform must build every component itself, but dependencies, integration responsibilities and failure ownership must be clear.&lt;/p&gt;

&lt;p&gt;Evaluation should return to real projects: are delivery time, defect rework, maintenance costs and human intervention improving? A demonstration or model benchmark can establish a limited capability, but cannot replace project-level delivery evidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Specialist software SMEs: deepen expertise in a valuable, bounded process
&lt;/h3&gt;

&lt;p&gt;Smaller firms need not treat a comprehensive general-purpose agent platform as the inevitable destination. More practical opportunities often lie in narrow but demanding work such as procurement inquiries, equipment-fault analysis, quality tracking or research-document processing.&lt;/p&gt;

&lt;p&gt;They can use existing models and platforms to package domain knowledge, tool interfaces and procedures into specialist agents or Skills. What can be sold, however, is more than a prompt: it includes a defined scope, data requirements, permission boundaries, test examples, an upgrade method and a service commitment.&lt;/p&gt;

&lt;p&gt;Productization depends on whether the capability can be reused across customers and whether deployment and maintenance costs remain manageable. If each customer requires substantial redevelopment, it must still be accounted for as project work. Using an agent does not automatically turn it into a software product.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F86vqmpecfh05kyogqmq7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F86vqmpecfh05kyogqmq7.png" alt="Firms should choose transformation opportunities from their existing strengths and business boundaries, then address missing capabilities." width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 6. Firms should choose transformation opportunities from their existing strengths and business boundaries, then address missing capabilities. Source: Author's enterprise transformation analysis informed by Document No. 209.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is a business analysis informed by the policy, not an official classification or a prescribed transformation route. A company can occupy several positions. The point is to identify a boundary within which it can create value, rather than relabel every existing product with policy terminology.&lt;/p&gt;

&lt;h2&gt;
  
  
  XIII. Making Transformation Work: From One Use Case to Continuous Delivery
&lt;/h2&gt;

&lt;p&gt;After choosing a direction, a company still needs to answer a practical question: what should it do first, and what evidence justifies further investment?&lt;/p&gt;

&lt;p&gt;Transformation need not begin with replacing every system, buying models in bulk or establishing a large AI department. A more measured approach starts with a verifiable business task and proceeds through choosing a use case, piloting it, establishing a service and expanding reuse. The following is implementation advice, not a process or timetable prescribed by Document No. 209.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1. Select the use case and establish the business baseline
&lt;/h3&gt;

&lt;p&gt;Prioritize work with genuine, recurring demand, relatively stable inputs and inspectable results. Confirm the rights to use the data, permission to access the necessary systems, and the ability to hand control back to a person when something goes wrong.&lt;/p&gt;

&lt;p&gt;For example, start with organizing procurement inquiries and flagging exceptions, then consider drafting transactions for approval. Do not begin with automatic payments whose authorization boundaries are undefined.&lt;/p&gt;

&lt;p&gt;Before the pilot, record processing time, volume, errors, rework, labor costs and the acceptance method, and assign a business owner. Otherwise it will be difficult to tell whether any improvement came from AI, a process redesign or additional human effort.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2. Pilot within controlled boundaries and record failures as well as successes
&lt;/h3&gt;

&lt;p&gt;Agree on permitted data, tools, permission limits and human-intervention conditions before introducing the agent into the process. Begin with read-only access, recommendations or actions after human review, then expand the scope as evidence warrants.&lt;/p&gt;

&lt;p&gt;Validation should include missing data, tool timeouts, expired permissions, model changes and duplicate requests. Can the system stop, notify, hand off to a person or recover safely? Operations with external effects need particular care: an absent success response is not proof that an operation did not execute.&lt;/p&gt;

&lt;p&gt;Proceed only when quality, efficiency and controllability meet the previously agreed acceptance criteria. If they do not, narrow the use case, revise the approach or stop the pilot.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3. Establish a sustainable service and reorganize roles and costs
&lt;/h3&gt;

&lt;p&gt;A successful technical pilot does not establish commercial readiness. Companies must also assign responsibility for model and knowledge updates, incident handling, permission changes and version rollback, and define how customers raise disputes and accept results.&lt;/p&gt;

&lt;p&gt;Roles need to change accordingly. Business staff own goals and acceptance; engineers own interfaces, execution and recovery; security staff review consequential permissions and data boundaries. One person may hold several roles, but important authorization and acceptance decisions cannot become ownerless because AI performs the work. Training, necessary independent review and role-transition costs also belong in the investment calculation.&lt;/p&gt;

&lt;p&gt;Account for the full cost of delivery. Model calls and compute are only part of it: data preparation, integration, human review, monitoring, incidents, upgrades and customer support must also be paid for. Only comparison with measurable customer benefits can establish whether subscriptions, implementation fees or operating fees are sustainable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4. Expand across customers and use cases only when evidence supports it
&lt;/h3&gt;

&lt;p&gt;Before scaling, separate shared capabilities from industry- or customer-specific configuration and identify permissions requiring individual authorization. Installation packages, skill versions, compatibility documentation, test examples, service documentation and rollback plans should become part of delivery.&lt;/p&gt;

&lt;p&gt;Success with one customer does not establish success with different data, models or business systems. New customers and environments require revalidation of critical assumptions; the first successful case is not a universal conclusion.&lt;/p&gt;

&lt;p&gt;Further investment should depend jointly on reuse, customer benefit, service costs and risk. Persistent dependence on extensive human intervention, insufficient benefits to cover costs, or unresolved consequential side effects are reasons not to force expansion merely to catch a policy window.&lt;/p&gt;

&lt;p&gt;Software transformation therefore comes down to three questions: &lt;strong&gt;Does the customer receive a verifiable improvement? Can the company deliver it sustainably? Do human accountability and risk boundaries remain clear?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk5xxy6bf7d25g8ub2d5m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk5xxy6bf7d25g8ub2d5m.png" alt="Establish a baseline, run a bounded pilot, and use evidence to justify ongoing service and broader reuse." width="800" height="498"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 7. Establish a baseline, run a bounded pilot, and use evidence to justify ongoing service and broader reuse. Source: Author's enterprise transformation analysis informed by Document No. 209.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  XIV. Five Common Misreadings
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Misreading 1: The government will subsidize every AI software company
&lt;/h3&gt;

&lt;p&gt;The plan discusses coordinated funding, first-version software policies and dedicated fiscal funds, but creates no universal subsidy entitlement. Specific support depends on later projects and local rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misreading 2: Connecting a large model completes software transformation
&lt;/h3&gt;

&lt;p&gt;The policy looks at development effectiveness, code quality, product upgrading, service value and real applications. Calling a model is an input, not an outcome.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misreading 3: An agent is just a chat interface
&lt;/h3&gt;

&lt;p&gt;The document also addresses execution frameworks, complex tasks, multi-agent collaboration, development and operating toolchains, system integration and behavior verification. Its scope goes well beyond chatbots.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misreading 4: A Skills market is a place to sell prompts
&lt;/h3&gt;

&lt;p&gt;Enterprise skills need versions, interfaces, dependencies, permissions, security, review and ongoing operations. Without these conditions, they are unlikely to enter critical business systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  Misreading 5: The policy already supplies unified agent standards
&lt;/h3&gt;

&lt;p&gt;The document initiates and advances standards development. Companies cannot present their own concepts as national standards or claim certification before such arrangements exist.&lt;/p&gt;

&lt;h2&gt;
  
  
  XV. The Opportunity: Software as a Continuously Working Organizational Capability
&lt;/h2&gt;

&lt;p&gt;Traditional software helped people record information, execute commands and manage processes. Generative AI enabled software to understand natural language. Agents take it further: software can receive goals, use tools and complete successive tasks.&lt;/p&gt;

&lt;p&gt;As those capabilities enter enterprises, purchasing decisions may change.&lt;/p&gt;

&lt;p&gt;An enterprise may buy a digital capability that continually delivers results, rather than just another system account. It may add an intelligent working role to procurement, R&amp;amp;D, sales, customer service, compliance or operations, rather than simply installing a feature set.&lt;/p&gt;

&lt;p&gt;This helps explain why Document No. 209 discusses software production, agent products, intelligent services and the organization of employment together. It addresses a reconfiguration of software, services and work.&lt;/p&gt;

&lt;p&gt;Models provide intelligence, but do not automatically establish a job. Tools enable actions, but do not establish accountability. Workflows provide a route, but do not replace business judgment. A functioning digital employee system must organize models, tools, skills, identity, responsibilities, authorization, collaboration, verification and recovery into something that can operate over time.&lt;/p&gt;

&lt;p&gt;Document No. 209 therefore opens more than a market for AI-written software.&lt;/p&gt;

&lt;p&gt;It also opens a market in which software becomes agents, agents enter organizations, and organizations reconfigure their productive capacity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: Toward an Agent Software Industry
&lt;/h2&gt;

&lt;p&gt;In 2025, the State Council's &lt;a href="https://www.mee.gov.cn/zcwj/gwywj/202508/t20250827_1126207.shtml" rel="noopener noreferrer"&gt;Opinion on Deepening the “AI Plus” Initiative&lt;/a&gt; established the national direction for integrating AI across the economy and society.&lt;/p&gt;

&lt;p&gt;In 2026, policy has moved progressively into implementation.&lt;/p&gt;

&lt;p&gt;Document No. 414 organizes application service providers and addresses who delivers. The entrepreneurship support plan cultivates AI SMEs and addresses who innovates. Document No. 209 transforms the software industry itself: how software is produced, how it runs continuously, how products and markets develop, and how behavior is verified and governed.&lt;/p&gt;

&lt;p&gt;Across Document No. 414, the entrepreneurship support plan and Document No. 209, policy attention is extending beyond model capabilities to application delivery, entrepreneurs, software products, ongoing services and industrial infrastructure.&lt;/p&gt;

&lt;p&gt;For software firms, this policy path calls attention to upgrading existing software and development methods, cultivating agent products and specialist Skills, and incorporating ongoing operations, behavior verification and workforce transitions into delivery capabilities. Whether these opportunities become viable businesses still has to be demonstrated through products, customers and operating results.&lt;/p&gt;

&lt;p&gt;Document No. 209 thus advances the question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Once agents truly become software, how does China intend to build an industry around them?&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Main Sources
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Ministry of Industry and Information Technology: &lt;a href="https://www.miit.gov.cn/jgsj/xxjsfzs/wjfb/art/2026/art_91a26793271e4c77aca49ecde4472999.html" rel="noopener noreferrer"&gt;Notice Issuing the Implementation Plan for the “AI Plus Software” Special Action&lt;/a&gt;, Gong Xin Bu Xin Fa [2026] No. 209, published September 11, 2026.&lt;/li&gt;
&lt;li&gt;General Office of MIIT: &lt;a href="https://www.miit.gov.cn/jgsj/kjs/wjfb/art/2026/art_331b6f1d28ad410aa9df711acf42dbfe.html" rel="noopener noreferrer"&gt;Notice on the Special Action to Cultivate AI Application Service Providers&lt;/a&gt;, Gong Xin Ting Ke Han [2026] No. 414, published August 31, 2026.&lt;/li&gt;
&lt;li&gt;General Office of MIIT: &lt;a href="https://www.miit.gov.cn/jgsj/qyj/wjfb/art/2026/art_86c400b4473849818629663a94a6d44b.html" rel="noopener noreferrer"&gt;Artificial Intelligence SME Entrepreneurship Support Plan (2026–2028)&lt;/a&gt;, Gong Xin Ting Qi Ye [2026] No. 31, published September 4, 2026.&lt;/li&gt;
&lt;li&gt;State Council: &lt;a href="https://www.mee.gov.cn/zcwj/gwywj/202508/t20250827_1126207.shtml" rel="noopener noreferrer"&gt;Opinion on Deepening the “AI Plus” Initiative&lt;/a&gt;, Guo Fa [2025] No. 11.&lt;/li&gt;
&lt;li&gt;Digital Employee Workshop: &lt;a href="https://joinwell52-ai.github.io/joinwell52/en/industry/2026-09-04-miit-414-ai-application-delivery" rel="noopener noreferrer"&gt;MIIT Document No. 414 Explained: From Model Supply to Application Delivery&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Digital Employee Workshop: &lt;a href="https://joinwell52-ai.github.io/joinwell52/en/industry/2026-09-08-miit-ai-sme-support-plan" rel="noopener noreferrer"&gt;MIIT's AI SME Entrepreneurship Support Plan Explained: An AI Startup Ecosystem Is Taking Shape&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;MIIT's Department of Information Technology Development: &lt;a href="https://www.miit.gov.cn/jgsj/xxjsfzs/gzdt/art/2026/art_45df9463e105446c996568f299e07bca.html" rel="noopener noreferrer"&gt;Seven Questions and an Infographic: Understanding the “AI Plus Software” Implementation Plan&lt;/a&gt;, published September 11, 2026; see questions 3 and 6 in particular.&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/industry/2026-09-15-miit-209-ai-software-policy" rel="noopener noreferrer"&gt;Chinese original&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/en/industry/2026-09-15-miit-209-ai-software-policy" rel="noopener noreferrer"&gt;English version&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>softwaredevelopment</category>
      <category>career</category>
    </item>
    <item>
      <title>Resume the conversation. Keep the old credentials too?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Wed, 16 Sep 2026 04:34:34 +0000</pubDate>
      <link>https://dev.to/joinwell52/resume-the-conversation-keep-the-old-credentials-too-5251</link>
      <guid>https://dev.to/joinwell52/resume-the-conversation-keep-the-old-credentials-too-5251</guid>
      <description>&lt;h1&gt;
  
  
  Resume the conversation. Keep the old credentials too?
&lt;/h1&gt;

&lt;p&gt;A resume test can pass while a persistence check still finds an old credential. We inspected actual temporary JSON files to separate those two outcomes.&lt;/p&gt;

&lt;p&gt;Saving an AI conversation lets you continue tomorrow. Keeping its history is easy to understand.&lt;/p&gt;

&lt;p&gt;But should the key used to launch the assistant stay with it?&lt;/p&gt;

&lt;p&gt;Here, a key is a credential the application uses to access a service. After an account or connection change, the current run may already use a different credential. &lt;strong&gt;Our real-file experiment showed an easily missed state: resume used the current value while the old value remained in the file on disk.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;“It works again” does not answer “were the earlier copies cleaned up?”&lt;/p&gt;

&lt;h2&gt;
  
  
  How did a key get into a conversation record?
&lt;/h2&gt;

&lt;p&gt;The source is &lt;a href="https://github.com/paperclipai/paperclip/pull/13498" rel="noopener noreferrer"&gt;Paperclip #13498&lt;/a&gt;, submitted by electrumnz. Paperclip is an open-source application for coordinating AI assistants. This proposed fix was still open when researched.&lt;/p&gt;

&lt;p&gt;When launching an assistant, an application supplies settings called environment variables. Some can contain service credentials. Upstream reports that this launch environment was saved along with session settings, turning values prepared for a run into potentially long-lived copies.&lt;/p&gt;

&lt;p&gt;That interested us because “save more so we can resume later” sounds reasonable. Yet history and credentials serve different purposes. Continuing yesterday's conversation need not require yesterday's credential.&lt;/p&gt;

&lt;p&gt;The candidate separates them: remove the launch environment when saving, and supply the current environment when resuming. We narrowed our experiment to a practical question: does that preserve the conversation, and what happens to old values already on disk?&lt;/p&gt;

&lt;h2&gt;
  
  
  Synthetic keys, real files
&lt;/h2&gt;

&lt;p&gt;We used recognizable synthetic strings in place of credentials and read no real keys. The candidate's save and load code was connected to temporary files. We supplied an old value, resumed with another current value, and inspected what remained in the file.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Operation&lt;/th&gt;
&lt;th&gt;Old synthetic value remains on disk&lt;/th&gt;
&lt;th&gt;Resume uses current value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Save the original record without filtering&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes; load still uses candidate code&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Save a new record through candidate code&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read a legacy record without saving again&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read legacy record, then save through candidate code&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All four preserved conversation history. The original object supplied to save also stayed unchanged: removing the environment from the disk copy did not modify data the running session might still need in memory.&lt;/p&gt;

&lt;p&gt;The third row is worth pausing over. Resume received the current value, while inspection still found the old value on disk. Both observations were correct; they looked in different places.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbzfuqvvurcxvshikocq5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbzfuqvvurcxvshikocq5.png" alt="Separate persisted history from current launch credentials" width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. Supplying current values during load and removing old environment fields during save affect different locations. Source: our real temporary-file experiment; diagram by the authors.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading an old file does not rewrite it
&lt;/h2&gt;

&lt;p&gt;Reading and changing a file are separate operations.&lt;/p&gt;

&lt;p&gt;After reading the old record, the candidate returns an object carrying the current environment. It does not rewrite the original file during that load. Only when we explicitly saved again did the old value disappear from the tested location.&lt;/p&gt;

&lt;p&gt;That matches the proposal's explanation that the environment field disappears once a record is resaved. An upgrade still leaves a practical question: what happens to old records that are never opened or saved again? If backups retain older copies, who handles them?&lt;/p&gt;

&lt;p&gt;The fix prevented new saves from persisting launch environments and supplied current values on resume. Our experiment clarified the distance between those outcomes and cleaning up every existing copy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does removing a copy make its credential unusable?
&lt;/h2&gt;

&lt;p&gt;That does not follow. “Old” means previously used; it does not mean the service has invalidated the credential. Deleting a local copy does not revoke access.&lt;/p&gt;

&lt;p&gt;Developers therefore need three separate answers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Will future saves keep writing it?&lt;/strong&gt; Inspect newly saved records.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happens to earlier copies?&lt;/strong&gt; Assign responsibility for legacy files and backups.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Have credentials that should no longer work been invalidated?&lt;/strong&gt; That requires credential management, not just file deletion.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our experiment answers how the tested save and load paths behave. It did not rotate or revoke credentials, or clean backups.&lt;/p&gt;

&lt;p&gt;Users can turn this into a concrete question too: after changing an account or connection, does the application explain what history remains and how old connection material is handled? Successful resume is useful, but it should not create an assumption that every old copy has also been addressed.&lt;/p&gt;

&lt;p&gt;That led to our maintainer question: could upgrade guidance distinguish preventing new persistence from handling existing records, and name the responsibility for the latter? It affects how the fix is used, as well as how its save function is written.&lt;/p&gt;

&lt;p&gt;For technical readers: filtering scope, the fifth control, and limitations&lt;/p&gt;

&lt;p&gt;Source author: electrumnz. PR #13498 was OPEN when researched. Candidate: &lt;code&gt;bfe7f568976a2c671184722a6e6294927d6e6fae&lt;/code&gt;; source baseline: &lt;code&gt;9cfa7fd2d13213b73396cf34b9fc95912857ff69&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;We ran the complete candidate &lt;code&gt;session-store.ts&lt;/code&gt;. It wraps the storage interface, removing &lt;code&gt;acpx.session_options.env&lt;/code&gt; on save and injecting the current launchEnv on load. Our underlying store used real temporary JSON files. Directly saving the raw record is a negative control, not the complete old ACPX implementation.&lt;/p&gt;

&lt;p&gt;A fifth boundary case deliberately placed the same synthetic value in an unrelated debug field. The candidate removed the designated environment field but retained that arbitrary copy. This tests filtering of a known structure, not general sensitive-value scanning; the artificial field does not establish another production disclosure path.&lt;/p&gt;

&lt;p&gt;Upstream also discusses multiple seats sharing an operating-system account and proposes disabling shell snapshots. We did not reproduce those paths or start a real ACPX child process. Cross-seat access, backup cleanup, credential rotation, and full product resume were not tested. Upstream's reported 195 tests are not counted as our own execution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-16-current-context" rel="noopener noreferrer"&gt;All five observations, probe, and pinned source&lt;/a&gt;. This does not establish a corresponding CodeFlowMu disclosure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-16-session-without-old-keys" rel="noopener noreferrer"&gt;English original&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-16-session-without-old-keys" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Back in the same workspace, but which visit owns the result?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Wed, 16 Sep 2026 04:33:22 +0000</pubDate>
      <link>https://dev.to/joinwell52/back-in-the-same-workspace-but-which-visit-owns-the-result-39l9</link>
      <guid>https://dev.to/joinwell52/back-in-the-same-workspace-but-which-visit-owns-the-result-39l9</guid>
      <description>&lt;h1&gt;
  
  
  Back in the same workspace, but which visit owns the result?
&lt;/h1&gt;

&lt;p&gt;An A→B→A navigation can make an old request look current again. We controlled its completion time and compared a generation check with a version deliberately missing that check.&lt;/p&gt;

&lt;p&gt;You open workspace A while its file list is loading. You switch to B, then return to A.&lt;/p&gt;

&lt;p&gt;Only now does the first request for A finish. The workspace name is right. It is still a file-list request. &lt;strong&gt;But does the answer belong to this visit or the previous one?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The name matches again, but the first request has not become a new request. We wanted to test whether the application can distinguish those visits.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give each visit a distinguishable marker
&lt;/h2&gt;

&lt;p&gt;The idea comes from &lt;a href="https://github.com/stablyai/orca/pull/20914" rel="noopener noreferrer"&gt;Orca #20914&lt;/a&gt;, submitted by Jinwoo-H. Orca's mobile client requests file information from the computer running an AI assistant, then retains results for later queries.&lt;/p&gt;

&lt;p&gt;This change gives one object responsibility for pending requests and temporarily stored results. Think of it as a record keeper for the current period of use: it needs to know which results still belong and which have become obsolete.&lt;/p&gt;

&lt;p&gt;The implementation attaches a valid-at-the-time marker to a request. A switch or reset retires the old marker. Returning to A does not reactivate the first visit's marker. The code calls these distinct periods “generations.”&lt;/p&gt;

&lt;p&gt;That interested us because persistent assistants encounter switching, reconnection, and recovery. Recognizing the workspace's name alone can miss changes it has gone through. Upstream explicitly tests A→B→A; we used that scenario to build a controlled timing experiment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the first result arrive late
&lt;/h2&gt;

&lt;p&gt;We executed the complete module at a fixed version: start a request for A, hold its answer back, change the scope, and only then deliver the old result. We checked whether it could be stored and read again.&lt;/p&gt;

&lt;p&gt;For comparison, we removed only the check that asks whether the request's marker still belongs to the current generation. This deliberately modified version helps explain the mechanism; &lt;strong&gt;it is not a historical Orca release&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Constructed scenario&lt;/th&gt;
&lt;th&gt;Original module&lt;/th&gt;
&lt;th&gt;Generation check removed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A→B→A; first A result arrives late&lt;/td&gt;
&lt;td&gt;Refuses storage&lt;/td&gt;
&lt;td&gt;Stores old result; later readable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Same name, explicit reset&lt;/td&gt;
&lt;td&gt;Refuses storage&lt;/td&gt;
&lt;td&gt;Stores old result; later readable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tell the owner that the identity epoch changed&lt;/td&gt;
&lt;td&gt;Refuses old request marker&lt;/td&gt;
&lt;td&gt;Stores result, but new scope cannot read it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Identity changes without informing the owner&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Old result remains storable and readable&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Old result remains storable and readable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pass a request marker to another owner&lt;/td&gt;
&lt;td&gt;Refuses: marker belongs elsewhere&lt;/td&gt;
&lt;td&gt;Still refuses&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The first two rows make the central point: rejecting a stale result requires more than the workspace name.&lt;/p&gt;

&lt;p&gt;The fourth row shows that protection also needs an input. If the caller never reports an identity change, the owner cannot know that a new period has begun. We deliberately omitted that signal in a synthetic case; this does not establish a defect in the real login flow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcwto8h6jxfxpuih9j0g1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcwto8h6jxfxpuih9j0g1.png" alt="Two visits to the same workspace" width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. The first request for A arrives after the returning visit has begun a new generation. Source: our controlled schedules; diagram by the authors, not a recording of the mobile UI.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Stored, read, and displayed are separate steps
&lt;/h2&gt;

&lt;p&gt;Another result was less obvious. We changed the outside scope without yet informing the owner. Submitting the old result then returned “committed.”&lt;/p&gt;

&lt;p&gt;On the next read with the new scope, the owner noticed the change and cleared the previously stored contents. That read did not return the old value.&lt;/p&gt;

&lt;p&gt;These temporarily stored contents are the cache. The implementation checks both when accepting a result and when a later reader supplies a scope. Reading only the “committed” verdict misses the second part of the protection.&lt;/p&gt;

&lt;p&gt;The screen is another step. A caller can still hold an old asynchronous result. If it displays that value directly, without another cache read, it needs its own ordering check. Orca's file-search code retains such a display check. Our experiment did not run the mobile UI, so it cannot guarantee what every screen shows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;To understand whether stale content can appear, follow the result: who accepts it, who reads it, and who puts it on screen?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Which changes must the application report?
&lt;/h2&gt;

&lt;p&gt;Workspace selection is an obvious change. Reauthentication or moving execution to another computer can also make earlier results unsuitable.&lt;/p&gt;

&lt;p&gt;The next questions are therefore specific: which events make this data obsolete, where does the application obtain reliable signals of those events, and who supplies them to the request owner? Naming a field “version” does not establish its source.&lt;/p&gt;

&lt;p&gt;Users can contribute concrete observations: after returning to a workspace, did the file list show new content and then briefly revert? Was there a reconnection or account switch? Those details help developers reproduce the right ordering.&lt;/p&gt;

&lt;p&gt;Developers can also consider how many updates one request produces. If it first publishes “loading” and later “ready,” how should a request marker constrain both updates? That is a follow-up design question, not an extension tested in this experiment.&lt;/p&gt;

&lt;p&gt;Returning to the same place does not reverse time. The application still needs to know which visit owns the result.&lt;/p&gt;

&lt;p&gt;For technical readers: additional schedules, boundaries, and pinned version&lt;/p&gt;

&lt;p&gt;Source author: Jinwoo-H. Candidate: &lt;code&gt;89711d6f55670781d23fc1a2d4e2aecf4d725758&lt;/code&gt;; merged on 2026-09-16 UTC. The complete &lt;code&gt;generation-scoped-request-owner.ts&lt;/code&gt; was executed. Eight schedules or boundary cases were run against the original and a mutation removing only the generation comparison: 16 observations.&lt;/p&gt;

&lt;p&gt;The request marker described above is the lease; the caller supplies a scope. The implementation checks both lease ownership and generation. Removing the generation comparison preserves the ownership check. An identity epoch included in scope changes the key, so an accepted old value in the ablation does not necessarily become readable under the new scope.&lt;/p&gt;

&lt;p&gt;Additional cases cover normal same-generation publication and stale-request cleanup. When an old request settled while a new one remained pending, cleanup preserved the new in-flight entry; another load shared its promise and started zero duplicate loads.&lt;/p&gt;

&lt;p&gt;Upstream states that mobile lacks an available negotiated capability epoch; the experiment does not invent one. Its authentication-epoch boundary case does not establish a real authentication path. A separate query sequence protects display, and the cache commit verdict is not a freshness guarantee for a value already held by the caller. Upstream identifies loading/ready publication in a future status-loader migration as unfinished work.&lt;/p&gt;

&lt;p&gt;No real RPC, mobile UI, host migration, or login flow was run. The private TypeScript lease brand is not claimed to isolate hostile JavaScript in the same process.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-16-current-context" rel="noopener noreferrer"&gt;Pinned source, all observations, and runnable probes&lt;/a&gt;. These observations do not establish a corresponding CodeFlowMu defect.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-16-same-place-new-visit" rel="noopener noreferrer"&gt;English original&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-16-same-place-new-visit" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The check failed. Why did the AI remember the answer?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Wed, 16 Sep 2026 04:33:20 +0000</pubDate>
      <link>https://dev.to/joinwell52/the-check-failed-why-did-the-ai-remember-the-answer-2gl8</link>
      <guid>https://dev.to/joinwell52/the-check-failed-why-did-the-ai-remember-the-answer-2gl8</guid>
      <description>&lt;h1&gt;
  
  
  The check failed. Why did the AI remember the answer?
&lt;/h1&gt;

&lt;p&gt;A thrown check is an exception in the current run. To find out whether it also changes the next run, we compared what the model actually received after the error.&lt;/p&gt;

&lt;p&gt;Imagine using an AI assistant: it produces an answer, but the program meant to check that answer fails.&lt;/p&gt;

&lt;p&gt;On your next question, you might expect the unchecked answer to stay out of the conversation. &lt;strong&gt;In our controlled comparison, the old logic raised an error and saved the answer anyway. The next model call received it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The error ended the current attempt without keeping that answer out of the next one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does “remember” mean here?
&lt;/h2&gt;

&lt;p&gt;It means something specific: the application saves conversation history and sends some of it back to the model with the next question. We did not test model training or a chat product's long-term memory feature.&lt;/p&gt;

&lt;p&gt;Once an unchecked answer enters that history, a later model call can receive it as context. An unconfirmed statement might then become a premise for another answer. That possible consequence motivated us; what our experiment directly established was whether the earlier answer was sent again.&lt;/p&gt;

&lt;p&gt;The source was &lt;a href="https://github.com/openai/openai-agents-js/pull/1938" rel="noopener noreferrer"&gt;OpenAI Agents JS #1938&lt;/a&gt;, submitted by jbeckwith-oai. This library connects models, tools, and conversations. Developers can arrange checks on an answer: a check may pass, reject the answer, or fail before reaching a verdict.&lt;/p&gt;

&lt;p&gt;The old implementation already handled “the check says no.” This patch addressed “the check could not finish.” For our research on assistants that keep working across turns, the important question was whether that interruption could leave something behind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Break the checker, then ask another question
&lt;/h2&gt;

&lt;p&gt;Our test model always returned the same distinctive marker text. We made the checks pass, reject, throw an error, or combine one passing check with another that threw. Then we ran another turn and inspected the history actually sent to its model.&lt;/p&gt;

&lt;p&gt;A script supplied the model's answers; real library code ran the conversation and handled storage. We used four execution/storage combinations and compared the old decision with the fixed one:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check outcome&lt;/th&gt;
&lt;th&gt;Old logic: answer replayed&lt;/th&gt;
&lt;th&gt;Fixed logic: answer replayed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;All pass&lt;/td&gt;
&lt;td&gt;4/4&lt;/td&gt;
&lt;td&gt;4/4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explicit rejection&lt;/td&gt;
&lt;td&gt;0/4&lt;/td&gt;
&lt;td&gt;0/4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Check throws&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4/4&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0/4&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One passes, another throws&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4/4&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0/4&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four means two ways of returning an answer—streaming or all at once—crossed with two session-storage implementations. These are constructed cases, not production incident rates. Both versions together produced 32 observations.&lt;/p&gt;

&lt;p&gt;The last two rows show the change: &lt;strong&gt;an answer without completed validation was replayed under the old decision and withheld under the fix.&lt;/strong&gt; Passing answers still remained, as did previously accepted history and the current user question.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3wb4ox9wpxm9rxokms18.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3wb4ox9wpxm9rxokms18.png" alt="Checks and conversation memory" width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. Whether an answer reaches the next turn after a check fails. Source: our pinned-source experiments and saved observations; diagram by the authors. Counts describe constructed cases.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  An unfinished check has not proved the answer wrong
&lt;/h2&gt;

&lt;p&gt;A check can fail because its service is unavailable. That does not establish that the answer itself is bad. Equally, one passing check cannot supply the conclusion of another check that failed.&lt;/p&gt;

&lt;p&gt;The fixed behavior preserves that distinction. It does not declare the answer disproved, and it does not automatically admit the answer as context for later work.&lt;/p&gt;

&lt;p&gt;This suggests a practical design choice. An application can retain what the model produced to help diagnose a failed attempt, while keeping that record out of the next model call. &lt;strong&gt;Retaining an attempt and admitting it into subsequent context need separate rules.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Even completed checks establish acceptance only under their configured requirements. They do not guarantee real-world truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  When checking recovers, which answer should it check?
&lt;/h2&gt;

&lt;p&gt;Suppose the service recovers after the answer has been regenerated or the user has changed the question. Which piece of content does a late “pass” actually approve?&lt;/p&gt;

&lt;p&gt;Neither this patch nor our experiment implements deferred verification. The comparison nevertheless leads to a concrete design question: a future retry would need to associate answer content, conversation turn, and checking rules, so that an old verdict cannot be attached to a new answer.&lt;/p&gt;

&lt;p&gt;Developers can start with a small experiment: make the checker throw, run another turn, and inspect what the model receives. Users can provide useful leads too: does an answer marked as failed later reappear as an already-established fact? Saving the surrounding conversation and error message gives an investigator more to work with than “the AI remembered incorrectly.”&lt;/p&gt;

&lt;p&gt;For technical readers: versions, full method, and reproduction&lt;/p&gt;

&lt;p&gt;The PR by jbeckwith-oai merged on 2026-09-15 UTC. Candidate: &lt;code&gt;457dfff8ce30d19ccbd4a3796ec482356dc19dfa&lt;/code&gt;. The control replaces only &lt;code&gt;runner/guardrails.ts&lt;/code&gt; with its exact predecessor from &lt;code&gt;8ac97dfedd0395beeb38edb0a17b83a9b89c3354&lt;/code&gt;. This was the PR's only changed production module; other candidate code stayed constant.&lt;/p&gt;

&lt;p&gt;The probe uses public &lt;code&gt;run&lt;/code&gt;, ScriptedModel, MemorySession, and an append-only session. It observes both stored items and subsequent model input in streaming and non-streaming modes: four check outcomes × two run modes × two session types × two versions = 32 observations. The mixed-check case establishes a failing batch, not a controlled order of check completion.&lt;/p&gt;

&lt;p&gt;We separately ran the corresponding upstream test file: 71 tests passed, including additional tool-output behavior. That count is not added to the custom observations as a reliability score. No live model service or external database was used, and correctness for every custom store is not established.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-16-current-context" rel="noopener noreferrer"&gt;Pinned source, saved observations, and runnable probes&lt;/a&gt;. This does not establish a corresponding CodeFlowMu defect.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-16-failed-check-memory" rel="noopener noreferrer"&gt;English original&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-16-failed-check-memory" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why did pausing work turn into a task failure?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Tue, 15 Sep 2026 03:29:38 +0000</pubDate>
      <link>https://dev.to/joinwell52/why-did-pausing-work-turn-into-a-task-failure-807</link>
      <guid>https://dev.to/joinwell52/why-did-pausing-work-turn-into-a-task-failure-807</guid>
      <description>&lt;h1&gt;
  
  
  Why did pausing work turn into a task failure?
&lt;/h1&gt;

&lt;p&gt;After pressing pause and then resume, you would probably expect the original work to return to its queue.&lt;/p&gt;

&lt;p&gt;If tasks instead appear as problems requiring intervention, it can look as though pausing broke them.&lt;/p&gt;

&lt;p&gt;A proposed fix in an AI team-management tool exposed that surprising possibility. &lt;strong&gt;Many things can prevent work from starting. Lose the reason, and a wait can turn into a recorded problem.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The same stop needs different next steps
&lt;/h2&gt;

&lt;p&gt;Paperclip organizes AI agents and tasks. It calls a group containing them a “company”; here, think of that as an AI team.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://github.com/paperclipai/paperclip/pull/13443" rel="noopener noreferrer"&gt;this proposal&lt;/a&gt;, MrBlackTongue reports that checks for unfinished work could escalate tasks as budget-blocked while their company was paused. The budget is a spending allowance for a team or agent.&lt;/p&gt;

&lt;p&gt;Both a pause and an exhausted allowance can prevent new work. But they call for different responses. An intentional pause usually means resume later; an exhausted allowance may require adjusting the budget or deciding whether to continue.&lt;/p&gt;

&lt;p&gt;We are building a collaboration system too. This source gave us a precise question to test: how did “wait for now” become “there is a problem to handle”?&lt;/p&gt;

&lt;h2&gt;
  
  
  The reason existed before it was discarded
&lt;/h2&gt;

&lt;p&gt;The initial check returned a specific reason that work could not start, such as “company paused” or “budget exhausted.”&lt;/p&gt;

&lt;p&gt;The recovery helper kept only a yes/no answer: something was blocking the action, or nothing was. Any block entered the budget-handling branch.&lt;/p&gt;

&lt;p&gt;Imagine forwarding a note that says “paused today; continue tomorrow” as simply “cannot work today.” If someone then fills in “because we ran out of money,” they will choose the wrong response. That analogy describes only the information loss; in the actual program, the discarded information was the cause of an invocation block.&lt;/p&gt;

&lt;p&gt;The candidate carries that cause forward and checks whether the company is paused before processing each task.&lt;/p&gt;

&lt;p&gt;Checking once at the beginning, however, is not enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three moments, two very different interpretations
&lt;/h2&gt;

&lt;p&gt;The company can change while a check is underway. We deliberately arranged this sequence:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The company is active at the initial check.&lt;/li&gt;
&lt;li&gt;It pauses; a later check observes that pause as the reason work cannot start.&lt;/li&gt;
&lt;li&gt;It resumes before the result is handled.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Looking up company state at step three returns “active.” But that cannot overturn step two: &lt;strong&gt;the earlier block really was caused by a pause.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The candidate uses the cause returned by that observation, rather than reconstructing an earlier reason from a later state.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff7skdwe17sijnkfxppxb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff7skdwe17sijnkfxppxb.png" alt="A later resume cannot rewrite the observed pause cause" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. Keep the reason with the check that observed it. Source: controlled timing inputs in paperclip.json, not measurements of live concurrent users.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Record which branch the program chooses
&lt;/h2&gt;

&lt;p&gt;We used the original budget-checking code and relevant recovery branches. A test program supplied the planned states and recorded whether the code chose to skip or request budget handling. &lt;strong&gt;No task was actually changed in a database, and this was not a full recovery-system test.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arranged condition&lt;/th&gt;
&lt;th&gt;Original choice&lt;/th&gt;
&lt;th&gt;Candidate choice&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Company already paused&lt;/td&gt;
&lt;td&gt;Request budget handling&lt;/td&gt;
&lt;td&gt;Skip for now&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pause after the initial check&lt;/td&gt;
&lt;td&gt;Request budget handling&lt;/td&gt;
&lt;td&gt;Skip for now&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resume just after a pause cause is read&lt;/td&gt;
&lt;td&gt;Request budget handling&lt;/td&gt;
&lt;td&gt;Skip using that pause cause&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Active company exceeds its allowance&lt;/td&gt;
&lt;td&gt;Request budget handling&lt;/td&gt;
&lt;td&gt;Still request budget handling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent budget-paused within an active company&lt;/td&gt;
&lt;td&gt;Request budget handling&lt;/td&gt;
&lt;td&gt;Still request budget handling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nothing prevents invocation&lt;/td&gt;
&lt;td&gt;Continue to later checks&lt;/td&gt;
&lt;td&gt;Continue to later checks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Entire company paused for a budget reason&lt;/td&gt;
&lt;td&gt;Request budget handling&lt;/td&gt;
&lt;td&gt;Respect the company pause; skip&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The seven comparisons show that the candidate did not simply turn every problem into waiting. An actual budget block within an active company retained its handling.&lt;/p&gt;

&lt;p&gt;The last row adds a distinction: when the entire company is paused, the candidate respects that overall state instead of escalating its individual tasks during this check.&lt;/p&gt;

&lt;h2&gt;
  
  
  After skipping, when does the system return?
&lt;/h2&gt;

&lt;p&gt;We established that the tested code stopped treating these pause scenarios as budget problems.&lt;/p&gt;

&lt;p&gt;But “skip this time” is not the end of the story. After resume, are the tasks reconsidered promptly? Could eliminating a false problem leave work waiting indefinitely instead? Answering that requires testing the complete recovery process.&lt;/p&gt;

&lt;p&gt;The next question is therefore: &lt;strong&gt;how can a system preserve why it stopped earlier while promptly checking again when conditions change?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Developers can examine each handoff: was the reason retained, and did handling infer it from state observed at another time? Users need a clear explanation of which tasks are waiting, which require intervention, and why, after they press resume.&lt;/p&gt;

&lt;p&gt;Preserving the reason is a first step toward making a pause understandable and its next steps manageable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Versions, scope, and reproduction
&lt;/h2&gt;

&lt;p&gt;We pinned predecessor 0e9b24c and candidate 1db5d93. The probe extracts the original getInvocationBlock, budget-block recovery branch, and candidate pause precheck. Database queries return scripted rows; an escalation recorder captures requested handling. Seven scenarios across two versions produced fourteen observations.&lt;/p&gt;

&lt;p&gt;“Request budget handling” means invoking the escalation path, “skip” means skipping the current candidate task, and “continue to later checks” means passing this branch only. None establishes an actual persisted task change or eventual recovery. The causes are company_paused and budget_exhausted.&lt;/p&gt;

&lt;p&gt;We did not run PostgreSQL integration, live concurrent workers, transactions, or the next recovery scan after resume. The author's reported batch incident is not our sample, and no corresponding CodeFlowMu defect was established.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-15-execution-facts" rel="noopener noreferrer"&gt;Source provenance, probes, results, and reproduction instructions&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-15-pause-is-not-failure" rel="noopener noreferrer"&gt;English original&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-15-pause-is-not-failure" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why did an AI tool error leave an extra database row?</title>
      <dc:creator>joinwell52</dc:creator>
      <pubDate>Tue, 15 Sep 2026 03:28:47 +0000</pubDate>
      <link>https://dev.to/joinwell52/why-did-an-ai-tool-error-leave-an-extra-database-row-nnf</link>
      <guid>https://dev.to/joinwell52/why-did-an-ai-tool-error-leave-an-extra-database-row-nnf</guid>
      <description>&lt;h1&gt;
  
  
  Why did an AI tool error leave an extra database row?
&lt;/h1&gt;

&lt;p&gt;If a page says “submission failed,” would you try again?&lt;/p&gt;

&lt;p&gt;Imagine that the first submission was saved, but the success message never made it back. Clicking again could perform the same action twice.&lt;/p&gt;

&lt;p&gt;We tested that possibility with a local database, rather than real orders or payments. We deliberately made a tool save a record and then raise an error. The result was more surprising than expected: &lt;strong&gt;with only one outer attempt allowed, two records appeared.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Did failure mean the action never happened—or that it happened but the reply failed?&lt;/p&gt;

&lt;h2&gt;
  
  
  How one attempt became two writes
&lt;/h2&gt;

&lt;p&gt;The source was &lt;a href="https://github.com/crewAIInc/crewAI/pull/7458" rel="noopener noreferrer"&gt;a proposed CrewAI fix&lt;/a&gt; by Ritiky23. CrewAI helps agents work together and use tools.&lt;/p&gt;

&lt;p&gt;We are building a collaboration system too, so retry behavior matters to us. A tool can return an error after changing a file, database, or external service. This source supplied actual code through which we could count those actions.&lt;/p&gt;

&lt;p&gt;The old implementation handled preparing arguments and running the tool inside the same error handler. If something failed, fallback handling called the tool again with the original arguments.&lt;/p&gt;

&lt;p&gt;A fallback intended to help when argument preparation failed could therefore also repeat a tool that had already saved its result before raising an error.&lt;/p&gt;

&lt;p&gt;The candidate separates the steps. Argument preparation retains its fallback, while an execution error no longer triggers that extra invocation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare attempts with saved records
&lt;/h2&gt;

&lt;p&gt;Our test tool inserted and committed a row in SQLite, which stores a database in a local file, then deliberately raised an error.&lt;/p&gt;

&lt;p&gt;Using the original tool-call methods with the same inputs produced:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Original: calls / saved records&lt;/th&gt;
&lt;th&gt;Candidate: calls / saved records&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Save then error; at most one attempt&lt;/td&gt;
&lt;td&gt;2 / 2&lt;/td&gt;
&lt;td&gt;1 / 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Save then error; at most three attempts&lt;/td&gt;
&lt;td&gt;6 / 6&lt;/td&gt;
&lt;td&gt;3 / 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Normal success; at most three attempts&lt;/td&gt;
&lt;td&gt;1 / 1&lt;/td&gt;
&lt;td&gt;1 / 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argument-description lookup fails; tool then succeeds&lt;/td&gt;
&lt;td&gt;1 / 1&lt;/td&gt;
&lt;td&gt;1 / 1&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;We checked both the synchronous and asynchronous implementations. Their results agreed.&lt;/p&gt;

&lt;p&gt;The first row shows that the fix removed an extra call inside one attempt. The second shows that &lt;strong&gt;three allowed attempts could still save three records&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F87vgluiyj99e38rjttnu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F87vgluiyj99e38rjttnu.png" alt="A later error does not erase a record already saved" width="800" height="400"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Figure 1. Saving a record and returning a result are separate steps. Source: SQLite observations in crew.json. Six writes becoming three does not mean the action happened only once.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The fix addresses a specific repeat. Making repeated business requests take effect only once requires something more: recognizing them as the same operation and finding its earlier result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can we simply check before trying again?
&lt;/h2&gt;

&lt;p&gt;That raises a natural follow-up: before repeating an action, would checking its history be enough?&lt;/p&gt;

&lt;p&gt;Orca, an application for managing AI sessions, provided an adjacent example. In &lt;a href="https://github.com/stablyai/orca/pull/20723" rel="noopener noreferrer"&gt;this proposal&lt;/a&gt;, brennanb2025 addresses whether recovery can treat a message absent from history as never delivered.&lt;/p&gt;

&lt;p&gt;We supplied the relevant code with the same empty-history input. The predecessor concluded “not delivered”; the candidate kept the answer unknown. In a separate case, the candidate also stopped requesting a new message identifier based on a rejection obtained during recovery.&lt;/p&gt;

&lt;p&gt;Before absence is useful evidence, we need to know what the available records can establish. &lt;strong&gt;No success record might mean no success—or an incomplete record.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This experiment tested how code interprets inputs; it did not actually resend a message. Together with the database experiment, it shows why an error or an empty lookup cannot, by itself, justify repeating an action.&lt;/p&gt;

&lt;h2&gt;
  
  
  Avoiding blind retries should not mean waiting forever
&lt;/h2&gt;

&lt;p&gt;Keeping “we cannot tell yet” is more honest. But what happens to subsequent work if the system never resolves the uncertainty?&lt;/p&gt;

&lt;p&gt;That is the most useful next experiment suggested by these results: save the effect, lose its reply, then restart the program. Can recovery use the same operation identifier to find the existing result rather than write again? If the result cannot be found, how can a user investigate or decide what happens next?&lt;/p&gt;

&lt;p&gt;We have not run that complete scenario and cannot announce that it is solved.&lt;/p&gt;

&lt;p&gt;For developers: who stores the operation's identifier and final result, and will they remain queryable after a restart? For users: if a failed-looking submission turns out to have succeeded, would a “check this submission” option help more than another “try again” button?&lt;/p&gt;

&lt;h2&gt;
  
  
  What each experiment actually covered
&lt;/h2&gt;

&lt;p&gt;CrewAI was pinned at a328710 and 2ab3821. We extracted the four complete original use/ause/_use/_ause methods and retained their control flow. Tool selection, telemetry, caching, formatting, and other surrounding dependencies used doubles. SQLite file writes and commits were real. Four scenarios, two routes, and two versions produced 16 observations. There was no full CrewAI integration or real email, order, or payment.&lt;/p&gt;

&lt;p&gt;Orca was pinned at 742a7ad and e6f789b. Complete production modules and their real imports processed five history/submission cases and five disposition cases per version: 20 observations. The two input groups were not joined into an actual send path. The empty-history case also set a consistent boundary and no turn in flight. A recovered rejection remained unconfirmed without requesting a fresh identity.&lt;/p&gt;

&lt;p&gt;A limitation remains: an unmarked synthetic rejection with reason not_delivered still requested a new identity. We did not establish that a real producer creates this combination. Reviews concerning rejection reasons and mobile/orchestration consumers were not reproduced here. Nor did we verify eventual resolution of waiting after an unknown result.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52/tree/main/research/manual-runs/2026-09-15-execution-facts" rel="noopener noreferrer"&gt;Sources, inputs, results, and reproduction instructions&lt;/a&gt;. These observations establish neither production incident frequency nor a corresponding CodeFlowMu defect.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://joinwell52-ai.github.io/joinwell52/en/engineering/2026-09-15-retry-after-error" rel="noopener noreferrer"&gt;English original&lt;/a&gt; · &lt;a href="https://joinwell52-ai.github.io/joinwell52/zh/engineering/2026-09-15-retry-after-error" rel="noopener noreferrer"&gt;中文版&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/joinwell52-AI/joinwell52" rel="noopener noreferrer"&gt;Research repository&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>programming</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
