<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ben Greenberg</title>
    <description>The latest articles on DEV Community by Ben Greenberg (@bengreenberg).</description>
    <link>https://dev.to/bengreenberg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F29526%2Fab3873ff-b15d-48ee-90c2-0006c40df4a1.jpg</url>
      <title>DEV Community: Ben Greenberg</title>
      <link>https://dev.to/bengreenberg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bengreenberg"/>
    <language>en</language>
    <item>
      <title>What stateless MCP changes for gateways</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Mon, 05 Oct 2026 07:26:11 +0000</pubDate>
      <link>https://dev.to/bengreenberg/what-stateless-mcp-changes-for-gateways-74h</link>
      <guid>https://dev.to/bengreenberg/what-stateless-mcp-changes-for-gateways-74h</guid>
      <description>&lt;p&gt;The MCP spec went stateless on July 28. If you run one MCP server, that mostly means deleting your session handling and adding a &lt;code&gt;server/discover&lt;/code&gt; method. If you run a gateway that sits in front of a dozen MCP servers and presents them to clients as one, it changes how the whole thing works.&lt;/p&gt;

&lt;p&gt;I've been following how &lt;a href="https://github.com/theagentrouter/agent-router" rel="noopener noreferrer"&gt;Agent Router&lt;/a&gt; (the AAIF project formerly known as Envoy AI Gateway) is handling this. In September the team merged &lt;a href="https://github.com/theagentrouter/agent-router/pull/2545" rel="noopener noreferrer"&gt;PR #2545&lt;/a&gt;, which adds stateless fan-out for discovery and list requests. It isn't serving traffic yet, and that's on purpose. But it's the clearest look so far at what an MCP gateway looks like once sessions go away.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the gateway used to lean on
&lt;/h2&gt;

&lt;p&gt;Under the old spec, a client opened a session with &lt;code&gt;initialize&lt;/code&gt; and carried an &lt;code&gt;Mcp-Session-Id&lt;/code&gt; on everything after that. Agent Router used that session as its backbone. When a client initialized, the gateway sent &lt;code&gt;initialize&lt;/code&gt; to each backend on the route, collected each backend's session ID and capabilities, and packed all of it into one encrypted session ID that it handed back to the client.&lt;/p&gt;

&lt;p&gt;From then on, each request carried what the gateway needed: which route this was, who the caller was, the upstream session ID for each backend, and what each backend could do. The gateway itself didn't have to remember anything. The client was carrying the gateway's memory around for it.&lt;/p&gt;

&lt;p&gt;The 2026-07-28 spec removes all of that. There's no &lt;code&gt;initialize&lt;/code&gt; handshake, no &lt;code&gt;Mcp-Session-Id&lt;/code&gt;, and no session. Each request carries its own protocol version, client info, and client capabilities in &lt;code&gt;_meta&lt;/code&gt;, plus &lt;code&gt;Mcp-Method&lt;/code&gt; and &lt;code&gt;Mcp-Name&lt;/code&gt; headers so infrastructure can see what the request is without parsing the body. If a client wants to know a server's capabilities up front, it calls &lt;code&gt;server/discover&lt;/code&gt;. Server-initiated requests like elicitation and sampling become multi round-trip requests, where the server returns an &lt;code&gt;input_required&lt;/code&gt; result and the client retries with the answer.&lt;/p&gt;

&lt;p&gt;So the gateway had to find a new home for each piece of information that used to live in that session ID.&lt;/p&gt;

&lt;p&gt;The same &lt;code&gt;tools/list&lt;/code&gt; call looks different under each spec. Under 2025-11-25, the client has to open a session first and carry its ID from then on:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /mcp
Content-Type: application/json

{"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {
  "protocolVersion": "2025-11-25",
  "clientInfo": {"name": "my-agent", "version": "1.4.0"},
  "capabilities": {}
}}

# Response header: Mcp-Session-Id: &amp;lt;encrypted gateway session&amp;gt;
# Client then sends notifications/initialized with that session ID

POST /mcp
MCP-Protocol-Version: 2025-11-25
Mcp-Session-Id: &amp;lt;encrypted gateway session&amp;gt;

{"jsonrpc": "2.0", "id": 2, "method": "tools/list"}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under 2026-07-28, there's no first step. The request names its own method in a header and brings its version, identity, and capabilities in &lt;code&gt;_meta&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /mcp
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/list
Authorization: Bearer &amp;lt;token&amp;gt;

{"jsonrpc": "2.0", "id": 7, "method": "tools/list", "params": {
  "_meta": {
    "io.modelcontextprotocol/protocolVersion": "2026-07-28",
    "io.modelcontextprotocol/clientInfo": {"name": "my-agent", "version": "1.4.0"},
    "io.modelcontextprotocol/clientCapabilities": {}
  }
}}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faym4kn05e7ifk2m63y0z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Faym4kn05e7ifk2m63y0z.png" alt="Agent Router before and after the 2026-07-28 spec. Before, the client opens a session and each backend holds one. After, each request stands alone, and a failed backend is skipped while the others' results are merged." width="800" height="507"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where each piece goes now
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://github.com/theagentrouter/agent-router/blob/main/docs/proposals/012-mcp-new-spec-rc/proposal.md" rel="noopener noreferrer"&gt;design proposal&lt;/a&gt; for this work walks through it piece by piece, and most of the answers turned out to be things Agent Router already had.&lt;/p&gt;

&lt;p&gt;The route name already arrives on each request in a header that Envoy sets. The old code only read it during &lt;code&gt;initialize&lt;/code&gt;. Now it reads it on each request.&lt;/p&gt;

&lt;p&gt;The backend for a single-target call like &lt;code&gt;tools/call&lt;/code&gt; is already encoded in the tool name. Agent Router exposes tools as &lt;code&gt;backend__toolname&lt;/code&gt;, so the name itself tells the gateway where to send the call. Under the new spec that name also shows up in the &lt;code&gt;Mcp-Name&lt;/code&gt; header.&lt;/p&gt;

&lt;p&gt;Per-backend session IDs aren't needed anymore, since modern backends don't have sessions either.&lt;/p&gt;

&lt;p&gt;Identity comes from the bearer token on each request, which the gateway was already checking. With no session, there's also no session to hijack.&lt;/p&gt;

&lt;p&gt;Capabilities are the one piece that needed new work. The gateway now asks each backend with &lt;code&gt;server/discover&lt;/code&gt; and merges the answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What #2545 adds
&lt;/h2&gt;

&lt;p&gt;The PR builds the stateless path for the five requests that have to go to more than one backend: &lt;code&gt;server/discover&lt;/code&gt;, &lt;code&gt;tools/list&lt;/code&gt;, &lt;code&gt;resources/list&lt;/code&gt;, &lt;code&gt;resources/templates/list&lt;/code&gt;, and &lt;code&gt;prompts/list&lt;/code&gt;. The gateway sends the request to each selected backend, merges the responses, and returns one JSON-RPC result.&lt;/p&gt;

&lt;p&gt;A few details stood out to me.&lt;/p&gt;

&lt;p&gt;Backend selection happens on each request now. Under sessions, the set of backends a client could see got fixed at &lt;code&gt;initialize&lt;/code&gt;. In the review thread, the author pointed out that re-evaluating on each request is the right stateless behavior, because a new token can change which backends a caller is allowed to see. Revoke someone's access to a backend and their next &lt;code&gt;tools/list&lt;/code&gt; reflects it. Under the old model, they'd keep seeing it until their session ended.&lt;/p&gt;

&lt;p&gt;Capability merging uses one shared helper for both the old session path and the new stateless path, so a capability advertised by any backend that answered shows up in the merged result the same way in both.&lt;/p&gt;

&lt;p&gt;Cache hints from the new spec (&lt;code&gt;ttlMs&lt;/code&gt; and &lt;code&gt;cacheScope&lt;/code&gt;) get merged by taking the most restrictive values. If one backend says its tool list goes stale immediately, the merged list goes stale immediately. If any backend marks its results private, the merged result is private.&lt;/p&gt;

&lt;p&gt;Put together, a merged response from a route with a &lt;code&gt;docs&lt;/code&gt; backend and a &lt;code&gt;tickets&lt;/code&gt; backend looks roughly like this. The tool names carry their backend as a prefix, which is how a later &lt;code&gt;tools/call&lt;/code&gt; finds its way back. If &lt;code&gt;docs&lt;/code&gt; said its list was good for 60 seconds and &lt;code&gt;tickets&lt;/code&gt; said 30, the merged list gets 30:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"docs__search"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Search the docs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"inputSchema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tickets__create"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Open a ticket"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"inputSchema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"object"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ttlMs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;30000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cacheScope"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"public"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"resultType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"complete"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Partial failure doesn't take down the route. If one backend is unreachable or sends back something malformed, the gateway skips it, logs a warning with the backend's name, records an error metric for it, and returns what the others sent. If none of them answer, you get a 500 instead of an empty list that looks like a healthy server with no tools. One reviewer, mohitgurnani, pushed for that all-failed check to cover the four list handlers, not only discovery, and it went in before merge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it isn't live yet
&lt;/h2&gt;

&lt;p&gt;If you go looking for where the new code gets called, you won't find a caller outside the tests. That's the plan, not an oversight.&lt;/p&gt;

&lt;p&gt;The proposal splits the work into phases. Phase 0 builds the entire stateless path as unreferenced code, one stage of the request lifecycle at a time: classifying the incoming request, selecting and discovering backends, forwarding and merging, then &lt;code&gt;subscriptions/listen&lt;/code&gt;. The proposal says outright that nothing in Phase 0 is reachable and existing behavior doesn't change. Phase 1 is where the dispatcher gets flipped so requests from modern clients start going down the new path.&lt;/p&gt;

&lt;p&gt;So #2545 is staged code waiting for activation. &lt;a href="https://github.com/theagentrouter/agent-router/pull/2692" rel="noopener noreferrer"&gt;PR #2692&lt;/a&gt;, which covers single-target calls like &lt;code&gt;tools/call&lt;/code&gt; and the integration work, was still open when I checked. &lt;code&gt;subscriptions/listen&lt;/code&gt; is a separate follow-up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9qkk8uthj986oigvva1j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9qkk8uthj986oigvva1j.png" alt="Agent Router's rollout of the 2026-07-28 spec as of October 5, 2026. Ingress and classification (#2518) and discovery and list fan-out (#2545) are merged. Single-target calls (#2692) are open, subscriptions/listen is not on main yet, and activation comes after. Mixed-version translation, result caching, and auth hardening are deferred." width="799" height="373"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What this enables once it's switched on
&lt;/h2&gt;

&lt;p&gt;The biggest change is for whoever operates the gateway and the servers behind it. Agent Router's encrypted session ID already let any gateway instance handle any request, but the setup still started with an &lt;code&gt;initialize&lt;/code&gt; sent to each backend on the route, and each backend kept its own session that later requests had to land on. Under the new model, nothing in the chain holds a session. A request carries what it needs, the gateway forwards it, and a stateless backend can answer it from whichever instance the load balancer picks. You can scale gateway instances and backend instances behind plain round-robin balancing without sticky routing, and the proposal explicitly rules out adding a shared session store like Redis for the modern path.&lt;/p&gt;

&lt;p&gt;Routing gets cheaper too. One of the proposal's goals is to let Envoy route on the &lt;code&gt;Mcp-Method&lt;/code&gt; and &lt;code&gt;Mcp-Name&lt;/code&gt; headers without parsing the JSON-RPC body. That opens up per-tool rate limits, per-method policies, and routing decisions made at the edge from headers alone.&lt;/p&gt;

&lt;p&gt;Approval flows get easier to run through a gateway. Under the old spec, a server that wanted to ask the user something mid-call needed an open stream back to the client, and the gateway had to rewrite request IDs to route answers back to the right backend. With multi round-trip requests, the backend returns &lt;code&gt;input_required&lt;/code&gt;, the client retries with the answer, and the gateway routes the retry by tool name like any other call. The proposal plans to pass these through untouched.&lt;/p&gt;

&lt;p&gt;Access changes take effect on the next request. Because backend selection runs per request, revoking a caller's access to a backend takes effect on their next call, not the next time they reconnect.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd still watch
&lt;/h2&gt;

&lt;p&gt;The fan-out talks to backends one after another, so a route with several slow backends adds up. And a partial result looks identical to a complete one from the client's side. If eight tools come back from a ten-backend route, the client can't tell whether that's everything or whether two backends dropped out. The author agreed that surfacing which backends failed would help but deferred it, since the old session path would need the same change to stay consistent. Until that lands, the per-backend logs and metrics are the only place to see it, so make sure your monitoring is watching them.&lt;/p&gt;

&lt;p&gt;There's also a small inconsistency in the discovery response. It includes a line saying how many backends the gateway is aggregating, and that number counts the backends selected, not the ones that answered. With one of two backends down, it still says two.&lt;/p&gt;

&lt;p&gt;The larger open piece is mixed deployments. Phase 1 covers modern clients talking to modern backends. A modern client reaching a backend that still uses sessions, or an older client reaching a stateless backend, needs the gateway to translate between the two models, and the proposal defers that to Phase 2. Its reasoning is that most new backends will ship on the new spec from day one. Whether that holds depends on how fast the MCP servers you already depend on upgrade.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>tutorial</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Is Structured Human Input the Missing Link in Agentic Work?</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Thu, 24 Sep 2026 07:08:43 +0000</pubDate>
      <link>https://dev.to/bengreenberg/is-structured-human-input-the-missing-link-in-agentic-work-2bh1</link>
      <guid>https://dev.to/bengreenberg/is-structured-human-input-the-missing-link-in-agentic-work-2bh1</guid>
      <description>&lt;p&gt;Have you ever told your agent to go do something and walked away? You set the spec clearly. You defined what done should look like. You provided reference materials. You expected to come back later and find it finished. Yet, you come back and its paused waiting for your input.&lt;/p&gt;

&lt;p&gt;The times it needs input from you in unpredictable and uneven. This clearly does not scale.&lt;/p&gt;

&lt;p&gt;A long-running task can pause for hours. By the time you return, you need to know what it’s waiting for, what answers are valid, and how your response reconnects to the right piece of work. Free-form text can carry an answer, but it can’t reliably describe the interaction that produced it.&lt;/p&gt;

&lt;p&gt;A pause needs a contract.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqxptcbx6eieolaybigo4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqxptcbx6eieolaybigo4.png" alt="A2A task pause and resume" width="800" height="68"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A2A already has the right lifecycle shape
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://a2a-protocol.org/latest/" rel="noopener noreferrer"&gt;Agent2Agent protocol&lt;/a&gt; models long-running work as a stateful Task. An agent can move a task into &lt;code&gt;input-required&lt;/code&gt;, and a client can continue the interaction with the same &lt;code&gt;contextId&lt;/code&gt; and, where appropriate, the same &lt;code&gt;taskId&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That gives you a durable handoff point. The task has an identity, a state, and a history.&lt;/p&gt;

&lt;p&gt;A2A messages can also carry structured JSON in a &lt;code&gt;data&lt;/code&gt; Part. So the protocol already has room for a client to return something more dependable than, “yes, please proceed.”&lt;/p&gt;

&lt;p&gt;What’s missing is a standard way for the agent to describe the input it needs when it pauses.&lt;/p&gt;

&lt;p&gt;That gap is the focus of &lt;a href="https://github.com/a2aproject/A2A/discussions/1016" rel="noopener noreferrer"&gt;A2A Discussion #1016&lt;/a&gt;, opened on August 29, 2025. The proposal is direct: when a task enters &lt;code&gt;input-required&lt;/code&gt;, the agent can provide an input schema for the client. A confirmation can become a boolean. A choice can become a list. The client no longer has to infer a form from an English sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Text is a weak contract at the moment you need precision
&lt;/h2&gt;

&lt;p&gt;Why does this break down? Because the text prompt is doing two different jobs.&lt;/p&gt;

&lt;p&gt;It tries to explain a decision to a person, and it tries to define a data contract for software. Those are separate needs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz5hxn8kyc43x6iaoxsiu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz5hxn8kyc43x6iaoxsiu.png" alt="Two jobs, one weak interface" width="800" height="435"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A framework-owned UI can make this work by knowing its own agent runtime. It can render a custom approval component, attach a callback, and translate the response into the format its agent expects. That’s fine inside one application.&lt;/p&gt;

&lt;p&gt;The friction appears when the client and server are independent.&lt;/p&gt;

&lt;p&gt;A generic A2A client can discover an agent through its Agent Card and run a task without knowing the agent’s framework. When that task pauses, the client should not need an adapter for every remote agent’s approval flow. It should be able to receive a declarative request, render an appropriate control, validate the answer, and send structured data back.&lt;/p&gt;

&lt;p&gt;That’s the difference between a pause that works in a demo and one that travels across clients.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fricwf6jk3bgttfqq1eya.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fricwf6jk3bgttfqq1eya.png" alt="Portable human input across clients" width="800" height="538"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Human input is part of the workflow state
&lt;/h2&gt;

&lt;p&gt;Treating human participation as unstructured chat also makes operational questions harder than they need to be.&lt;/p&gt;

&lt;p&gt;Can an operator see which tasks are blocked on a decision? Can a client distinguish a request for approval from a request for missing account data? Can it route an approval to the person who has authority to make it?&lt;/p&gt;

&lt;p&gt;Without a structured request, every client invents its own answer.&lt;/p&gt;

&lt;p&gt;With one, the task can say what it needs at the same layer where it reports that it is waiting. The client can then present the request as a form, an approval screen, or an API action, while preserving the same task identity and response shape.&lt;/p&gt;

&lt;p&gt;That does not require A2A to dictate a single UI. It gives the UI enough information to act without guessing.&lt;/p&gt;

&lt;p&gt;There’s also a practical boundary here. A schema tells a client the shape of input, not whether the client should trust the requested action. The client still owns its policy: who may approve it, what must be shown before approval, and whether the request can proceed. A portable input contract makes those controls easier to build because the request is legible to software.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fekufvdt1h63b96jwz7dc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fekufvdt1h63b96jwz7dc.png" alt="Schema and client policy" width="800" height="151"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  This belongs in an open agent ecosystem
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://aaif.io" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt; brings projects together around the protocols and infrastructure agents need to work across implementation boundaries. A2A’s task model is already a useful example of that work: agents can remain opaque while exposing enough state for another system to collaborate with them.&lt;/p&gt;

&lt;p&gt;Human participation needs the same treatment.&lt;/p&gt;

&lt;p&gt;An agent may call tools through MCP, delegate work through A2A, and run inside a client built by someone else. If it stops for a decision, that decision should not collapse back into framework-specific glue code. The request should move with the task.&lt;/p&gt;

&lt;p&gt;This conversation is a serious next step in how we build agentic applications. The starting point is the question centered in &lt;a href="https://github.com/a2aproject/A2A/discussions/1016" rel="noopener noreferrer"&gt;Discussion #1016&lt;/a&gt; and the work to advance it.&lt;/p&gt;

</description>
      <category>a2a</category>
      <category>ai</category>
      <category>discuss</category>
      <category>news</category>
    </item>
    <item>
      <title>Jev vs Claude: Who Wins?</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Fri, 18 Sep 2026 11:15:07 +0000</pubDate>
      <link>https://dev.to/bengreenberg/jev-vs-claude-who-wins-4mln</link>
      <guid>https://dev.to/bengreenberg/jev-vs-claude-who-wins-4mln</guid>
      <description>&lt;p&gt;I did not need Jev to beat Claude or Kimi on a benchmark. I needed to know whether I could trust it with a decision I actually make regularly, where a false pass matters and uncertainty cannot just be hidden behind confident prose.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;That is a much harder test of the product claim.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision I actually needed
&lt;/h2&gt;

&lt;p&gt;A surprising amount of AI work starts with the assumption that the answer should come from a large language model.&lt;/p&gt;

&lt;p&gt;I wanted to challenge that assumption.&lt;/p&gt;

&lt;p&gt;I had a useful test case in a judging workflow that I have iterated upon numerous times and use multiple times a year. The dataset I used for this experiment was from the 2026 Arbitrum Open House London Online Buildathon. One part of that workflow is an &lt;em&gt;Arbitrum Alignment&lt;/em&gt; gate. Given the evidence already collected for a project, the system has to make one bounded decision:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;satisfied, not_satisfied, or insufficient_evidence.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;That is not a writing task.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The broader judging rubric absolutely contains work that benefits from a frontier LLM. There are 0 to 5 scores that require reading code, interpreting implementation quality, and weighing technical evidence. There are also prose fields where useful explanations need to be generated.&lt;/p&gt;

&lt;p&gt;The full workflow never ends with a final score, rather a brief is given back to me, the human reviewer, to thoroughly analyze, source check and make a final call on.&lt;/p&gt;

&lt;p&gt;I was not testing those.&lt;/p&gt;

&lt;p&gt;I isolated the part of the workflow where the model is not being asked to write, brainstorm, explain, or synthesize an open-ended answer. It is being asked to apply a defined policy to a bounded evidence packet and choose one of three states.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is almost exactly the territory TypeSafe claims Jev is built for.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the question was not whether Jev could replace Claude, GLM, Kimi or Qwen.&lt;/p&gt;

&lt;p&gt;It was whether those models were overkill for this decision in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  A fair fight, deliberately narrow
&lt;/h2&gt;

&lt;p&gt;I tested 102 archived submissions, anonymized as P-numbers.&lt;/p&gt;

&lt;p&gt;Each system received the same JSON evidence packet and the same four-step written decision procedure. Every configuration ran three times across all 102 submissions, producing 306 decisions per variant. I ran it against Claude Sonnet.&lt;/p&gt;

&lt;p&gt;There was no Jev specific simplification of the policy and no additional context given to Sonnet. Both systems had to answer the same question from the same evidence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkv6zg4t11l86wn1qmkdx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkv6zg4t11l86wn1qmkdx.png" alt="A controlled decision test" width="795" height="72"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The narrowness of the test is important here because it relates exactly to what TypeSafe claims Jev is all about.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Jev is not a general-purpose text model. TypeSafe positions it as a System One model for structured decisions, with primitives such as Choice, Score, and Noul rather than free-form generation. Choice, the relevant primitive here, selects among a predefined set of outcomes and returns probabilities and confidence alongside the decision.&lt;/p&gt;

&lt;p&gt;The workload also fit within Jev's current constraints. I was passing a structured evidence packet for one verification gate, not asking it to ingest an entire repository or execute the complete judging workflow.&lt;/p&gt;

&lt;p&gt;That makes this a deliberately unfair place to make a sweeping model comparison.&lt;/p&gt;

&lt;p&gt;It also makes it a very fair place to test Jev's actual claim.&lt;/p&gt;

&lt;p&gt;If a model built specifically for bounded decisions cannot hold up here, then everything else about Jev pretty much falls to the sidelines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same accuracy band, radically different operating cost
&lt;/h2&gt;

&lt;p&gt;No more suspense, the tl;dr is it held up.&lt;/p&gt;

&lt;p&gt;Jev choice plus four diagnostic Nouls reached 100.0% accuracy against the existing labels across the 306 decisions. It produced zero false passes, zero false flags, and was unanimous across all three runs.&lt;/p&gt;

&lt;p&gt;Claude Sonnet 5 at high reasoning reached 99.0%.&lt;/p&gt;

&lt;p&gt;On the headline metric, that is effectively the same accuracy band: 100.0% versus 99.0%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then the economics diverge sharply.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjpbxdjc8l63u1rqo82q9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjpbxdjc8l63u1rqo82q9.png" alt="Same accuracy band, different operating model" width="799" height="411"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Jev's median latency was 378 milliseconds.&lt;/p&gt;

&lt;p&gt;Sonnet high's was 3,554 milliseconds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That made Sonnet roughly 9.4 times slower on this task.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At the measured usage and pricing, 10,000 evaluations would cost approximately $2.27 with Jev and $129.74 with Sonnet high.&lt;/p&gt;

&lt;p&gt;That is roughly a 57x difference.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fulpu5pqx88wro99h5t71.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fulpu5pqx88wro99h5t71.png" alt="Jev vs Sonnet high" width="512" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is where the experiment stops being an interesting model comparison and starts becoming a systems-design question.&lt;/p&gt;

&lt;p&gt;If the output I need is one of three known states, and a decision-specific model can deliver comparable accuracy for roughly one-fiftieth the operating cost, what exactly am I buying from the generative model?&lt;/p&gt;

&lt;p&gt;But raw accuracy gives good marks to both systems.&lt;/p&gt;

&lt;p&gt;Of the 102 gold labels, 77 are satisfied, seven are not_satisfied, and 18 are insufficient_evidence.&lt;/p&gt;

&lt;p&gt;A classifier that simply answered satisfied every time would already look pretty good on an accuracy chart while being completely unacceptable for the purpose of this gate.&lt;/p&gt;

&lt;p&gt;The hard part is not recognizing the obvious passes. It is knowing what to do when the evidence is incomplete, ambiguous, or negative.&lt;/p&gt;

&lt;p&gt;That is where the failures became much more interesting than the headline score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Confidence that changes the workflow
&lt;/h2&gt;

&lt;p&gt;This was the result that changed how I thought about Jev.&lt;/p&gt;

&lt;p&gt;Jev Choice alone made two incorrect decisions across the 306 runs. Both landed in its 0.2 to 0.3 confidence range.&lt;/p&gt;

&lt;p&gt;Sonnet high made three incorrect decisions. All three landed in its 0.9 to 1.0 confidence bin.&lt;/p&gt;

&lt;p&gt;Jev's Expected Calibration Error was 0.037. Sonnet high's was 0.058.&lt;/p&gt;

&lt;p&gt;Those numbers matter, but the workflow implication matters more.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxnrajk6zic3zumvqxpba.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxnrajk6zic3zumvqxpba.png" alt="Confidence changes the workflow" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At a Jev confidence threshold of 0.5, I could have automated 98% of the decisions in this dataset while retaining 100% accuracy among the automated decisions.&lt;/p&gt;

&lt;p&gt;The remaining 2% could have gone to human review.&lt;/p&gt;

&lt;p&gt;That is a far more useful property than simply being right slightly more often.&lt;/p&gt;

&lt;p&gt;A model does not need to be perfectly accurate to be useful in an automated decision pipeline. It needs its uncertainty to correlate with the places where automation becomes dangerous.&lt;/p&gt;

&lt;p&gt;The two systems also failed differently.&lt;/p&gt;

&lt;p&gt;Jev Choice's majority error was in an ambiguous row labeled insufficient_evidence that Jev passed as satisfied.&lt;/p&gt;

&lt;p&gt;But Jev also assigned low confidence to the decision. A simple confidence gate would have stopped it from being automated.&lt;/p&gt;

&lt;p&gt;Sonnet's failures centered on an empty-repository row labeled insufficient_evidence that it classified as not_satisfied.&lt;/p&gt;

&lt;p&gt;That is a less serious outcome operationally, but it exposes another distinction in the policy: "the evidence shows the requirement was not met" and "there is not enough evidence to decide" are not the same state.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sonnet recognized that something was wrong, but it was highly confident in the wrong category.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is why I find the confidence result more serious than the 100.0% accuracy result.&lt;/p&gt;

&lt;p&gt;The interesting question is not just whether a model can make the decision. It is whether the model gives the surrounding system enough information to know when not to let that decision through.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the policy whole
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;What happened when I tried to make the system more deterministic?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Jev exposes smaller decision primitives that make decomposition appealing. Alongside the final Choice, I used four Nouls as diagnostic sub-decisions. A Noul evaluates whether a statement is true and returns a probability.&lt;/p&gt;

&lt;p&gt;That gives you something very helpful: inspectable intermediate state.&lt;/p&gt;

&lt;p&gt;My instinct was to take those four individual judgments and implement the final four-step policy myself in code.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Ask the model the atomic questions.&lt;/li&gt;
&lt;li&gt;Get four probabilities.&lt;/li&gt;
&lt;li&gt;Write the conditionals.&lt;/li&gt;
&lt;li&gt;Remove as much model judgment from the final step as possible.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It sounded safer. It was actually significantly worse.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzi7wy0t6k7r0v7u9o5yf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzi7wy0t6k7r0v7u9o5yf.png" alt="Whole-policy judgment beats reconstruction" width="799" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Jev Choice alone reached 99.3%.&lt;/p&gt;

&lt;p&gt;Choice plus the four diagnostic Nouls reached 100.0%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;My hand-coded composite of those same four Nouls fell to 94.1% and produced six false passes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm0mx4wllgobtgphdl58p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm0mx4wllgobtgphdl58p.png" alt="How the evaluation method changed accuracy" width="512" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is beyond a rounding error, it is the worst failure mode for this gate.&lt;/p&gt;

&lt;p&gt;The individual sub-decisions were useful. My reconstruction of the policy from them was not.&lt;/p&gt;

&lt;p&gt;I learned something important from that:&lt;/p&gt;

&lt;p&gt;Sub-decisions are valuable for diagnosis, auditability, and understanding why a result occurred. But an ordered decision policy is not necessarily equivalent to a bag of independent Boolean answers. The sequence matters. The interaction between conditions matters. The meaning of one piece of evidence can depend on what has already been established elsewhere in the procedure.&lt;/p&gt;

&lt;p&gt;In this experiment, asking Jev to apply the written policy as a whole worked better than asking it for individually reasonable facts and assuming I could perfectly reassemble the judgment afterward.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jev does not replace the rest of this judging workflow.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It accepts text rather than repositories. Its current context constraints make it unsuitable for simply dumping an entire codebase into the model. The judging rubric still includes code-reading scores, broader analysis, and generated prose where a frontier LLM remains the more appropriate tool.&lt;/p&gt;

&lt;p&gt;The result is not "Jev replaces Claude." It is that in this very real workflow where there is a consistent bounded decision gate, a general LLM may not be the right tool anymore.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpi768zkq5a7d4f18aan8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpi768zkq5a7d4f18aan8.png" alt="Different jobs, different models" width="799" height="453"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once I separated those two things in what is needed for a generative LLM model and a System One model, the economics changed by roughly 57x, the latency changed by nearly an order of magnitude, and the confidence signal gave me a credible way to automate almost the entire workload while escalating the uncertain edge cases.&lt;/p&gt;

&lt;p&gt;I'm still figuring out where and how to apply this new model into my workflows, but this experiment has given me a good starting point in understanding its strengths and its weaknesses. More importantly, it's opened my eyes up to the potential of AI systems that incorporate more than the current ways of doing things.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>typescript</category>
      <category>todayilearned</category>
    </item>
    <item>
      <title>The Redirect Is Part of the Threat Model: Hardening MCP Client Connections</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Mon, 14 Sep 2026 15:50:20 +0000</pubDate>
      <link>https://dev.to/bengreenberg/the-redirect-is-part-of-the-threat-model-hardening-mcp-client-connections-3oa0</link>
      <guid>https://dev.to/bengreenberg/the-redirect-is-part-of-the-threat-model-hardening-mcp-client-connections-3oa0</guid>
      <description>&lt;p&gt;I was reading the release notes for the MCP Python SDK while planning this month’s AAIF Ambassador contribution, and one change stopped me: clients on 2.x now follow HTTP redirects only when they remain within the endpoint’s origin.&lt;/p&gt;

&lt;p&gt;That’s a good default.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5o5bmz3x42ftwep2zs9w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5o5bmz3x42ftwep2zs9w.png" alt="Redirect decision tree" width="800" height="1016"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A redirect can move a client from the server it was configured to trust to somewhere else. If your client carries an authenticated session, OAuth state, or tool-discovery requests along for that move, you’ve expanded the set of endpoints that can receive them.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/modelcontextprotocol/python-sdk/releases/tag/v2.2.0" rel="noopener noreferrer"&gt;MCP Python SDK v2.2.0 release&lt;/a&gt;, published September 7, makes the boundary explicit. &lt;code&gt;Client("https://...")&lt;/code&gt;, &lt;code&gt;streamable_http_client&lt;/code&gt;, and &lt;code&gt;sse_client&lt;/code&gt; follow redirects only when the scheme, host, and port stay the same. An &lt;code&gt;http&lt;/code&gt; to &lt;code&gt;https&lt;/code&gt; upgrade on the same host is allowed. A redirect anywhere else fails, and the session remains usable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsx3bzppg2ptzlaqvbgq3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsx3bzppg2ptzlaqvbgq3.png" alt="What an origin-bound redirect check compares" width="799" height="288"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat the configured origin as an authority boundary
&lt;/h2&gt;

&lt;p&gt;MCP clients connect to endpoints that can expose tools and prompts, then ask users and agents to act on the results. That makes the endpoint URL more than a convenience setting. It identifies where your client can establish a session and where it can send authenticated requests.&lt;/p&gt;

&lt;p&gt;Why does a redirect change that? Because HTTP redirects carry authority in a way application code can easily overlook. A client starts with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://mcp.example.com/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the server responds with a redirect to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://other.example.net/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A general-purpose HTTP client may be happy to follow it. An MCP client needs a stricter answer: &lt;code&gt;other.example.net&lt;/code&gt; was never the configured server.&lt;/p&gt;

&lt;p&gt;The Python SDK now applies this rule to OAuth provider requests too. That closes a gap where the MCP transport might enforce an origin boundary while the authentication flow followed a different redirect policy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frq1xfkbd689fu35g61sx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frq1xfkbd689fu35g61sx.png" alt="Redirect policy must cover every client path" width="800" height="137"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is a small rule with a direct consequence: if the server you intended to use lives at another origin, configure that URL explicitly. Don’t let a redirect decide it for you.&lt;/p&gt;

&lt;h2&gt;
  
  
  Redirects are also deployment changes
&lt;/h2&gt;

&lt;p&gt;Same-origin redirects still have valid uses. A server may redirect &lt;code&gt;/mcp&lt;/code&gt; to &lt;code&gt;/mcp/&lt;/code&gt;, or route traffic through a path that preserves the same scheme, host, and port. The SDK release notes call out the trailing-slash case: clients no longer need an &lt;code&gt;httpx.AsyncClient&lt;/code&gt; configured with &lt;code&gt;follow_redirects&lt;/code&gt; for MCP requests.&lt;/p&gt;

&lt;p&gt;That means maintainers should treat endpoint changes as part of their security review, even when the change looks like routing cleanup.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbaxutn9poe3zud6qxfe3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbaxutn9poe3zud6qxfe3.png" alt="Endpoint migration decision" width="800" height="672"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Ask one question before shipping an endpoint redirect: does the final URL have the same scheme, host, and port as the URL users configure?&lt;/p&gt;

&lt;p&gt;If the answer is no, publish the new endpoint and let clients opt into it. A redirect is the wrong migration mechanism when it crosses an origin boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  A review checklist for MCP client maintainers
&lt;/h2&gt;

&lt;p&gt;When you change MCP connection handling, check the following:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkdqevyc63kzcdwnoiw1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdkdqevyc63kzcdwnoiw1.png" alt="MCP redirect review areas" width="800" height="639"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Endpoint normalization:&lt;/strong&gt; Confirm that adding or removing a trailing slash stays within the configured origin. Don’t rewrite a user-provided endpoint to a different host or port.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Redirect behavior:&lt;/strong&gt; Enforce the origin check in every MCP transport you support. The Python SDK’s release names standard client connections, Streamable HTTP, and SSE connections. Your implementation should not leave one transport with a looser policy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;OAuth flows:&lt;/strong&gt; Apply the same redirect rule to authorization, token, and metadata requests that your MCP client makes. Authentication code often uses a separate HTTP client or redirect setting.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Session state:&lt;/strong&gt; When a redirect is rejected, keep the existing session state usable where your transport allows it. The Python SDK reports an &lt;code&gt;MCPError&lt;/code&gt; for disallowed redirects and keeps the session available. An SSE connection fails with &lt;code&gt;httpx2.HTTPStatusError&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Authenticated tool discovery:&lt;/strong&gt; Don’t send credentials or session-bound headers to an origin that wasn’t explicitly configured. This includes requests used to discover which tools a server offers.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Error messages:&lt;/strong&gt; Tell the user the redirect was rejected because it left the endpoint’s origin, and show the target URL only when it is safe to expose. The useful remediation is clear: configure the intended endpoint directly.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tests:&lt;/strong&gt; Add cases for a same-origin path redirect, an &lt;code&gt;http&lt;/code&gt; to &lt;code&gt;https&lt;/code&gt; upgrade on the same host, and a redirect to a different host. Also test OAuth requests separately from transport requests.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;a href="https://aaif.io" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt; gives projects such as MCP a neutral home for the protocols and open-source software that let agents work across tools and frameworks. Interoperability depends on clients agreeing on predictable behavior at boundaries like this one.&lt;/p&gt;

&lt;p&gt;A configured MCP endpoint should remain the authority boundary.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>news</category>
      <category>security</category>
    </item>
    <item>
      <title>Where MCP Ends and A2A Begins: Building a Two-Agent Support Workflow Without Tool-Wrapping</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Fri, 11 Sep 2026 13:20:22 +0000</pubDate>
      <link>https://dev.to/bengreenberg/where-mcp-ends-and-a2a-begins-building-a-two-agent-support-workflow-without-tool-wrapping-3l20</link>
      <guid>https://dev.to/bengreenberg/where-mcp-ends-and-a2a-begins-building-a-two-agent-support-workflow-without-tool-wrapping-3l20</guid>
      <description>&lt;p&gt;When an agent needs help from another service, there is an architectural question to answer first: should that service be exposed as a tool, or should the agent delegate work to another agent?&lt;/p&gt;

&lt;p&gt;The distinction matters when the service on the other side is itself autonomous.&lt;/p&gt;

&lt;p&gt;A diagnostic agent, for example, may need to request missing context, investigate across several systems, maintain state across multiple exchanges, and eventually return a report. Exposing that agent as a function such as &lt;code&gt;run_diagnostics()&lt;/code&gt; can flatten those behaviors into a tool-shaped interface.&lt;/p&gt;

&lt;p&gt;MCP and A2A provide a cleaner separation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; standardizes how models and agents interact with tools, APIs, data sources, and other capabilities. &lt;a href="https://a2a-protocol.org/latest/topics/a2a-and-mcp/" rel="noopener noreferrer"&gt;A2A&lt;/a&gt; standardizes communication between independent agents that need to discover one another, exchange context, delegate work, and manage stateful tasks.&lt;/p&gt;

&lt;p&gt;For the support workflow in this tutorial, the boundary is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MCP is how the support agent directly uses tools and resources available within its operating environment.&lt;/li&gt;
&lt;li&gt;A2A is how the support agent delegates work to another agent that owns its own execution process.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That distinction changes the interface you build.&lt;/p&gt;

&lt;p&gt;A2A and MCP now also share a governance home. A2A became a Growth Stage project of the &lt;a href="https://aaif.io/blog/a2a-joins-aaif" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt;, joining MCP and other open agentic infrastructure projects under the Linux Foundation.&lt;/p&gt;

&lt;p&gt;Let’s build the smallest useful version of that boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  The support workflow
&lt;/h2&gt;

&lt;p&gt;Consider two agents:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A support agent receives a developer issue: “My deployment completed, but the API returns 401.”&lt;/li&gt;
&lt;li&gt;A diagnostic agent knows how to inspect deployment configuration and identity-provider state.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The support agent has MCP tools for operations it performs directly: searching the support knowledge base, reading the ticket, and retrieving deployment metadata from systems it can access.&lt;/p&gt;

&lt;p&gt;Those tools do not have to run on the same machine as the support agent. The important distinction is that the support agent invokes them as capabilities and controls how they are composed into its workflow.&lt;/p&gt;

&lt;p&gt;The diagnostic agent is different. It owns its own process. It may inspect several systems, request additional context, perform a longer-running investigation, or produce a report after several interactions.&lt;/p&gt;

&lt;p&gt;That is where A2A fits.&lt;/p&gt;

&lt;p&gt;The handoff looks like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwt1xf1a640msfysa7r42.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwt1xf1a640msfysa7r42.png" alt="A developer reports an API authentication problem to a support agent, which directly uses MCP tools and delegates the investigation to an A2A diagnostic agent." width="800" height="326"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The useful boundary is not simply “local versus remote.” It is capability use versus agent delegation.&lt;/p&gt;

&lt;p&gt;If an agent needs to invoke a defined capability directly, MCP is usually the appropriate interface. If it needs to delegate a goal to an independently operating agent and let that agent manage its own process, A2A provides the protocol primitives for that interaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the Agent Card
&lt;/h2&gt;

&lt;p&gt;Before the support agent delegates anything, it needs to know whether the diagnostic agent can handle the request.&lt;/p&gt;

&lt;p&gt;A2A uses an &lt;a href="https://a2a-protocol.org/latest/topics/agent-discovery/" rel="noopener noreferrer"&gt;Agent Card&lt;/a&gt; for this. It is a JSON metadata document describing an agent’s identity, service interfaces, supported capabilities, security requirements, and skills.&lt;/p&gt;

&lt;p&gt;For agents using well-known discovery, the card can be published at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://your-agent-domain/.well-known/agent-card.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The official &lt;a href="https://github.com/a2aproject/a2a-cli" rel="noopener noreferrer"&gt;A2A CLI&lt;/a&gt; provides a direct way to inspect it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;a2a card get https://agent.example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace &lt;code&gt;https://agent.example.com&lt;/code&gt; with an A2A agent endpoint you operate or can access.&lt;/p&gt;

&lt;p&gt;The agent card is the client’s description of how the remote agent can be used.&lt;/p&gt;

&lt;p&gt;For this workflow, the support agent needs answers to a few practical questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does this agent advertise a skill related to deployment diagnostics?&lt;/li&gt;
&lt;li&gt;Does it support streaming?&lt;/li&gt;
&lt;li&gt;Which security schemes does it declare?&lt;/li&gt;
&lt;li&gt;Which A2A interfaces and endpoints does it expose?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The support agent can use that information to determine whether the agent is appropriate before sending ticket or deployment context across the boundary.&lt;/p&gt;

&lt;p&gt;This is different from a conventional tool call. With an MCP tool, the client can discover a defined tool schema and invoke that capability. With A2A, the client discovers an agent capable of accepting broader work and interacting through the A2A message and task model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep directly operated capabilities in MCP
&lt;/h2&gt;

&lt;p&gt;Before delegating the investigation, the support agent might use MCP tools such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;search_knowledge_base("deployment completed API 401")
get_ticket_context(ticket_id)
get_deployment_metadata(deployment_id)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are capabilities the support agent is operating directly. It chooses the calls, controls their sequence, and consumes the results as part of its own reasoning process.&lt;/p&gt;

&lt;p&gt;It might be tempting to expose the diagnostic agent as another tool:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;diagnose_deployment(deployment_id, error)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That can work when the interaction really is equivalent to a discrete capability invocation.&lt;/p&gt;

&lt;p&gt;The abstraction becomes less useful when the system behind it needs to behave as an agent. It may need information that was not available when the call started. It may need the caller to authorize access. It may perform work long enough to require lifecycle tracking. It may generate one or more artifacts as the investigation proceeds.&lt;/p&gt;

&lt;p&gt;Those cases can lead the tool wrapper to accumulate custom state, polling, callbacks, and continuation mechanisms.&lt;/p&gt;

&lt;p&gt;A2A already provides protocol concepts for those interactions.&lt;/p&gt;

&lt;p&gt;Use MCP to gather the information needed for a useful delegation request. Then send the goal and relevant context to the diagnostic agent through A2A.&lt;/p&gt;

&lt;h2&gt;
  
  
  Delegate the investigation
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi0ph0aifg46jxzf1rmrm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi0ph0aifg46jxzf1rmrm.png" alt="The support agent fetches the diagnostic agent's Agent Card, gathers local context using MCP tools, and sends an A2A delegation request that returns either a message or task." width="799" height="445"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once the Agent Card has been checked and the initial context collected, the support agent can send a message:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;a2a send &lt;span class="nt"&gt;-a&lt;/span&gt; https://agent.example.com &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Investigate deployment dep_123. The deployment completed, but requests return 401. The affected API is orders."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The official CLI can negotiate among the A2A interfaces advertised by the agent, including JSON-RPC, HTTP+JSON/REST, and gRPC.&lt;/p&gt;

&lt;p&gt;That means the client can interact through the A2A abstraction instead of maintaining a different application-level integration for every agent implementation.&lt;/p&gt;

&lt;p&gt;There is one important detail here: sending an A2A message does not always create a task.&lt;/p&gt;

&lt;p&gt;According to the protocol, the remote agent can return either:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a &lt;code&gt;Message&lt;/code&gt;, for an immediate interaction that does not require task tracking, or&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;Task&lt;/code&gt;, for stateful work that needs lifecycle management.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a diagnostic investigation that takes time or may require additional input, a &lt;code&gt;Task&lt;/code&gt; is the more relevant model.&lt;/p&gt;

&lt;p&gt;A task gives the agents a shared unit of stateful work. It has an identifier, a lifecycle, status information, and potentially artifacts representing outputs of the work.&lt;/p&gt;

&lt;p&gt;A task can remain active while the diagnostic agent investigates. It can move into a state indicating that more input or authorization is required. It can produce artifacts such as a diagnostic report. The client can later retrieve the task or request cancellation.&lt;/p&gt;

&lt;p&gt;That is a different interaction model from invoking a function and waiting for its return value.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stream updates when the user is waiting
&lt;/h2&gt;

&lt;p&gt;Support workflows become difficult to reason about when a remote investigation starts and the calling application receives no information until completion.&lt;/p&gt;

&lt;p&gt;A2A supports streaming for agents that advertise the capability in their Agent Card.&lt;/p&gt;

&lt;p&gt;With the CLI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;a2a send &lt;span class="nt"&gt;-a&lt;/span&gt; https://agent.example.com &lt;span class="nt"&gt;--stream&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Investigate deployment dep_123. The deployment completed, but requests return 401."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For task-based interactions, the A2A protocol can stream task status updates and artifact updates while the work progresses.&lt;/p&gt;

&lt;p&gt;The support agent can translate those events into useful information for the developer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The deployment details have been sent to the diagnostic agent.

The deployment is reachable. The diagnostic agent is now checking identity-provider configuration.

The diagnostic report is ready.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those messages should correspond to real information received from the remote agent. The support agent should not manufacture intermediate progress simply to make the interface appear responsive.&lt;/p&gt;

&lt;p&gt;Streaming is only one option for following work.&lt;/p&gt;

&lt;p&gt;A2A also defines task operations for retrieving task state, listing tasks, canceling active work, subscribing to task updates, and configuring push notifications. These mechanisms support workflows where the client cannot or should not keep one streaming connection open for the entire investigation.&lt;/p&gt;

&lt;p&gt;The A2A CLI exposes task-oriented commands as well, including task inspection and cancellation, which makes the lifecycle visible while developing and debugging an integration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Let the remote agent ask for context
&lt;/h2&gt;

&lt;p&gt;The initial diagnostic request may not contain enough information to finish the investigation.&lt;/p&gt;

&lt;p&gt;That is a normal part of a stateful agent interaction.&lt;/p&gt;

&lt;p&gt;Suppose the diagnostic agent determines that it needs the deployment region before it can continue. It can return the task in an &lt;code&gt;input-required&lt;/code&gt; state and explain what information is missing.&lt;/p&gt;

&lt;p&gt;The support agent can then use an MCP tool it already controls:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;get_deployment_metadata("dep_123")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Suppose the result contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;region = us-east-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The support agent can send that information back while referencing the existing A2A task rather than beginning an unrelated investigation.&lt;/p&gt;

&lt;p&gt;This preserves an important boundary.&lt;/p&gt;

&lt;p&gt;The diagnostic agent does not need direct access to the support agent’s deployment metadata tool. The support agent remains responsible for its own systems and decides what context crosses the agent boundary.&lt;/p&gt;

&lt;p&gt;That can also reduce unnecessary privilege sharing. Instead of giving the diagnostic agent standing access to another system, the support agent can provide the specific piece of context needed for the current task.&lt;/p&gt;

&lt;p&gt;A2A uses task and context identifiers to support these continued interactions. A &lt;code&gt;taskId&lt;/code&gt; identifies the stateful unit of work, while a &lt;code&gt;contextId&lt;/code&gt; can group related interactions.&lt;/p&gt;

&lt;p&gt;The result is a multi-turn collaboration between agents without requiring either side to expose its internal tools, memory, or implementation to the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  A decision rule you can use
&lt;/h2&gt;

&lt;p&gt;Before adding another integration to an agent, ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this agent need to invoke a capability, or delegate a goal to another independently operating agent?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If it needs to invoke a capability, expose that capability through MCP.&lt;/p&gt;

&lt;p&gt;If it needs to delegate a goal, discover the other agent through its Agent Card, verify that its advertised skills and security requirements fit the request, and communicate through A2A.&lt;/p&gt;

&lt;p&gt;For this support workflow, that gives a clear split:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Read ticket context&lt;/td&gt;
&lt;td&gt;MCP&lt;/td&gt;
&lt;td&gt;The support agent directly operates the ticket-system integration.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Search troubleshooting documentation&lt;/td&gt;
&lt;td&gt;MCP&lt;/td&gt;
&lt;td&gt;The support agent directly invokes a retrieval capability.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Investigate a deployment through a diagnostic agent&lt;/td&gt;
&lt;td&gt;A2A&lt;/td&gt;
&lt;td&gt;The diagnostic agent owns the investigation and its execution process.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Request additional diagnostic context&lt;/td&gt;
&lt;td&gt;A2A + MCP&lt;/td&gt;
&lt;td&gt;A2A carries the request between agents; the support agent can use MCP to retrieve information from its own systems.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Return a diagnostic report&lt;/td&gt;
&lt;td&gt;A2A&lt;/td&gt;
&lt;td&gt;The report can be represented as an artifact produced by the delegated work.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two agents do not need to belong to different companies, frameworks, or repositories for this boundary to be useful.&lt;/p&gt;

&lt;p&gt;They can initially live in the same codebase.&lt;/p&gt;

&lt;p&gt;The protocol boundary becomes especially valuable when the diagnostic agent later moves to another framework, team, service, or organization. The support agent can continue using MCP for its tools while communicating with the diagnostic system as an agent rather than reducing it to a tool-shaped wrapper.&lt;/p&gt;

&lt;p&gt;That is the practical boundary between MCP and A2A: MCP equips an agent with capabilities. A2A gives independently operating agents a standard way to work together.&lt;/p&gt;

&lt;p&gt;The official A2A CLI is a useful place to start experimenting with that boundary:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/a2aproject/a2a-cli" rel="noopener noreferrer"&gt;https://github.com/a2aproject/a2a-cli&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>mcp</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Wiring a Reachy Mini into OpenClaw without trusting the robot</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Tue, 08 Sep 2026 11:25:41 +0000</pubDate>
      <link>https://dev.to/bengreenberg/wiring-a-reachy-mini-into-openclaw-without-trusting-the-robot-3lgh</link>
      <guid>https://dev.to/bengreenberg/wiring-a-reachy-mini-into-openclaw-without-trusting-the-robot-3lgh</guid>
      <description>&lt;p&gt;I put a &lt;a href="https://pollen-robotics.com/reachy-mini/" rel="noopener noreferrer"&gt;Reachy Mini&lt;/a&gt; in the living room.&lt;/p&gt;

&lt;p&gt;It is a small desktop robot from Pollen Robotics, now part of Hugging Face. A head on a moving body, two antennas, a camera, a speaker, and a microphone array. Inside is a Raspberry Pi CM4 running a daemon that exposes the motors, the audio and the camera over an HTTP API on port 8000, plus an app system: you install a Python package onto the robot and the daemon runs it as the current app.&lt;/p&gt;

&lt;p&gt;Out of the box it ships with demos. It waves, it dances, it follows a face. &lt;/p&gt;

&lt;p&gt;What I wanted was a voice touchpoint for the family, and a way to interact with the family while I am traveling. I already run OpenClaw in the house. I call my instance Jeeves, which is my sarcastic but helpful British butler. It holds my calendar, the home automation devices in every room, a Jewish holiday and Sabbath scheduler, a knowledge base, and the skills that act on all of it. I talk to it through Telegram and a dashboard.&lt;/p&gt;

&lt;p&gt;So the robot is a face and a microphone in the room where my family sits, and Jeeves is everything worth saying back. Wiring the two together was my last weekend's project.&lt;/p&gt;

&lt;h2&gt;
  
  
  The obvious wiring, and why I did not ship it
&lt;/h2&gt;

&lt;p&gt;The obvious version is to put the agent behind the robot. Install an app on the Reachy that captures audio, sends the transcript to the OpenClaw gateway, and speaks the answer. The robot becomes a client of my agent. &lt;/p&gt;

&lt;p&gt;I got that working and then took it apart, for two reasons.&lt;/p&gt;

&lt;p&gt;The first is that the robot is not a machine I can trust. Its own daemon API has no authentication of any kind. I checked this rather than assumed it: fetch &lt;code&gt;/openapi.json&lt;/code&gt; off the robot and you get 100 endpoints and zero security schemes. Anything on the LAN can drive the motors, open the camera, or stop the running app. The security threat is low, but still not tolerable. We maintain separate guest WiFi, but even with that, I didn't like that exposure.&lt;/p&gt;

&lt;p&gt;If the agent runs on the robot, then the robot holds a gateway token. The gateway token reaches an agent with my calendar, my house and my shell. Not acceptable to me.&lt;/p&gt;

&lt;p&gt;The second reason showed up in the audit trail once it was answering questions in the room. A general question was a full agent turn: a system prompt around 30k tokens carrying tool profiles, the skills index and the workspace bootstrap, on a persistent session that had grown to 51k, with a reasoning model spending about 500 tokens of thought before its first word.&lt;/p&gt;

&lt;p&gt;Median 15 seconds to answer "what's the weather today?". &lt;/p&gt;

&lt;p&gt;At fifteen seconds nobody in the room waits for the answer. They go find a phone instead, and the robot goes back to being just a cute toy and an ornament on the shelf.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape I landed on
&lt;/h2&gt;

&lt;p&gt;One rule drives the whole design: the robot is untrusted, and everything that could leak lives behind all the security I invested in my OpenClaw setup.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fauarf1a96n4ov3py6ikg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fauarf1a96n4ov3py6ikg.png" alt="The untrusted Reachy Mini can communicate only with a scoped broker, which routes approved requests to deterministic home handlers or a lean OpenClaw agent. Appears after: " width="799" height="136"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Living room                          Mac mini
Reachy Mini (CM4)                    reachy-broker :8092
  reachy-mini-daemon :8000   ---&amp;gt;      one scoped bearer token
  jeeves_hub app             Tailscale 
    mic, speaker, motors               redaction at the boundary
    one token, no data                 whisper / OpenClaw gateway
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The robot runs one app, &lt;code&gt;jeeves_hub&lt;/code&gt;. It does wake-word matching, voice activity detection, motion and expressions, and it holds exactly one bearer token scoped to the broker. It never sees a credential, a model, or a calendar entry. Audio goes up, a policy and a reply come down.&lt;/p&gt;

&lt;p&gt;The broker is a small Python HTTP server on the machine. It is the only path between the robot and my data, and it is the only thing that talks to OpenClaw.&lt;/p&gt;

&lt;p&gt;If the robot were fully compromised tomorrow, what the attacker gets the highly restricted intent allowlist and nothing else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tailscale, not the LAN
&lt;/h2&gt;

&lt;p&gt;The broker binds the tailnet interface and not &lt;code&gt;0.0.0.0&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;The robot is on my tailnet, so the broker addresses it by its tailnet address too, never its LAN address. Nothing else in the house can reach the broker, and the traffic between the two is WireGuard rather than plaintext HTTP across a shared network.&lt;/p&gt;

&lt;p&gt;Tailscale also handles the remote case as well. When I "teleport" in from a hotel, that is direct WireGuard, not the vendor's WebRTC path. &lt;/p&gt;

&lt;h2&gt;
  
  
  The intent allowlist
&lt;/h2&gt;

&lt;p&gt;This is an essential part of the design.&lt;/p&gt;

&lt;p&gt;The robot cannot phrase a request. It names an intent and passes typed arguments. Free text only ever reaches one intent, &lt;code&gt;general.ask&lt;/code&gt;, which is also the least privileged one: no house access, no calendar, no files.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;The intent allowlist.

This is the closed set of things a shared-room robot may ask for. An unknown
intent is refused before any data source is touched, so the blast radius of a
fully compromised robot is exactly what is listed here and nothing else.
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Routing a transcript to an intent is done with rules, not a model. That was deliberate. An LLM router adds a second model round trip to every turn, and it can be talked into picking a different intent by whatever is said in the room.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;AC_MENTION&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\b(?:a\.?[/ ]?c\.?|air[ -]?condition(?:er|ing|ers)?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;|air[ -]?cons?|cooling)\b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;I&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The doctrine there is one line: an imperative actuates, a question never does, and a command that names no known room asks which one. "Is the ac on" is a read. "Turn on the ac" is a write. Anything ambiguous falls to the read.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fapj440ajgh9pk91zhuud.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fapj440ajgh9pk91zhuud.png" alt="A deterministic router distinguishes commands, questions, ambiguous requests, and unknown intents before any home data or device is accessed." width="799" height="502"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Redaction happens on the broker side of the boundary, before a word is spoken. My calendar status is coarse: busy or free, until when, and a whereabouts word from a closed vocabulary. Never event titles, locations, attendees or company names. &lt;/p&gt;

&lt;h2&gt;
  
  
  The lean OpenClaw agent
&lt;/h2&gt;

&lt;p&gt;This is how I dropped the response time on Reachy from 15 seconds to about 3 seconds.&lt;/p&gt;

&lt;p&gt;Instead of routing general questions at &lt;code&gt;main&lt;/code&gt;, I gave the robot its own agent in &lt;code&gt;config/openclaw.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Reachy family hub"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Lean voice agent for the living-room robot: no skills, no tools, no workspace context. The persona rides in each message."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"workspace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"/Users/you/.openclaw/workspace-reachy"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"FILL_IN_YOUR_LLM_MODEL_HERE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"fallbacks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"FALLBACK_MODEL_1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"FALLBACK_MODEL_2"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"skills"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"profile"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"minimal"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"heartbeat"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"every"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0m"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contextInjection"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"never"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"thinkingDefault"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"off"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reasoningDefault"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"off"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"memory"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"search"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every line there is removing something. No skills means no skills index in the prompt. The minimal tool profile means no coding tools. &lt;code&gt;contextInjection: never&lt;/code&gt; means the workspace is not injected, so the persona files on disk are documentation for me rather than tokens I buy on every question. Thinking and reasoning off, because a living room answer is not a research task.&lt;/p&gt;

&lt;p&gt;The same question that measured 15 seconds now measures about 2.9 seconds for the model leg and 3.4 to 3.9 seconds for the whole turn including speech synthesis, at roughly 5k prompt tokens instead of 30k. About a tenth of the cost.&lt;/p&gt;

&lt;p&gt;The persona is composed per turn by the broker and rides in the message, so every request is stateless. No &lt;code&gt;user&lt;/code&gt; field, no session header. A persistent session would re-buy the persona on every question anyway.&lt;/p&gt;

&lt;p&gt;The broker POSTs to the gateway's OpenAI-compatible endpoint and addresses the agent by name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openclaw/reachy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# the agent, not a model id
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_completion_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;GATEWAY_URL&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
             &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;method&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;POST&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is a second agent, &lt;code&gt;openclaw/reachy-deep&lt;/code&gt;, with the same lean shape and a stronger model, reached only when somebody says "think carefully". Everything else stays on the fast path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping the model out of the house
&lt;/h2&gt;

&lt;p&gt;The lean agent handles general knowledge. It does not touch the house, and that is the other half of the speed story.&lt;/p&gt;

&lt;p&gt;Before the deterministic router existed, "turn off the kitchen light" reached the model, which rediscovered the automation skill with its shell tool on every single request. 43 seconds when it worked, and a timeout when it did not.&lt;/p&gt;

&lt;p&gt;Now every device is a closed table and a deterministic intent, and a light takes one to two seconds. Same for the time, for greetings, for "what can you help with", which is spoken from a command guide generated out of the router itself so it can never advertise a phrase the robot does not understand.&lt;/p&gt;

&lt;p&gt;Zero tokens for any of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it under launchd
&lt;/h2&gt;

&lt;p&gt;Both the broker and the speech service are launchd agents on the Mac.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;
...

&lt;span class="nb"&gt;cd&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/path/to/reachy-broker"&lt;/span&gt;
&lt;span class="nb"&gt;exec&lt;/span&gt; /opt/homebrew/bin/python3 &lt;span class="nt"&gt;-m&lt;/span&gt; api.server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's all it takes it boot it up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it feels like now
&lt;/h2&gt;

&lt;p&gt;You say "hey Jeeves" and then a sentence. The hub matches the wake phrase locally, streams the audio to the broker, whisper transcribes it, the router picks an intent, and either a deterministic handler or the lean agent answers. &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk216exb6knqp2at5usl1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk216exb6knqp2at5usl1.png" alt="The request path transcribes audio, routes house requests to deterministic handlers, and sends general questions to a stateless lean agent." width="799" height="469"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;House questions and house commands never leave the Mac. General questions go to the LLM through the gateway and never carry a fact about my family. The robot holds one token that can do exactly the things on a list I wrote.&lt;/p&gt;

&lt;p&gt;The part I did not expect to care about is that setting thatg boundary made the fun parts possible. Once the robot could only ever name an intent, I stopped worrying about what it might be talked into doing and started adding things: games, ambient motion randomly making the kids laugh. None of that needed a new trust decision, because there is only one, and it is enforced in a single file.&lt;/p&gt;

&lt;p&gt;If you are wiring a device you cannot trust into an agent that can act on your life, put a broker between them and give the device a closed list of things it may ask for. Then, once you take care of that bit, you can just get down to building it for both productivity and joy.&lt;/p&gt;

</description>
      <category>openclaw</category>
      <category>ai</category>
      <category>productivity</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Vector Search Is Still the Memory Layer Agents Actually Need</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Thu, 27 Aug 2026 07:37:45 +0000</pubDate>
      <link>https://dev.to/bengreenberg/vector-search-is-still-the-memory-layer-agents-actually-need-50dn</link>
      <guid>https://dev.to/bengreenberg/vector-search-is-still-the-memory-layer-agents-actually-need-50dn</guid>
      <description>&lt;p&gt;When I was working on &lt;a href="https://pragprog.com/titles/bgvector/vector-search-with-javascript/" rel="noopener noreferrer"&gt;&lt;em&gt;Vector Search with JavaScript&lt;/em&gt;&lt;/a&gt;, vector search was a hot topic. By the time the  book was published some people had begun saying that because of LLMs and their advances, we have moved beyond vector search.&lt;/p&gt;

&lt;p&gt;This couldn't be farther from the truth. LLMs and agentic development is amazing, but it often gets things wrong. They don't fail because the model is weak always, but they fail because the right context can be sitting somewhere else and they had no idea that it existed.&lt;/p&gt;

&lt;p&gt;Your docs are in one place. Tool outputs are in another. Prior decisions are in chat history, issue comments, &lt;code&gt;AGENTS.md&lt;/code&gt;, local files, and half a dozen API responses. You can paste more into the prompt, but that gets expensive and messy fast.&lt;/p&gt;

&lt;p&gt;Vector search gives agents a memory layer they can inspect, query, move, and rebuild.&lt;/p&gt;

&lt;p&gt;That still matters in an LLM-first world.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://aaif.io" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt; is a good place to frame this because AAIF is about open agentic infrastructure: MCP, goose, AGENTS.md, agentgateway, and the protocols around them. If agents are going to work across tools and runtimes, memory can’t live as a hidden feature inside one hosted product. It needs to be part of the system you can reason about.&lt;/p&gt;

&lt;h3&gt;
  
  
  The prompt is the wrong database
&lt;/h3&gt;

&lt;p&gt;A prompt is a request. It’s not a storage layer.&lt;/p&gt;

&lt;p&gt;Once you treat the prompt as storage, every workflow starts to rot. You add summaries. Then summaries of summaries. Then a “context” block, and then a "context" block for the original context block.&lt;/p&gt;

&lt;p&gt;That doesn’t scale for project-specific agents.&lt;/p&gt;

&lt;p&gt;You need retrieval that can answer questions like:&lt;/p&gt;

&lt;p&gt;Which migration introduced this column?&lt;/p&gt;

&lt;p&gt;What did the tool return the last time this failed?&lt;/p&gt;

&lt;p&gt;Which internal doc explains this service boundary?&lt;/p&gt;

&lt;p&gt;What did we decide about auth in the previous session?&lt;/p&gt;

&lt;p&gt;Why does that happen? Because agents need working memory and reference memory at the same time. The model can reason over the current task, but your project context lives outside the model. Vector search gives you a way to fetch the few pieces that match the current intent instead of dragging the whole project into every turn.&lt;/p&gt;

&lt;h3&gt;
  
  
  MCP makes retrieval a first-class interface
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://modelcontextprotocol.io/docs/2026-07-28/getting-started/intro" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; gives AI applications a standard way to connect to external systems. MCP servers can expose tools and resources, and resources are identified by URIs in the spec.&lt;/p&gt;

&lt;p&gt;That maps cleanly to vector search.&lt;/p&gt;

&lt;p&gt;You can build an MCP server with tools like:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;search_project_context(query, filters)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;fetch_context_chunk(uri)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;upsert_tool_result(source, content, metadata)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;list_context_sources(project_id)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The vector database doesn’t need to know about the agent. The agent doesn’t need to know about the vector database. MCP becomes the contract between them.&lt;/p&gt;

&lt;p&gt;That contract matters when you want portability. Today your agent might run in an IDE. Tomorrow it might run in a local runtime like &lt;a href="https://aaif.io/projects/goose" rel="noopener noreferrer"&gt;goose&lt;/a&gt;. The retrieval layer should move with you.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should go into agent memory?
&lt;/h3&gt;

&lt;p&gt;Start with the things you already look up manually.&lt;/p&gt;

&lt;p&gt;Index your docs, READMEs, runbooks, schema notes, generated API references, issue threads, and selected tool outputs. Store the raw text or clean markdown. Keep metadata with every chunk: source URI, file path, repo, commit SHA when you have it, timestamp, author if useful, and content type.&lt;/p&gt;

&lt;p&gt;Then be strict about retrieval.&lt;/p&gt;

&lt;p&gt;Don’t return anonymous chunks. Return chunks with source links.&lt;/p&gt;

&lt;p&gt;Don’t rely on similarity alone. Use metadata filters.&lt;/p&gt;

&lt;p&gt;Don’t treat old context and new context equally. Add recency where the domain changes.&lt;/p&gt;

&lt;p&gt;Don’t make the agent trust memory blindly. Give it enough source data to quote the file, open the URI, or ask for confirmation before making a risky change.&lt;/p&gt;

&lt;p&gt;Vector search is useful because it’s probabilistic. Agent memory is useful when that probability is wrapped in provenance.&lt;/p&gt;

&lt;h3&gt;
  
  
  A small useful pattern
&lt;/h3&gt;

&lt;p&gt;A practical agent memory loop can stay straightforward.&lt;/p&gt;

&lt;p&gt;First, chunk source material by meaning, not by arbitrary token count. Function-level chunks work better than splitting every thousand characters in code-heavy repos. Section-level chunks work better for docs.&lt;/p&gt;

&lt;p&gt;Then embed each chunk and store it with metadata.&lt;/p&gt;

&lt;p&gt;At runtime, the agent turns the current task into a retrieval query. The MCP server searches the vector index, filters by project or source type, and returns a small set of candidates with scores and URIs. The agent fetches the best chunks, reads them, and decides what to do next.&lt;/p&gt;

&lt;p&gt;That’s enough for many workflows.&lt;/p&gt;

&lt;p&gt;You can add hybrid search when exact identifiers matter. You can add reranking when your top results are noisy. You can add write-back when tool results become useful future context. But the base shape stays the same: retrieve, inspect, act.&lt;/p&gt;

&lt;h3&gt;
  
  
  Vector search also makes memory debuggable
&lt;/h3&gt;

&lt;p&gt;When an agent gives a bad answer, you need to know whether the reasoning failed or retrieval failed.&lt;/p&gt;

&lt;p&gt;Those are different problems.&lt;/p&gt;

&lt;p&gt;If retrieval returned the wrong chunks, fix chunking, filters, metadata, or ranking. If retrieval returned the right chunks and the model ignored them, fix the prompt or tool policy. If the index is stale, fix ingestion.&lt;/p&gt;

&lt;p&gt;Without an inspectable retrieval layer, all of that collapses into “the agent was wrong.”&lt;/p&gt;

&lt;p&gt;You can log the query, returned chunk IDs, scores, metadata filters, and final sources used. You can replay the retrieval step without running the full agent. You can delete bad documents from the index. You can rebuild from source.&lt;/p&gt;

&lt;p&gt;That is what I would call operational memory.&lt;/p&gt;

&lt;p&gt;Vector search didn’t become obsolete because models got better. It became more useful because agents now have more places to look.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
      <category>vectordatabase</category>
    </item>
    <item>
      <title>What a good Agents.md should teach an agent on day one</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Mon, 03 Aug 2026 15:38:29 +0000</pubDate>
      <link>https://dev.to/bengreenberg/what-a-good-agentsmd-should-teach-an-agent-on-day-one-3nen</link>
      <guid>https://dev.to/bengreenberg/what-a-good-agentsmd-should-teach-an-agent-on-day-one-3nen</guid>
      <description>&lt;p&gt;I hit this last week while working inside my own OpenClaw workspace: the agent had access to the right files, the right tools, and the right project context, but the useful behavior didn't come from any one magic prompt. It came from a small stack of durable instructions.&lt;/p&gt;

&lt;p&gt;The root &lt;code&gt;AGENTS.md&lt;/code&gt; said what to read first. &lt;code&gt;SOUL.md&lt;/code&gt; defined the assistant's operating posture. &lt;code&gt;USER.md&lt;/code&gt; gave personal context. &lt;code&gt;TOOLS.md&lt;/code&gt; separated reusable tool behavior from local machine details. Skill docs explained when to load specialized workflows.&lt;/p&gt;

&lt;p&gt;That structure has proven useful for me time and time again.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2e2un8p08dzc7s2vp2iv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2e2un8p08dzc7s2vp2iv.png" alt="A stack showing AGENTS.md as the routing file that points agents to posture, user context, local tool notes, and deeper skill file" width="799" height="249"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AGENTS.md, now part of the &lt;a href="https://aaif.io" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt; ecosystem hosted by the Linux Foundation, gives developers a plain Markdown place to tell coding agents how to work in a repo. The format is intentionally simple. The hard part isn't the file. The hard part is deciding what deserves to live in it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the first five minutes
&lt;/h2&gt;

&lt;p&gt;A good &lt;code&gt;AGENTS.md&lt;/code&gt; should answer one question first: what should the agent do before touching code?&lt;/p&gt;

&lt;p&gt;In my workspace, the startup path is explicit:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Read &lt;code&gt;SOUL.md&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Read &lt;code&gt;USER.md&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Read today's and yesterday's daily memory files&lt;/li&gt;
&lt;li&gt;In a main session, read &lt;code&gt;MEMORY.md&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That gives the agent a boot order. It doesn't need to guess which file matters, whether memory is allowed, or whether private context belongs in a shared chat.&lt;/p&gt;

&lt;p&gt;Most repo instructions skip this. They say "follow project conventions" and then bury the conventions across a README, package scripts, CI config, old PRs, and comments. An agent can search, but search isn't the same as orientation.&lt;/p&gt;

&lt;p&gt;Give it a first route through the repo.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9o7pzcjrf4gqqwd127oh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9o7pzcjrf4gqqwd127oh.png" alt="A boot flow for an agent: read AGENTS.md, load context, route by task type, act within boundaries, and update durable docs when needed." width="800" height="1182"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate identity from operating rules
&lt;/h2&gt;

&lt;p&gt;Your repo probably doesn't need a &lt;code&gt;SOUL.md&lt;/code&gt;, but the pattern is useful. One file can define working posture, while &lt;code&gt;AGENTS.md&lt;/code&gt; defines project behavior.&lt;/p&gt;

&lt;p&gt;For a software repo, that might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Working posture&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Read the existing code before proposing new abstractions.
&lt;span class="p"&gt;-&lt;/span&gt; Prefer local helpers over new dependencies.
&lt;span class="p"&gt;-&lt;/span&gt; Keep changes scoped to the user request.
&lt;span class="p"&gt;-&lt;/span&gt; Run the narrowest useful test first, then broaden if the change touches shared behavior.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those are judgment rules. They belong near the top because they shape every later decision.&lt;/p&gt;

&lt;p&gt;Then put repo-specific mechanics somewhere else:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Commands&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Install dependencies: &lt;span class="sb"&gt;`pnpm install`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Run unit tests: &lt;span class="sb"&gt;`pnpm test`&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; Run type checks: &lt;span class="sb"&gt;`pnpm typecheck`&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why split them? Because commands change faster than principles. If you mix everything together, the file turns into a junk drawer. Agents will still read it, but you won't know which instruction is steering behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put boundaries where the agent will trip over them
&lt;/h2&gt;

&lt;p&gt;The best line in my workspace &lt;code&gt;AGENTS.md&lt;/code&gt; is short: &lt;code&gt;trash &amp;gt; rm&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That teaches a local safety rule in three tokens. It says destructive deletion should be recoverable. It doesn't explain Unix philosophy. It doesn't lecture. It gives the agent a rule it can apply while acting.&lt;/p&gt;

&lt;p&gt;Your &lt;code&gt;AGENTS.md&lt;/code&gt; should include boundaries like that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Red lines&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Don't edit generated files directly.
&lt;span class="p"&gt;-&lt;/span&gt; Don't change public API behavior without updating tests.
&lt;span class="p"&gt;-&lt;/span&gt; Don't run migrations against shared databases.
&lt;span class="p"&gt;-&lt;/span&gt; Use &lt;span class="sb"&gt;`trash`&lt;/span&gt; instead of &lt;span class="sb"&gt;`rm`&lt;/span&gt; when deleting local files.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice the shape: concrete verbs, concrete objects, concrete limits.&lt;/p&gt;

&lt;p&gt;"Be careful with data" is too vague. "Don't run migrations against shared databases" gives the agent something it can obey.&lt;/p&gt;

&lt;h2&gt;
  
  
  Teach context access rules
&lt;/h2&gt;

&lt;p&gt;Agents often fail by reading too little or too much. Repo instructions can fix both.&lt;/p&gt;

&lt;p&gt;In my workspace, &lt;code&gt;MEMORY.md&lt;/code&gt; is only loaded in main sessions, not shared contexts. That's a privacy rule and a context rule at the same time. Daily notes are raw logs. Long-term memory is curated. &lt;code&gt;TOOLS.md&lt;/code&gt; is for environment-specific notes, while skills are reusable.&lt;/p&gt;

&lt;p&gt;That structure avoids a common problem: durable instructions become a dumping ground for every fact anyone might need someday.&lt;/p&gt;

&lt;p&gt;For a team repo, you can use the same split:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Context files&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="sb"&gt;`README.md`&lt;/span&gt;: human setup and project overview.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`AGENTS.md`&lt;/span&gt;: agent workflow and repo norms.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`docs/architecture.md`&lt;/span&gt;: current service boundaries.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`docs/runbooks/`&lt;/span&gt;: production procedures. Read only when the task touches operations.
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`.env.example`&lt;/span&gt;: allowed environment variable names. Never read real &lt;span class="sb"&gt;`.env`&lt;/span&gt; files unless asked.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last sentence matters. It tells the agent where the map ends.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep local details out of shared skills
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foebi39ptrnil8hpc1n9g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Foebi39ptrnil8hpc1n9g.png" alt="A decision tree for placing instructions in reusable skills, AGENTS.md, local tool notes, private memory, or temporary task notes." width="798" height="211"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;TOOLS.md&lt;/code&gt; in my workspace makes a clean distinction: skills define how tools work, and &lt;code&gt;TOOLS.md&lt;/code&gt; stores local specifics like camera names, SSH aliases, speakers, or preferred voices.&lt;/p&gt;

&lt;p&gt;That maps well to engineering teams.&lt;/p&gt;

&lt;p&gt;A reusable instruction might say:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;When debugging CI, inspect the failing job logs before changing code.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A local instruction might say:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;The staging dashboard is at &lt;span class="nt"&gt;&amp;lt;internal&lt;/span&gt; &lt;span class="na"&gt;URL&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those shouldn't live in the same place. Reusable instructions can move across projects. Local details shouldn't leak, and they age faster.&lt;/p&gt;

&lt;p&gt;This is one reason AGENTS.md fits naturally inside the AAIF project set. MCP describes how agents connect to tools. agentgateway works on routing and governing agent traffic. AGENTS.md handles repo-level behavior. You need all of those layers if agents are going to work across projects without each tool inventing its own private convention.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use skills for depth, not bulk
&lt;/h2&gt;

&lt;p&gt;My workspace says: "Skills provide your tools. When you need one, check its &lt;code&gt;SKILL.md&lt;/code&gt;."&lt;/p&gt;

&lt;p&gt;That's the right division of labor. &lt;code&gt;AGENTS.md&lt;/code&gt; should route the agent to deeper instructions. It shouldn't contain the full manual for every workflow.&lt;/p&gt;

&lt;p&gt;Bad:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Release process&lt;/span&gt;

[900 lines of release rules, changelog policy, package registry notes, rollback steps, comms templates, and edge cases]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Better:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Release process&lt;/span&gt;

For release work, read &lt;span class="sb"&gt;`skills/release/SKILL.md`&lt;/span&gt; before making changes. Do not publish packages or create GitHub releases unless the user explicitly asks.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why does this work? The root file stays readable, and the agent loads detail only when the task needs it.&lt;/p&gt;

&lt;p&gt;That matters more as context grows. An instruction file can hurt you if it forces every task to carry every workflow. A CSS fix doesn't need your incident response manual.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tell the agent when to speak and when to stay quiet
&lt;/h2&gt;

&lt;p&gt;This is an important part as well that shouldn't be ignored.&lt;/p&gt;

&lt;p&gt;The workspace &lt;code&gt;AGENTS.md&lt;/code&gt; has group chat rules. It tells the assistant to respond when directly mentioned, when it can add value, or when correcting meaningful misinformation. It also tells the assistant to stay quiet when the conversation is casual or already answered.&lt;/p&gt;

&lt;p&gt;That's repo-relevant too. Agents need communication norms.&lt;/p&gt;

&lt;p&gt;For a development repo, that might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## PR comments&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Comment when a change affects behavior users can observe.
&lt;span class="p"&gt;-&lt;/span&gt; Mention test gaps plainly.
&lt;span class="p"&gt;-&lt;/span&gt; Don't restate the diff.
&lt;span class="p"&gt;-&lt;/span&gt; Don't leave speculative security claims without a concrete path or file reference.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Agents generate a lot of text by default. Your instructions should define what useful text looks like in your project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make maintenance part of the contract
&lt;/h2&gt;

&lt;p&gt;The workspace instructions have a blunt rule: no "mental notes." If something should persist, write it to a file.&lt;/p&gt;

&lt;p&gt;That belongs in more repos.&lt;/p&gt;

&lt;p&gt;Agents learn project facts during a task: a flaky test command, a generated directory that shouldn't be edited, a local setup wrinkle, a service boundary that wasn't documented. If the agent only uses that knowledge once, the next run pays the same discovery cost.&lt;/p&gt;

&lt;p&gt;Add a maintenance rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Updating these instructions&lt;/span&gt;

When you learn a durable repo rule, update &lt;span class="sb"&gt;`AGENTS.md`&lt;/span&gt; or the relevant doc in the same PR. Keep task-specific notes out of this file.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then enforce the second sentence. Otherwise AGENTS.md becomes a chat transcript with headings.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical structure
&lt;/h2&gt;

&lt;p&gt;If I were starting a repo-level &lt;code&gt;AGENTS.md&lt;/code&gt; today, I'd use this shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# AGENTS.md&lt;/span&gt;

&lt;span class="gu"&gt;## Start here&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Read this file before making changes.
&lt;span class="p"&gt;-&lt;/span&gt; Read &lt;span class="sb"&gt;`README.md`&lt;/span&gt; for setup.
&lt;span class="p"&gt;-&lt;/span&gt; Read the nearest package-level &lt;span class="sb"&gt;`AGENTS.md`&lt;/span&gt; if one exists.

&lt;span class="gu"&gt;## Working posture&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Preserve existing patterns unless the task calls for changing them.
&lt;span class="p"&gt;-&lt;/span&gt; Keep edits scoped.
&lt;span class="p"&gt;-&lt;/span&gt; Prefer small tests close to the changed code.

&lt;span class="gu"&gt;## Commands&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Install:
&lt;span class="p"&gt;-&lt;/span&gt; Test:
&lt;span class="p"&gt;-&lt;/span&gt; Typecheck:
&lt;span class="p"&gt;-&lt;/span&gt; Lint:

&lt;span class="gu"&gt;## Repo map&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; &lt;span class="sb"&gt;`apps/web`&lt;/span&gt;: frontend
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`packages/api`&lt;/span&gt;: API client
&lt;span class="p"&gt;-&lt;/span&gt; &lt;span class="sb"&gt;`packages/db`&lt;/span&gt;: schema and migrations

&lt;span class="gu"&gt;## Boundaries&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Don't edit generated files.
&lt;span class="p"&gt;-&lt;/span&gt; Don't run destructive database commands.
&lt;span class="p"&gt;-&lt;/span&gt; Ask before publishing, emailing, posting, or deploying.

&lt;span class="gu"&gt;## Workflow routing&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; For releases, read &lt;span class="sb"&gt;`docs/release.md`&lt;/span&gt;.
&lt;span class="p"&gt;-&lt;/span&gt; For security changes, read &lt;span class="sb"&gt;`docs/security.md`&lt;/span&gt;.
&lt;span class="p"&gt;-&lt;/span&gt; For UI changes, inspect existing components first.

&lt;span class="gu"&gt;## Maintenance&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Add durable lessons here.
&lt;span class="p"&gt;-&lt;/span&gt; Remove stale instructions when the code changes.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's enough for day one. It creates the framework to build upon as you continue to iterate.&lt;/p&gt;

&lt;p&gt;The goal isn't to make the agent know everything. The goal is to make the first move reasonabe, the dangerous moves constrained, and the next file obvious.&lt;/p&gt;

&lt;p&gt;In an open agentic ecosystem, the shared convention doesn't need to be heavy to be useful. It needs to be predictable enough that any agent can arrive in your repo and know where to begin: &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;

</description>
      <category>openclaw</category>
      <category>agents</category>
      <category>ai</category>
    </item>
    <item>
      <title>From First-Run Drop-Off to First Useful Agent Run</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Thu, 16 Jul 2026 22:16:16 +0000</pubDate>
      <link>https://dev.to/bengreenberg/from-first-run-drop-off-to-first-useful-agent-run-mde</link>
      <guid>https://dev.to/bengreenberg/from-first-run-drop-off-to-first-useful-agent-run-mde</guid>
      <description>&lt;p&gt;I keep coming back to the same onboarding question: what happens in the first 10 minutes?&lt;/p&gt;

&lt;p&gt;For agent tools, that window is brutal. A developer opens a repo, starts the agent, asks for a change, and waits to see if the tool understands the project. If the agent guesses the package manager, misses the test path, edits generated files, or asks the developer to explain the repo from scratch, trust drops fast.&lt;/p&gt;

&lt;p&gt;That isn't an agent model problem every time. A lot of it is repo readiness.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://aaif.io" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt;, hosted by the Linux Foundation, is building an open home for projects like MCP, goose, AGENTS.md, and agentgateway. That work can sound big and infrastructural, but one of the most useful entry points is small: make your repo easier for an agent to understand on the first run.&lt;/p&gt;

&lt;p&gt;AGENTS.md is the repo-side context. goose is a practical runtime path. Together, they give you a way to move from "the agent is poking around" to "the agent made a useful first pass."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkxw71wmpdfdinbwzonrg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkxw71wmpdfdinbwzonrg.png" alt="A structure diagram showing AGENTS.md as repo context and goose as the runtime path that uses it for a first-run task." width="800" height="239"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Start With The First Useful Run
&lt;/h2&gt;

&lt;p&gt;Don't begin by asking, "What should our agent docs say?"&lt;/p&gt;

&lt;p&gt;Ask this instead: what should a developer be able to ask an agent to do in this repo within 10 minutes?&lt;/p&gt;

&lt;p&gt;Pick one task. Not the whole system. One useful first run.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Find the right entry point for a small bug&lt;/li&gt;
&lt;li&gt;Add a focused test around an existing function&lt;/li&gt;
&lt;li&gt;Update a docs page with a known source file nearby&lt;/li&gt;
&lt;li&gt;Explain how a specific package or module is wired&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That first run gives your AGENTS.md a job. It isn't a policy dump. It's the context an agent needs to avoid wasting the developer's first session.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put Repo Truth Where Agents Can Find It
&lt;/h2&gt;

&lt;p&gt;AGENTS.md is a simple open format for guiding coding agents, and the project site says it's already used by over 60k open-source projects: &lt;a href="https://agents.md" rel="noopener noreferrer"&gt;https://agents.md&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The reason it works is plain: agents need a predictable place for repo instructions. README files are written for humans. CI files are written for automation. AGENTS.md gives agents the details that usually live in maintainer heads.&lt;/p&gt;

&lt;p&gt;Your first version should answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What kind of project is this?&lt;/li&gt;
&lt;li&gt;Where does source code live?&lt;/li&gt;
&lt;li&gt;Where do tests live?&lt;/li&gt;
&lt;li&gt;Which files should agents avoid editing?&lt;/li&gt;
&lt;li&gt;What style or architecture choices should agents preserve?&lt;/li&gt;
&lt;li&gt;What should the agent do before claiming a task is done?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep it short enough that someone would maintain it. Stale agent instructions are worse than missing ones because they create confident mistakes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write Instructions Like Maintainer Notes
&lt;/h2&gt;

&lt;p&gt;An AGENTS.md file doesn't need brand language. It needs maintainer notes.&lt;/p&gt;

&lt;p&gt;Say things like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# AGENTS.md&lt;/span&gt;

&lt;span class="gu"&gt;## Project Shape&lt;/span&gt;

This repo contains a web app and supporting packages. App code lives in &lt;span class="sb"&gt;`apps/web`&lt;/span&gt;. Shared code lives in &lt;span class="sb"&gt;`packages`&lt;/span&gt;.

&lt;span class="gu"&gt;## Working Rules&lt;/span&gt;

Prefer small changes that match nearby patterns. Do not rewrite public APIs unless the task asks for it.

&lt;span class="gu"&gt;## Tests&lt;/span&gt;

When changing behavior, add or update the closest existing test. If you can't run the test locally, say what you inspected and why the test wasn't run.

&lt;span class="gu"&gt;## Files To Avoid&lt;/span&gt;

Do not edit generated files, lockfiles, or vendored code unless the task is specifically about dependency updates.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what's missing: fake certainty.&lt;/p&gt;

&lt;p&gt;Don't say "run the full test suite" unless that's realistic. Don't list commands you haven't checked. Don't tell the agent to use a package manager you don't use. Your agent instructions should be as true as your README.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design The Goose Path
&lt;/h2&gt;

&lt;p&gt;goose is an open-source AI agent runtime under AAIF. Its project page describes it as an agent that can install, execute, edit, and test with any LLM: &lt;a href="https://aaif.io/projects/goose" rel="noopener noreferrer"&gt;https://aaif.io/projects/goose&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7jcxs97gifz95dtndps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm7jcxs97gifz95dtndps.png" alt="An open source, extensible AI agent that goes beyond code suggestions. Install, execute, edit, and test with any LLM." width="800" height="226"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For onboarding, think of goose as the first-run path you can test against your repo instructions.&lt;/p&gt;

&lt;p&gt;A good first-run path has three pieces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A clear starting task&lt;/li&gt;
&lt;li&gt;A repo-level AGENTS.md&lt;/li&gt;
&lt;li&gt;A visible stopping point&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The stopping point matters. If the agent changes code, how does the developer know whether it did the right thing? Maybe the agent should point to the files it changed. Maybe it should explain the test it would run. Maybe it should stop before touching a migration, generated file, or public API.&lt;/p&gt;

&lt;p&gt;That belongs in AGENTS.md.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make The Agent Ask Better Questions
&lt;/h2&gt;

&lt;p&gt;A useful agent doesn't need to know everything. It needs to know when to stop guessing.&lt;/p&gt;

&lt;p&gt;Add guidance for uncertainty:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## When Unsure&lt;/span&gt;

If the requested change touches auth, billing, data deletion, or production configuration, ask before editing.

If there are multiple plausible implementations, describe the tradeoff and choose the smallest local change unless the user tells you otherwise.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why does that help? Because first-run drop-off often comes from surprise. The agent edits the wrong layer, takes a broad refactor path, or treats a risky area like ordinary code.&lt;/p&gt;

&lt;p&gt;Good instructions narrow the blast radius.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat Docs As Product Surface
&lt;/h2&gt;

&lt;p&gt;Developer onboarding isn't separate from product. The docs shape what users try, where they get stuck, and whether they come back.&lt;/p&gt;

&lt;p&gt;For agent-ready repos, AGENTS.md is part of that product surface. So review it the same way you'd review a quickstart:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the first task obvious?&lt;/li&gt;
&lt;li&gt;Are repo boundaries named?&lt;/li&gt;
&lt;li&gt;Are setup assumptions current?&lt;/li&gt;
&lt;li&gt;Are risky areas called out?&lt;/li&gt;
&lt;li&gt;Can a new contributor tell what "done" means?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where AAIF's open ecosystem angle becomes practical. If agent tools are going to work across projects, maintainers need shared conventions that don't depend on one vendor, one editor, or one model. AGENTS.md gives repos a portable instruction layer. goose gives developers an open way to run agent workflows against it.&lt;/p&gt;

&lt;p&gt;Small file. Real leverage.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Practical Checklist
&lt;/h2&gt;

&lt;p&gt;Use this before you point an agent at your repo:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filo2d3ixxv4x1wxobamu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Filo2d3ixxv4x1wxobamu.png" alt="An iteration loop for improving AGENTS.md by running a first task, observing wrong guesses, and editing the instructions." width="800" height="86"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick one first-run task a new developer would value.&lt;/li&gt;
&lt;li&gt;Add or update AGENTS.md with project shape, test expectations, and files to avoid.&lt;/li&gt;
&lt;li&gt;Remove commands you haven't verified.&lt;/li&gt;
&lt;li&gt;Tell the agent how to behave around risky code paths.&lt;/li&gt;
&lt;li&gt;Run the first task through goose or your agent runtime of choice.&lt;/li&gt;
&lt;li&gt;Edit AGENTS.md based on where the agent guessed wrong.&lt;/li&gt;
&lt;li&gt;Repeat until the first run produces something you would review seriously.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The goal isn't to make the agent perfect. The goal is to make the first session legible.&lt;/p&gt;

&lt;p&gt;A developer should be able to open the repo, start the agent, ask for one scoped task, and understand the result without becoming the repo tour guide.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agentskills</category>
      <category>tutorial</category>
      <category>architecture</category>
    </item>
    <item>
      <title>I built my first Robinhood Chain app as an index basket</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Sun, 12 Jul 2026 15:51:24 +0000</pubDate>
      <link>https://dev.to/arbitrum/i-built-my-first-robinhood-chain-app-as-an-index-basket-20li</link>
      <guid>https://dev.to/arbitrum/i-built-my-first-robinhood-chain-app-as-an-index-basket-20li</guid>
      <description>&lt;p&gt;I built a small index basket app on Robinhood Chain because I wanted to understand the developer path from the first contract deploy all the way to a working frontend.&lt;/p&gt;

&lt;p&gt;The app is intentionally plain: a user deposits Stock Tokens, which are blockchain tokens that represent real equity exposure, and receives an ERC-20 basket share. ERC-20 is Ethereum's standard token interface, so a compatible token exposes familiar methods like &lt;code&gt;balanceOf&lt;/code&gt;, &lt;code&gt;transfer&lt;/code&gt;, and &lt;code&gt;approve&lt;/code&gt;. The basket share is priced from live price feeds, and the user can redeem it back into the underlying Stock Tokens.&lt;/p&gt;

&lt;p&gt;That's the part that made this interesting to me. The chain is custom, but the app path is not. I still wrote Solidity, deployed with Foundry, read contract state with viem, and wrote transactions from React with wagmi.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm1c8mbe3qvudjo15x4gi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm1c8mbe3qvudjo15x4gi.png" alt="A user connects a wallet to the React frontend, which uses wagmi and viem to call Robinhood Chain contracts that interact with Stock Tokens and Chainlink feeds." width="600" height="76"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've built normal web apps, think of the chain's RPC endpoint as the API base URL. A wallet is login plus a signing key. A smart contract is backend code you deploy to the chain, except you should treat it like immutable infrastructure because you don't get to hot-patch it casually later.&lt;/p&gt;

&lt;p&gt;The demo and source are here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;App: &lt;a href="https://robinhood-chain-dapp.vercel.app/" rel="noopener noreferrer"&gt;https://robinhood-chain-dapp.vercel.app/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Code: &lt;a href="https://github.com/hummusonrails/robinhood-chain-dapp-example" rel="noopener noreferrer"&gt;https://github.com/hummusonrails/robinhood-chain-dapp-example&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The custom chain still feels like the EVM
&lt;/h2&gt;

&lt;p&gt;Robinhood Chain is a custom Arbitrum Chain, which means it runs as a dedicated chain on the stack of Arbitrum, an Ethereum scaling system. It is also EVM-compatible. EVM means Ethereum Virtual Machine, the runtime that executes Solidity contracts, so the tooling surface looks like the Ethereum developer flow many tutorials already teach.&lt;/p&gt;

&lt;p&gt;An L2, or rollup, is a chain that executes transactions separately and then posts compressed proof or transaction data back to Ethereum. Robinhood Chain uses Ethereum blobs for data availability, which is a cheaper Ethereum data lane for rollups to publish the data needed to reconstruct chain state. Gas, the metered compute fee you pay to run transactions, is paid in ETH.&lt;/p&gt;

&lt;p&gt;The first deploy looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;PRIVATE_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0x&amp;lt;your_private_key&amp;gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;RH_RPC_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://rpc.testnet.chain.robinhood.com

forge create src/MyContract.sol:MyContract &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--rpc-url&lt;/span&gt; &lt;span class="nv"&gt;$RH_RPC_URL&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--private-key&lt;/span&gt; &lt;span class="nv"&gt;$PRIVATE_KEY&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--broadcast&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Foundry is the contract build, test, and deploy CLI. &lt;code&gt;forge create&lt;/code&gt; compiles the contract, sends the deployment transaction, and broadcasts it to the RPC endpoint.&lt;/p&gt;

&lt;p&gt;This is the part I appreciate as a developer. Robinhood gets its own chain configuration, infrastructure, pricing, and product controls. I still get the contract model I know how to reason about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stock Tokens are ERC-20s with display accounting
&lt;/h2&gt;

&lt;p&gt;Stock Tokens are ERC-20s that represent real market assets. That means contracts can read balances, request approvals, and transfer them using the same functions they would use for any other ERC-20.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;IERC20 nvda = IERC20(NVDA_TOKEN_ADDRESS);

uint256 balance = nvda.balanceOf(user);
nvda.approve(spender, amount);
nvda.transfer(recipient, amount);
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The wrinkle is corporate actions. Stocks split. Dividends happen. The economic relationship between one token and one underlying share can change.&lt;/p&gt;

&lt;p&gt;Stock Tokens implement ERC-8056, the Scaled UI Amount extension. Raw token balances stay stable for contracts. A UI multiplier is how wallets and apps display the share-equivalent amount.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F97bik6ktzeo8cc3thtrs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F97bik6ktzeo8cc3thtrs.png" alt="Raw token balances stay stable for contract accounting while corporate actions update a UI multiplier used for displayed share-equivalent balances." width="600" height="112"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;interface IScaledUIAmount {
    function uiMultiplier() external view returns (uint256);
    function balanceOfUI(address account) external view returns (uint256);
    function totalSupplyUI() external view returns (uint256);
    function newUIMultiplier() external view returns (uint256);
    function effectiveAt() external view returns (uint256);
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The display conversion is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;underlyingShares = rawBalance * uiMultiplier / 1e18;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why not just mutate balances after a split?&lt;/p&gt;

&lt;p&gt;Because contracts depend on stable accounting. If my basket contract holds a raw token balance, I don't want a display-level corporate action to unexpectedly rewrite the reserve math inside the contract. The UI can show share-equivalent amounts, and the protocol can keep using raw ERC-20 units.&lt;/p&gt;

&lt;h2&gt;
  
  
  The basket contract does only a few things
&lt;/h2&gt;

&lt;p&gt;The sample app has two contracts and no owner.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Contract&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Control surface&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;BasketFactory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Deploys and records baskets&lt;/td&gt;
&lt;td&gt;Permissionless&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;BasketToken&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Holds components, mints, redeems, prices shares&lt;/td&gt;
&lt;td&gt;No owner or upgrade path&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The factory does not custody user funds. It deploys a basket and records the address.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;function createBasket(
    string calldata name,
    string calldata symbol,
    BasketToken.Component[] calldata components,
    uint256 maxPriceAge
) external returns (address basket) {
    basket = address(new BasketToken(name, symbol, components, maxPriceAge));
    _baskets.push(basket);
    isBasket[basket] = true;
    emit BasketCreated(basket, msg.sender, name, symbol);
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each basket stores a fixed list of components. A component is a token, a price feed, and the amount of that token backing one basket share.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;struct Component {
    address token;
    address feed;
    uint256 unitsPerShare;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A price feed is an external data source a contract or app can read. In this app, the feeds come from Chainlink, an oracle network that publishes market data onchain. An oracle is the bridge between offchain facts, like a stock price, and onchain code.&lt;/p&gt;

&lt;p&gt;On mainnet, the production chain where real assets move, the demo basket uses TSLA, NVDA, and AAPL Stock Tokens with their Chainlink feeds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;function _mainnetComponents()
    internal
    pure
    returns (BasketToken.Component[] memory c)
{
    c = new BasketToken.Component[](3);
    c[0] = BasketToken.Component(MAINNET_TSLA, MAINNET_TSLA_FEED, 0.4e18);
    c[1] = BasketToken.Component(MAINNET_NVDA, MAINNET_NVDA_FEED, 0.3e18);
    c[2] = BasketToken.Component(MAINNET_AAPL, MAINNET_AAPL_FEED, 0.3e18);
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One &lt;code&gt;TRIO&lt;/code&gt; share is backed by &lt;code&gt;0.4 TSLA&lt;/code&gt;, &lt;code&gt;0.3 NVDA&lt;/code&gt;, and &lt;code&gt;0.3 AAPL&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On testnet, which is a staging chain with no real funds at risk, the demo uses faucet Stock Tokens for TSLA, AMZN, and NFLX with mock feeds. A faucet is a service that gives you test tokens so you can build without spending real money.&lt;/p&gt;

&lt;h2&gt;
  
  
  Minting follows the ERC-20 approval pattern
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff9pnghj64spuyu4lhuaf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff9pnghj64spuyu4lhuaf.png" alt="Raw token balances stay stable for contract accounting while corporate actions update a UI multiplier used for displayed share-equivalent balances." width="528" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;From the frontend, minting is two steps: approve each component token, then call the basket.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;writeContract&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;address&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;stockTokenAddress&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;abi&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;erc20Abi&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;functionName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;approve&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;basketAddress&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;writeContract&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;address&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;basketAddress&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;abi&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;basketTokenAbi&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;functionName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;mint&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;args&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;shares&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;account&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;writeContract&lt;/code&gt; call is from wagmi, a React library for wallet connections and contract writes. viem is the TypeScript Ethereum client underneath it for typed reads, writes, and transaction handling.&lt;/p&gt;

&lt;p&gt;Onchain, the basket pulls the required component amounts and mints shares in the same transaction.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;function mint(uint256 shares, address to) external nonReentrant {
    if (shares == 0) revert ZeroShares();

    uint256 count = _components.length;
    for (uint256 i = 0; i &amp;lt; count; i++) {
        Component memory c = _components[i];
        uint256 amount = Math.mulDiv(
            c.unitsPerShare,
            shares,
            SHARE_UNIT,
            Math.Rounding.Ceil
        );

        IERC20(c.token).safeTransferFrom(msg.sender, address(this), amount);
    }

    _mint(to, shares);
    emit Minted(msg.sender, to, shares);
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The rounding direction matters. Mint rounds up so a user cannot underpay the basket reserves by tiny decimal leftovers.&lt;/p&gt;

&lt;p&gt;Redeem is the mirror image. Burn first, transfer components out, and round down so the reserves cannot be overdrawn.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;function redeem(uint256 shares, address to) external nonReentrant {
    if (shares == 0) revert ZeroShares();
    if (to == address(0)) revert ZeroAddress();

    _burn(msg.sender, shares);

    uint256 count = _components.length;
    for (uint256 i = 0; i &amp;lt; count; i++) {
        Component memory c = _components[i];
        uint256 amount = Math.mulDiv(
            c.unitsPerShare,
            shares,
            SHARE_UNIT,
            Math.Rounding.Floor
        );

        IERC20(c.token).safeTransfer(to, amount);
    }

    emit Redeemed(msg.sender, to, shares);
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Redeem skips the price feed and the factory. It burns shares and returns the component tokens the contract already holds.&lt;/p&gt;

&lt;p&gt;I keep pricing and redemption apart on purpose. You use the price for the UI. Redeem returns the collateral.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local to mainnet is the path I want rehearsed
&lt;/h2&gt;

&lt;p&gt;The repo gives you the whole loop.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone &lt;span class="nt"&gt;--recurse-submodules&lt;/span&gt; https://github.com/hummusonrails/robinhood-chain-dapp-example.git
&lt;span class="nb"&gt;cd &lt;/span&gt;robinhood-chain-dapp-example

pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;anvil
pnpm run deploy:local
pnpm run smoke
pnpm run dev
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;anvil&lt;/code&gt; is Foundry's local development chain. Think of it as a throwaway local server for contracts. The local deploy creates mock Stock Tokens, mock Chainlink feeds, the factory, and a demo &lt;code&gt;Tech Trio&lt;/code&gt; basket. It also writes &lt;code&gt;apps/frontend/.env.local&lt;/code&gt;, so the frontend knows which addresses to call.&lt;/p&gt;

&lt;p&gt;For Robinhood Chain testnet:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;PRIVATE_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nv"&gt;$YOUR_TESTNET_KEY&lt;/span&gt; pnpm run deploy:testnet
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One Foundry script handles local, testnet, and mainnet by checking the &lt;code&gt;chainid&lt;/code&gt;, which is the chain's network identifier.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;function run() external {
    vm.startBroadcast();

    BasketFactory factory = new BasketFactory();
    console2.log("FACTORY=%s", address(factory));

    BasketToken.Component[] memory components;
    if (block.chainid == 4663) {
        components = _mainnetComponents();
    } else if (block.chainid == 46630) {
        components = _testnetComponents();
    } else {
        components = _localComponents();
    }

    address basket =
        factory.createBasket("Tech Trio", "TRIO", components, MAX_PRICE_AGE);

    console2.log("DEMO_BASKET=%s", basket);
    console2.log("CHAIN_ID=%s", block.chainid);

    vm.stopBroadcast();
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The testnet deployment behind the live walkthrough is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Contract&lt;/th&gt;
&lt;th&gt;Address&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;BasketFactory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0xC1940D5fd58ce735A44a53f910852B12250F6a14&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;BasketToken&lt;/code&gt; (&lt;code&gt;TRIO&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0x7633e0920Ea46A8Ec54F61C95adECD391c01Edd4&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Before spending mainnet gas, I want fork tests. A fork test runs tests against a local copy of live chain state, so you can check integration assumptions without sending real transactions.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pnpm run &lt;span class="nb"&gt;test&lt;/span&gt;:contracts
pnpm run &lt;span class="nb"&gt;test&lt;/span&gt;:fork
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is the shape of the fork test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;function test_mintAndRedeem_withRealStockTokens() public {
    deal(TSLA, alice, 1e18);
    deal(NVDA, alice, 1e18);
    deal(AAPL, alice, 1e18);

    vm.startPrank(alice);
    IERC20Metadata(TSLA).approve(address(basket), type(uint256).max);
    IERC20Metadata(NVDA).approve(address(basket), type(uint256).max);
    IERC20Metadata(AAPL).approve(address(basket), type(uint256).max);

    basket.mint(2e18, alice);
    assertEq(basket.balanceOf(alice), 2e18);
    assertEq(IERC20Metadata(TSLA).balanceOf(address(basket)), 0.8e18);

    basket.redeem(2e18, alice);
    assertEq(basket.totalSupply(), 0);
    assertEq(IERC20Metadata(TSLA).balanceOf(alice), 1e18);
    vm.stopPrank();
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the development loop I want for this kind of app: mocks for speed, testnet for wallet flow, fork tests for live integration assumptions, and mainnet only after the path is rehearsed.&lt;/p&gt;

&lt;p&gt;A block explorer, which is basically hosted request logs for a chain, then gives you a way to inspect deployed contracts and transactions. The testnet contracts are verified on Blockscout, so you can read the source and calls after deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stock Tokens change the app assumptions
&lt;/h2&gt;

&lt;p&gt;Stock Tokens still fit the ERC-20 interface, but the surrounding assumptions are different from a generic token.&lt;/p&gt;

&lt;p&gt;For user balances, don't blindly show &lt;code&gt;balanceOf&lt;/code&gt;. Use &lt;code&gt;balanceOfUI&lt;/code&gt; or apply &lt;code&gt;uiMultiplier&lt;/code&gt; so the user sees the share-equivalent amount.&lt;/p&gt;

&lt;p&gt;For prices, read the per-token Chainlink feed on mainnet. For corporate actions, track multiplier updates and pending effective times. For valuation, remember that the feed price already includes the multiplier.&lt;/p&gt;

&lt;p&gt;Stock market hours matter too. Crypto feeds may update around the clock. Stock feeds follow market sessions, so stale data checks need to reflect that.&lt;/p&gt;

&lt;p&gt;The exit path is the one I care about most. If a user wants to redeem their basket share, I don't want that flow blocked because an oracle read is stale. The contract already holds the component tokens. Redemption should return the collateral.&lt;/p&gt;

&lt;h2&gt;
  
  
  Learn the chain later; start with the app
&lt;/h2&gt;

&lt;p&gt;Robinhood Chain runs on Arbitrum Nitro, the same underlying technology as Arbitrum One, deployed as a dedicated chain. Arbitrum One is the public shared L2. A custom Arbitrum Chain gives a team its own execution environment while keeping the Ethereum-style contract model.&lt;/p&gt;

&lt;p&gt;The mechanics under the hood are also why the fees are small. Transactions hit a sequencer, which is the service that orders transactions for the rollup, land in fast blocks, get batched, and settle back to Ethereum using blob data. The fee combines L2 execution gas with the data cost on L1, which is Ethereum itself.&lt;/p&gt;

&lt;p&gt;That's useful context, but I wouldn't start by trying to absorb the whole chain architecture.&lt;/p&gt;

&lt;p&gt;Start with the working system. Read a Stock Token balance. Approve a token. Mint a share. Redeem it. Check the price feed. Run the fork test. Look at the transaction in a block explorer.&lt;/p&gt;

&lt;p&gt;Plenty of apps can live on Robinhood Chain. This basket is a good first build because it touches the surfaces most apps using market assets will need: token reads, approvals, ERC-8056 display logic, Chainlink feeds, local mocks, testnet deployment, verified contracts, fork tests, and a Next.js frontend using wagmi and viem.&lt;/p&gt;

&lt;p&gt;What I learned from building it is where the real work sits: deciding where accounting belongs, where pricing belongs, and which assumptions deserve a test before real users and real assets touch the contract. The app stack itself is familiar.&lt;/p&gt;

</description>
      <category>web3</category>
      <category>solidity</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>The MCP Release Candidate Survival Guide: Apps, Auth, Deprecations, and Tool Schemas</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Thu, 02 Jul 2026 10:27:58 +0000</pubDate>
      <link>https://dev.to/bengreenberg/the-mcp-release-candidate-survival-guide-apps-auth-deprecations-and-tool-schemas-5da2</link>
      <guid>https://dev.to/bengreenberg/the-mcp-release-candidate-survival-guide-apps-auth-deprecations-and-tool-schemas-5da2</guid>
      <description>&lt;p&gt;The &lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/" rel="noopener noreferrer"&gt;MCP &lt;code&gt;2026-07-28&lt;/code&gt; release candidate&lt;/a&gt; is the largest major specification revision since MCP launched. It is also a compatibility test for everyone building clients, servers, SDKs, gateways, and developer tools around the protocol.&lt;/p&gt;

&lt;p&gt;The release candidate was locked on May 21, 2026. The final specification is scheduled for July 28, 2026. This window is the time to test real implementations and find migration pain points.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Check your transport assumptions
&lt;/h2&gt;

&lt;p&gt;The biggest change is that MCP is now stateless at the protocol layer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feowi3elz87y74aelwnps.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feowi3elz87y74aelwnps.png" alt="Stateless topology (Source: https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/)" width="710" height="340"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If your current Streamable HTTP implementation depends on &lt;code&gt;initialize&lt;/code&gt;, &lt;code&gt;initialized&lt;/code&gt;, or &lt;code&gt;Mcp-Session-Id&lt;/code&gt;, you have migration work. In the release candidate, each request carries the protocol version, client info, and capabilities in &lt;code&gt;_meta&lt;/code&gt;. The new &lt;code&gt;server/discover&lt;/code&gt; method covers cases where a client needs server capabilities up front.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;tools/call&lt;/code&gt; request over Streamable HTTP now includes headers such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That changes how infrastructure can handle MCP traffic. A gateway no longer needs to inspect JSON bodies just to route or rate-limit common operations. A load balancer can send requests to any server instance because the protocol no longer assumes a sticky session.&lt;/p&gt;

&lt;p&gt;The compatibility question is simple: does your server still hide required state in the connection?&lt;/p&gt;

&lt;p&gt;If yes, move that state into an explicit application handle. For example, a tool can return a &lt;code&gt;basket_id&lt;/code&gt;, &lt;code&gt;browser_id&lt;/code&gt;, or job handle, and the model can pass it back as a normal tool argument later. That makes state visible to the model and portable across server instances.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Test server-to-client request flows
&lt;/h2&gt;

&lt;p&gt;Stateless MCP still needs interaction during a call. The release candidate changes how that works.&lt;/p&gt;

&lt;p&gt;Server-initiated requests can only happen while the server is processing a client request. For elicitation, roots, or sampling flows, the server returns an &lt;code&gt;InputRequiredResult&lt;/code&gt;, and the client retries the original call with &lt;code&gt;inputResponses&lt;/code&gt; and &lt;code&gt;requestState&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That means clients need to preserve and replay the right data. Servers need to treat the retry as a continuation, even if it lands on another instance.&lt;/p&gt;

&lt;p&gt;A good test case is a destructive tool call that asks for confirmation. The server should return an input request, the client should collect the answer, and the retry should succeed without relying on connection memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Treat MCP Apps as real app surfaces
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/#mcp-apps-server-rendered-user-interfaces" rel="noopener noreferrer"&gt;MCP Apps&lt;/a&gt; let servers provide interactive HTML interfaces that hosts render in sandboxed iframes.&lt;/p&gt;

&lt;p&gt;This is a big developer-experience change, but it also has security and product implications. Tools can declare UI templates ahead of time, which lets hosts prefetch, cache, and review them before anything runs. The UI still talks back through MCP’s JSON-RPC protocol, so UI-driven actions go through the same consent path as tool calls.&lt;/p&gt;

&lt;p&gt;If you maintain a host, test your iframe isolation, permission prompts, and caching behavior. If you maintain a server, check that your UI template declarations are deterministic and do not depend on hidden session state.&lt;/p&gt;

&lt;p&gt;The Apps model will reward consistent discipline: explicit templates, clear tool boundaries, and no surprise network behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Harden authorization now
&lt;/h2&gt;

&lt;p&gt;The release candidate tightens MCP authorization around OAuth 2.0 and OpenID Connect deployments.&lt;/p&gt;

&lt;p&gt;Clients now need to validate the &lt;code&gt;iss&lt;/code&gt; parameter on authorization responses under RFC 9207. Authorization servers should begin sending &lt;code&gt;iss&lt;/code&gt; now because future clients are expected to reject responses without it.&lt;/p&gt;

&lt;p&gt;Dynamic Client Registration also changes. Clients declare OpenID Connect &lt;code&gt;application_type&lt;/code&gt;, which matters for desktop and CLI clients using localhost redirect URIs. Clients also bind registered credentials to the issuing authorization server’s &lt;code&gt;issuer&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The migration checklist here is direct:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Verify &lt;code&gt;iss&lt;/code&gt; handling in clients.&lt;/li&gt;
&lt;li&gt;Confirm authorization servers send &lt;code&gt;iss&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Check Dynamic Client Registration metadata.&lt;/li&gt;
&lt;li&gt;Re-register when a resource moves between authorization servers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is especially relevant for MCP because one client may connect to many servers. Mix-up risks are not theoretical in that shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Stop building new work on deprecated core features
&lt;/h2&gt;

&lt;p&gt;Roots, Sampling, and Logging are deprecated in the release candidate.&lt;/p&gt;

&lt;p&gt;They still work. The deprecation is annotation-only for this release, and the methods, types, and capability flags continue to work in every spec version published within a year of it. Removal would require a separate SEP.&lt;/p&gt;

&lt;p&gt;Still, new work should move elsewhere:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Roots should move to tool parameters, resource URIs, or server configuration.&lt;/li&gt;
&lt;li&gt;Sampling should move to direct integration with model provider APIs.&lt;/li&gt;
&lt;li&gt;Logging should move to &lt;code&gt;stderr&lt;/code&gt; for stdio or OpenTelemetry for structured observability.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you maintain SDK abstractions, this is the moment to add warnings without breaking users. If you maintain docs, stop teaching deprecated features as the default path.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Validate tool schemas against JSON Schema 2020-12
&lt;/h2&gt;

&lt;p&gt;Tool &lt;code&gt;inputSchema&lt;/code&gt; and &lt;code&gt;outputSchema&lt;/code&gt; now use full JSON Schema 2020-12.&lt;/p&gt;

&lt;p&gt;For input schemas, the root still has to be &lt;code&gt;type: "object"&lt;/code&gt;, but schemas can now use &lt;code&gt;oneOf&lt;/code&gt;, &lt;code&gt;anyOf&lt;/code&gt;, &lt;code&gt;allOf&lt;/code&gt;, conditionals, &lt;code&gt;$ref&lt;/code&gt;, and &lt;code&gt;$defs&lt;/code&gt;. Output schemas are unrestricted. &lt;code&gt;structuredContent&lt;/code&gt; can be any JSON value instead of only an object.&lt;/p&gt;

&lt;p&gt;That creates opportunity and risk.&lt;/p&gt;

&lt;p&gt;Servers should bound schema depth and validation time. Implementations should not auto-dereference external &lt;code&gt;$ref&lt;/code&gt; URIs. Clients that made assumptions about simple object-only schemas need tests against composed schemas.&lt;/p&gt;

&lt;p&gt;Also check error handling. The missing resource error changes from MCP’s custom &lt;code&gt;-32002&lt;/code&gt; to the JSON-RPC standard &lt;code&gt;-32602&lt;/code&gt; Invalid Params. If your client matches on the literal code, update it.&lt;/p&gt;

&lt;p&gt;As you work through the checklist, if you find any issues or major friction points bring them to the community. You can open an issue in the &lt;a href="https://github.com/modelcontextprotocol/modelcontextprotocol/issues" rel="noopener noreferrer"&gt;specification repository&lt;/a&gt;. For implementation questions, the relevant &lt;a href="https://modelcontextprotocol.io/community/working-interest-groups" rel="noopener noreferrer"&gt;Working Group&lt;/a&gt; channel in the &lt;a href="https://modelcontextprotocol.io/community/communication#discord" rel="noopener noreferrer"&gt;contributor Discord&lt;/a&gt; is the fastest path to an answer.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>opensource</category>
      <category>news</category>
    </item>
    <item>
      <title>MCP Server or CLI: A Decision Rubric for Developer Tooling</title>
      <dc:creator>Ben Greenberg</dc:creator>
      <pubDate>Wed, 01 Jul 2026 14:29:03 +0000</pubDate>
      <link>https://dev.to/bengreenberg/mcp-server-or-cli-a-decision-rubric-for-developer-tooling-2ch6</link>
      <guid>https://dev.to/bengreenberg/mcp-server-or-cli-a-decision-rubric-for-developer-tooling-2ch6</guid>
      <description>&lt;p&gt;Teams are rushing to make their internal tools available to agents. That is good. It is also where a lot of design mistakes begin.&lt;/p&gt;

&lt;p&gt;The question usually shows up like this:&lt;/p&gt;

&lt;p&gt;Should we expose this as an MCP server, or should the agent just use our CLI?&lt;/p&gt;

&lt;p&gt;That framing makes it sound like one option is more “agentic” than the other. I do not think that is the useful distinction. A CLI and an MCP server solve different problems. The better question is: where does this capability naturally live, and what contract does the agent need in order to use it well?&lt;/p&gt;

&lt;p&gt;MCP, now hosted by the &lt;a href="https://aaif.io" rel="noopener noreferrer"&gt;Agentic AI Foundation&lt;/a&gt; under the Linux Foundation, gives the open agentic AI ecosystem a shared protocol for connecting agents to tools, data, and applications. That shared protocol matters because teams should not have to rebuild the same integration patterns for every agent runtime. But MCP is not a reason to wrap every executable in a server. Sometimes a CLI is exactly the right interface. Sometimes an MCP server is. Often, the answer is both, with different responsibilities.&lt;/p&gt;

&lt;p&gt;Here is the rubric I use.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy3oiqpbqmdo9utn0add6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy3oiqpbqmdo9utn0add6.png" alt="A decision flow showing when to choose a CLI, an MCP server, both, or documentation only." width="200" height="300"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the workflow
&lt;/h2&gt;

&lt;p&gt;A CLI is strongest when the workflow already belongs to a human developer.&lt;/p&gt;

&lt;p&gt;If the task is repo-local, terminal-native, and already part of how developers build, test, debug, or ship software, start with the CLI. Developers know how to inspect it. CI can run it. Logs are usually visible. Failures can be reproduced outside the agent. The same interface works for humans, scripts, and automation.&lt;/p&gt;

&lt;p&gt;Good CLI-shaped examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Running a project-specific code generator&lt;/li&gt;
&lt;li&gt;Applying a migration in a local development environment&lt;/li&gt;
&lt;li&gt;Linting, formatting, testing, or packaging&lt;/li&gt;
&lt;li&gt;Inspecting repo state&lt;/li&gt;
&lt;li&gt;Scaffolding files inside a checked-out project&lt;/li&gt;
&lt;li&gt;Running one-off diagnostics where stdout and exit codes are enough&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An MCP server is strongest when the workflow does not naturally belong in a terminal session, or when the agent needs a structured, discoverable capability instead of a command string.&lt;/p&gt;

&lt;p&gt;Good MCP-shaped examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reading from or writing to a SaaS API&lt;/li&gt;
&lt;li&gt;Searching a private knowledge base&lt;/li&gt;
&lt;li&gt;Fetching typed records from an internal system&lt;/li&gt;
&lt;li&gt;Performing actions that need scoped authorization&lt;/li&gt;
&lt;li&gt;Exposing capabilities across multiple agent clients&lt;/li&gt;
&lt;li&gt;Providing context as resources, not just command output&lt;/li&gt;
&lt;li&gt;Giving the agent a constrained set of tool calls instead of broad shell access&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trap is assuming that “agent can run command” and “agent has a good tool interface” are the same thing. They are not.&lt;/p&gt;

&lt;p&gt;A CLI gives an agent a way to execute. MCP gives an agent a way to understand what capabilities exist, what inputs they accept, and what kind of result comes back.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8zupnsmjjyaqy67nr7d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx8zupnsmjjyaqy67nr7d.png" alt="A comparison of the responsibilities that belong to CLI surfaces versus MCP tool surfaces." width="600" height="220"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The five-question rubric
&lt;/h2&gt;

&lt;p&gt;When deciding between a CLI, an MCP server, or both, I like to ask five questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Is this primarily an interactive human workflow?
&lt;/h3&gt;

&lt;p&gt;If yes, prefer a CLI.&lt;/p&gt;

&lt;p&gt;Developers still need tools that work when no agent is involved. If a human would reasonably run the tool while sitting inside a repo, reading logs, adjusting flags, and retrying, a CLI is usually the right primary interface.&lt;/p&gt;

&lt;p&gt;That does not mean agents cannot use it. Agents are quite good at driving existing developer workflows when the commands are documented and the outputs are predictable. This is where &lt;a href="https://aaif.io/projects/agents-md/" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt; fits naturally: document which commands are safe, how to run tests, what directories are off-limits, and what failure modes are expected.&lt;/p&gt;

&lt;p&gt;The CLI should be boring in the best way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Clear subcommands&lt;/li&gt;
&lt;li&gt;Stable flags&lt;/li&gt;
&lt;li&gt;Machine-readable output where useful&lt;/li&gt;
&lt;li&gt;Non-zero exit codes on failure&lt;/li&gt;
&lt;li&gt;Dry-run modes for risky actions&lt;/li&gt;
&lt;li&gt;Good help text&lt;/li&gt;
&lt;li&gt;No hidden interactive prompts in automation paths&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the tool needs a human to make judgment calls mid-run, keep that interaction in the CLI. Do not hide it behind an MCP tool and pretend the workflow became autonomous.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Does the agent need to discover the capability?
&lt;/h3&gt;

&lt;p&gt;If yes, lean toward MCP.&lt;/p&gt;

&lt;p&gt;One of the real advantages of MCP is that tools can be described to the agent as tools. The agent does not need to infer everything from a README, shell history, or tribal knowledge. It can see available tool names, descriptions, schemas, and expected inputs.&lt;/p&gt;

&lt;p&gt;That matters when the capability is meant to be reused across agents or across teams.&lt;/p&gt;

&lt;p&gt;A CLI can be documented well, but discovery is still indirect. The agent needs to know the command exists, know where it is installed, know how to call it, and know how to interpret the output. MCP makes the capability part of the agent’s tool surface.&lt;/p&gt;

&lt;p&gt;Use MCP when the agent should be able to answer questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What tools are available to me?&lt;/li&gt;
&lt;li&gt;What arguments does this action require?&lt;/li&gt;
&lt;li&gt;What resources can I inspect?&lt;/li&gt;
&lt;li&gt;What shape will the result have?&lt;/li&gt;
&lt;li&gt;What actions are allowed in this environment?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is especially useful for APIs and internal systems where a raw CLI would either expose too much or force the agent to learn a human-oriented interface.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Where is the auth boundary?
&lt;/h3&gt;

&lt;p&gt;If the action crosses an authorization boundary, consider MCP carefully.&lt;/p&gt;

&lt;p&gt;A CLI often inherits the developer’s local environment: shell credentials, config files, tokens, SSH agents, cloud profiles. That can be fine for local workflows. It can also be too broad for agent access.&lt;/p&gt;

&lt;p&gt;MCP gives teams a cleaner place to define permission boundaries. The server can expose only the operations the agent should have. It can scope credentials server-side. It can validate inputs before touching the underlying system. It can log tool calls in a way that is easier to review than arbitrary shell execution.&lt;/p&gt;

&lt;p&gt;This does not make MCP magically safe. A poorly designed MCP server can still be dangerous. But the server boundary gives you a place to enforce policy.&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Should the agent inherit the user’s full shell environment?&lt;/li&gt;
&lt;li&gt;Should this action use delegated or scoped credentials?&lt;/li&gt;
&lt;li&gt;Do we need per-tool authorization?&lt;/li&gt;
&lt;li&gt;Do we need audit logs of agent actions?&lt;/li&gt;
&lt;li&gt;Do we need to prevent arbitrary command composition?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the answer to those questions is yes, a CLI alone may be too blunt.&lt;/p&gt;

&lt;p&gt;For production-facing systems, this is also where infrastructure projects like &lt;a href="https://aaif.io/projects/agentgateway/" rel="noopener noreferrer"&gt;agentgateway&lt;/a&gt; become relevant. Once agent traffic spans MCP servers, APIs, models, and services, teams need consistent policy, routing, and observability. That is a different layer than the individual tool decision, but the design choices connect.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Is the operation stateful?
&lt;/h3&gt;

&lt;p&gt;State changes raise the bar.&lt;/p&gt;

&lt;p&gt;A read-only diagnostic command is one thing. A tool that creates tickets, deploys services, updates customer data, rotates secrets, or changes infrastructure is another.&lt;/p&gt;

&lt;p&gt;For state-changing actions, the interface should make the action hard to misuse. That can be done in a CLI, an MCP server, or both. The question is which interface gives you the better control surface.&lt;/p&gt;

&lt;p&gt;A CLI might be right when the state change belongs in a developer-controlled workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Apply this migration to my local database&lt;/li&gt;
&lt;li&gt;Generate this file in my branch&lt;/li&gt;
&lt;li&gt;Create a release artifact after tests pass&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An MCP server might be right when the state change touches an external system:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Create an incident&lt;/li&gt;
&lt;li&gt;Update a CRM record&lt;/li&gt;
&lt;li&gt;Open a pull request with a structured payload&lt;/li&gt;
&lt;li&gt;Provision access for a user&lt;/li&gt;
&lt;li&gt;Trigger a workflow in a deployment platform&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For stateful MCP tools, I would avoid generic verbs like &lt;code&gt;run&lt;/code&gt;, &lt;code&gt;execute&lt;/code&gt;, or &lt;code&gt;update&lt;/code&gt; when the action can be modeled more specifically. The tool should say what it does. The input schema should constrain what can happen. The response should include enough structured data for the agent to verify the result.&lt;/p&gt;

&lt;p&gt;For stateful CLIs, I want dry-run support, confirmation controls that can be disabled only in explicit automation modes, and output that makes it clear what changed.&lt;/p&gt;

&lt;p&gt;The shared principle is the same: the agent should not be guessing.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. What is the maintenance cost?
&lt;/h3&gt;

&lt;p&gt;Every new interface becomes a product surface.&lt;/p&gt;

&lt;p&gt;A CLI needs packaging, versioning, docs, help text, examples, and compatibility guarantees. An MCP server needs all of that plus server lifecycle, transport decisions, schema design, client compatibility, authentication, deployment, monitoring, and operational ownership.&lt;/p&gt;

&lt;p&gt;That cost may be worth it. But it should buy something real.&lt;/p&gt;

&lt;p&gt;An MCP wrapper around a CLI can be useful when it adds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Better tool descriptions&lt;/li&gt;
&lt;li&gt;Safer input validation&lt;/li&gt;
&lt;li&gt;Structured outputs&lt;/li&gt;
&lt;li&gt;Scoped permissions&lt;/li&gt;
&lt;li&gt;Shared access across agent clients&lt;/li&gt;
&lt;li&gt;A stable abstraction over a messy underlying command&lt;/li&gt;
&lt;li&gt;Resource access the CLI does not provide well&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An MCP wrapper is probably not worth it when it only shells out to an existing command and returns the same text output the agent would have seen anyway.&lt;/p&gt;

&lt;p&gt;That does not mean “never wrap a CLI.” It means the wrapper should create leverage. If the MCP server is only a thinner, less debuggable path to the same command, keep the CLI and document it well.&lt;/p&gt;

&lt;h2&gt;
  
  
  The “both” pattern
&lt;/h2&gt;

&lt;p&gt;Many teams should build both, but not as duplicates.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F65tjx4yfo9fhfnfx03xh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F65tjx4yfo9fhfnfx03xh.png" alt="Humans and CI use the CLI, agents use an MCP server, and both call shared core logic while MCP adds policy and schemas." width="600" height="260"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A good pattern is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The CLI remains the human and CI interface.&lt;/li&gt;
&lt;li&gt;The MCP server exposes selected capabilities for agents.&lt;/li&gt;
&lt;li&gt;Shared core logic lives below both interfaces.&lt;/li&gt;
&lt;li&gt;The MCP server does not become a dumping ground for every CLI command.&lt;/li&gt;
&lt;li&gt;The CLI does not become an escape hatch for unsafe agent actions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, imagine an internal deployment tool.&lt;/p&gt;

&lt;p&gt;The CLI might support the full developer workflow: build, validate, preview, deploy, rollback, inspect logs, and run local checks. It assumes the user is a developer with repo access and deployment permissions.&lt;/p&gt;

&lt;p&gt;The MCP server might expose narrower tools: get deployment status, list services, create a preview environment, request a rollback plan, or fetch logs for a specific service. Those tools can have tighter schemas and safer defaults. They can also return structured data that an agent can reason over without parsing terminal output.&lt;/p&gt;

&lt;p&gt;Both interfaces can call the same underlying deployment library. They do not need to expose the same surface area.&lt;/p&gt;

&lt;p&gt;That separation is healthy. Humans need power tools. Agents need constrained capabilities with clear contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical decision guide
&lt;/h2&gt;

&lt;p&gt;Use a CLI when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The task is repo-local or terminal-native&lt;/li&gt;
&lt;li&gt;Humans need to run it directly&lt;/li&gt;
&lt;li&gt;CI should run the same interface&lt;/li&gt;
&lt;li&gt;Shell composition is a feature&lt;/li&gt;
&lt;li&gt;Text output and exit codes are enough&lt;/li&gt;
&lt;li&gt;The auth model is already appropriate for the local developer context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use an MCP server when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The agent needs discoverable tools or resources&lt;/li&gt;
&lt;li&gt;The task touches an external system&lt;/li&gt;
&lt;li&gt;Inputs and outputs should be typed&lt;/li&gt;
&lt;li&gt;Permissions need to be scoped&lt;/li&gt;
&lt;li&gt;Multiple agent clients should share the same integration&lt;/li&gt;
&lt;li&gt;The tool should hide implementation details&lt;/li&gt;
&lt;li&gt;Auditability and policy matter&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use both when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Humans and agents both need the capability&lt;/li&gt;
&lt;li&gt;The CLI is already valuable&lt;/li&gt;
&lt;li&gt;The MCP server can expose a safer or more structured subset&lt;/li&gt;
&lt;li&gt;Shared core logic can prevent drift&lt;/li&gt;
&lt;li&gt;The agent interface should be stable even if the CLI evolves&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do neither, at least for now, when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The workflow is not understood yet&lt;/li&gt;
&lt;li&gt;The tool would expose broad credentials without guardrails&lt;/li&gt;
&lt;li&gt;The “agent use case” is just novelty&lt;/li&gt;
&lt;li&gt;A README update would solve the immediate problem&lt;/li&gt;
&lt;li&gt;The maintenance owner is unclear&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What this means for the open agentic AI ecosystem
&lt;/h2&gt;

&lt;p&gt;The open agentic AI ecosystem needs standards like MCP. It also needs restraint.&lt;/p&gt;

&lt;p&gt;If every team turns every script into an MCP server, agents get a larger tool list but not necessarily better tools. Tool overload is real. Poor descriptions, loose schemas, unsafe side effects, and noisy outputs make agents worse, not better.&lt;/p&gt;

&lt;p&gt;The goal is not to make everything agent-accessible. The goal is to make the right capabilities available through the right contract.&lt;/p&gt;

&lt;p&gt;That is why this decision matters so much at this particular moment. MCP gives builders a common way to expose tools and context. AGENTS.md gives projects a common place to tell coding agents how to work inside a repo. agentgateway points toward the operational layer teams need when agent traffic becomes production traffic.&lt;/p&gt;

&lt;p&gt;These projects are stronger when we use them for the problems they actually solve.&lt;/p&gt;

&lt;p&gt;A CLI is not “less agentic” because it runs in a terminal. An MCP server is not “better architecture” because it speaks a protocol. The useful line is simpler:&lt;/p&gt;

&lt;p&gt;If the work belongs in the developer workflow, start with a CLI.&lt;/p&gt;

&lt;p&gt;If the agent needs a structured, discoverable, permissioned capability, build an MCP server.&lt;/p&gt;

&lt;p&gt;If both humans and agents need it, design both surfaces intentionally and keep the shared logic underneath.&lt;/p&gt;

&lt;p&gt;That is the rubric. Not CLI versus MCP. CLI where the workflow lives. MCP where the capability needs a contract.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>cli</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
