<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Roy Berris</title>
    <description>The latest articles on DEV Community by Roy Berris (@royberris).</description>
    <link>https://dev.to/royberris</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4177064%2F9f62e53c-5769-49b9-91b2-61fba16a0713.png</url>
      <title>DEV Community: Roy Berris</title>
      <link>https://dev.to/royberris</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/royberris"/>
    <language>en</language>
    <item>
      <title>The Three-Layer System for Consistent AI Specifications</title>
      <dc:creator>Roy Berris</dc:creator>
      <pubDate>Sun, 11 Oct 2026 13:37:00 +0000</pubDate>
      <link>https://dev.to/royberris/the-three-layer-system-for-consistent-ai-specifications-1b19</link>
      <guid>https://dev.to/royberris/the-three-layer-system-for-consistent-ai-specifications-1b19</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; We treat the specification as the single source of truth for user intent and business rules, compiling functional design down through three strict layers. Business rules stay in native prose, the domain model structures invariants in English, and the contract or wire schema is simply the final compiled layer. Splitting the process this way stops hallucinations and eliminates schema drift.&lt;/p&gt;

&lt;p&gt;When we first asked an AI model to generate a functional specification directly from rough requirements, the result looked convincing on the surface. But when we inspected the logic, the problems jumped out. The model quietly dropped business validation rules, invented defaults out of thin air, and hallucinated system behavior. Trying to fix all of that by stuffing more instructions into one giant prompt just confused the model and made the output drift even further.&lt;/p&gt;

&lt;p&gt;People often assume specification generation is just about producing API definitions or code stubs. In our team, we look at it differently. A specification is the functional design of your system, and it is the single source of truth for user intent and business rules. Today's language models cannot bridge human intent and rigid technical contracts in one jump. To solve this, we treat functional design like a compiler chain: three distinct, strictly one-way layers that compile human truth down into a verifiable contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Single-Shot Specification Lead to Drift?
&lt;/h2&gt;

&lt;p&gt;Most teams start by feeding user stories or meeting notes into a chat prompt and asking for an interface definition or code stubs immediately. Under the surface, the model tries to solve three completely different problems at the same time: understanding business policies, modeling domain relationships, and formatting technical contract syntax. In that single pass, it silently invents enum values, misinterprets domain invariants, and overlooks edge cases just to produce syntactically valid output.&lt;/p&gt;

&lt;p&gt;This creates silent drift between what business stakeholders expect and what engineers actually build. The trouble usually shows up late, often when validation fails in staging. When that happens, the temptation is to patch the generated schema or contract by hand.&lt;/p&gt;

&lt;p&gt;But the contract is only the final compiled artifact of your functional design. Hand-editing downstream contracts breaks the chain of truth. The next time you run an AI generation tool, it overwrites those manual fixes or diverges further because the source of truth was never updated.&lt;/p&gt;

&lt;p&gt;Splitting this workflow into sequential passes introduces a small latency trade-off. Generating intermediate representations takes a few minutes instead of a few seconds. In our experience, that extra time is worth it. You keep the context clean at each step and save yourself painful debugging sessions later.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does the Three-Layer Compiler Pipeline Work?
&lt;/h2&gt;

&lt;p&gt;Instead of treating specification generation as a single prompt, we treat it like a compiler chain. Each stage has a single, clear job, moving strictly from human and business intent down to machine contracts. Upstream artifacts remain the single source of truth, and downstream layers are generated automatically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBMMVsiTGF5ZXIgMTogSHVtYW4gYW5kIEJ1c2luZXNzIFRydXRoPGJyLz5OYXRpdmUgYnVzaW5lc3MgcnVsZXMgYW5kIHVzZXIgaW50ZW50Il0gLS0-IEwyWyJMYXllciAyOiBGdW5jdGlvbmFsIERlc2lnbjxici8-RG9tYWluIG1vZGVsIGFuZCBzeXN0ZW0gaW52YXJpYW50cyJdCiAgICBMMiAtLT4gTDNbIkxheWVyIDM6IFZlcmlmaWFibGUgQ29udHJhY3Q8YnIvPkNvbXBpbGVkIHdpcmUgc2NoZW1hcyBhbmQgaW50ZXJmYWNlcyJdCgogICAgc3R5bGUgTDEgZmlsbDojMTQ0MzJhLHN0cm9rZTojNGFkZTgwLGNvbG9yOiNkY2ZjZTcKICAgIHN0eWxlIEwyIGZpbGw6IzBlM2E0YSxzdHJva2U6IzY3ZThmOSxjb2xvcjojZTBmN2ZmCiAgICBzdHlsZSBMMyBmaWxsOiMzYjFmNWMsc3Ryb2tlOiNjMDg0ZmMsY29sb3I6I2YzZThmZg%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBMMVsiTGF5ZXIgMTogSHVtYW4gYW5kIEJ1c2luZXNzIFRydXRoPGJyLz5OYXRpdmUgYnVzaW5lc3MgcnVsZXMgYW5kIHVzZXIgaW50ZW50Il0gLS0-IEwyWyJMYXllciAyOiBGdW5jdGlvbmFsIERlc2lnbjxici8-RG9tYWluIG1vZGVsIGFuZCBzeXN0ZW0gaW52YXJpYW50cyJdCiAgICBMMiAtLT4gTDNbIkxheWVyIDM6IFZlcmlmaWFibGUgQ29udHJhY3Q8YnIvPkNvbXBpbGVkIHdpcmUgc2NoZW1hcyBhbmQgaW50ZXJmYWNlcyJdCgogICAgc3R5bGUgTDEgZmlsbDojMTQ0MzJhLHN0cm9rZTojNGFkZTgwLGNvbG9yOiNkY2ZjZTcKICAgIHN0eWxlIEwyIGZpbGw6IzBlM2E0YSxzdHJva2U6IzY3ZThmOSxjb2xvcjojZTBmN2ZmCiAgICBzdHlsZSBMMyBmaWxsOiMzYjFmNWMsc3Ryb2tlOiNjMDg0ZmMsY29sb3I6I2YzZThmZg%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" alt="Diagram" width="276" height="446"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The system organizes requirements into three distinct stages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Layer 1 (Human and Business Truth):&lt;/strong&gt; Written in native business prose to capture stakeholder rules directly. In our projects in the Netherlands, capturing policies in Dutch prevents premature translation errors and preserves regulatory nuances that non-technical domain experts care about. Stakeholders can read, verify, and own this layer directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 2 (Functional Design and Domain Model):&lt;/strong&gt; Expressed as an English architectural specification following &lt;a href="https://www.domainlanguage.com/ddd/" rel="noopener noreferrer"&gt;Domain-Driven Design&lt;/a&gt; principles. This layer normalizes business concepts into formal entities, &lt;a href="https://berris.dev/nodes/using-value-objects-in-net/" rel="noopener noreferrer"&gt;value objects&lt;/a&gt;, lifecycle states, and domain invariants without any transport or serialization baggage. It defines what the system does and why, independent of delivery protocols.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 3 (Verifiable Contract):&lt;/strong&gt; Defined in &lt;a href="https://typespec.io/" rel="noopener noreferrer"&gt;TypeSpec&lt;/a&gt;, OpenAPI, or schema definitions. This layer handles transport mechanics, status codes, query parameters, header definitions, and serialization, following the principles we described in &lt;a href="https://berris.dev/nodes/designing-apis-for-ai-agents/" rel="noopener noreferrer"&gt;Designing APIs for AI Agents&lt;/a&gt;. Wire contracts and API schemas are not the starting point. They are the final compiled layer of functional design.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This workflow is an automated compiler pipeline rather than a traditional waterfall process. We edit the upstream business rules, run the generation tool, and let the pipeline compile the downstream contracts. Nobody edits the compiled wire contract by hand.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do We Capture and Validate Business Rules?
&lt;/h2&gt;

&lt;p&gt;You cannot expect an AI model or an engineer sitting alone to invent business truth. Layer 1 requires collaborative discovery with domain experts before any code or prompt runs. In our projects, we use three discovery techniques to draw out rules from stakeholders:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://www.eventstorming.com/" rel="noopener noreferrer"&gt;Event Storming&lt;/a&gt;:&lt;/strong&gt; We gather domain experts and developers in a room to map domain events along a business timeline. We explore what happens across a process, what triggers each action, and which policies govern state changes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://cucumber.io/blog/bdd/example-mapping-introduction/" rel="noopener noreferrer"&gt;Example Mapping&lt;/a&gt;:&lt;/strong&gt; We take each user story and break it down into concrete business rules illustrated by realistic examples. Talking through concrete scenarios reveals edge cases and hidden assumptions that abstract bullet points conceal.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stakeholder interviews:&lt;/strong&gt; We talk directly with product managers, operational staff, and compliance officers. Capturing their exact words in their native language preserves legal and operational nuances that get lost when developers translate requirements straight into technical jargon.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Capturing fragments of rules is only half the battle. You also need to validate the requirements as a coherent whole before feeding them to an AI compiler.&lt;/p&gt;

&lt;p&gt;We run validation sessions with the business to check completeness. We walk through the end-to-end user journey to make sure every failure mode, boundary condition, and lifecycle transition has an explicit rule. Walking through realistic examples surfaces edge cases early, when changing a rule costs nothing.&lt;/p&gt;

&lt;p&gt;Once the rules are complete and unambiguous, we secure formal sign-off from business stakeholders. Because Layer 1 uses plain language without HTTP status codes, JSON fields, or database tables, non-technical experts can read every sentence and take full ownership. This sign-off freezes the baseline for Layer 1. Only after the business approves Layer 1 does the AI compiler pipeline turn those rules into downstream models and contracts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tracing an Example Through the Three Layers: A Library Book Loan
&lt;/h2&gt;

&lt;p&gt;To see how this works in practice, consider a classic scenario: a member borrowing a physical book from a library. Walking this feature through the three layers shows how business intent compiles into a technical contract while maintaining complete traceability.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 1: Business Rules
&lt;/h3&gt;

&lt;p&gt;Layer 1 captures the intent and rules in plain language approved by the library staff:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;User Intent:&lt;/strong&gt; A registered library member wants to borrow a physical book.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rule 1 (Loan limit):&lt;/strong&gt; A member can have at most 5 active book loans at any time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rule 2 (Good standing):&lt;/strong&gt; A member cannot borrow books if they have an unpaid fine.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rule 3 (Loan period):&lt;/strong&gt; The standard loan period is 21 days from the date of checkout.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;There are no endpoints, status codes, or database keys here. Any librarian can read this list and confirm whether it accurately describes how the library operates.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 2: Domain Model
&lt;/h3&gt;

&lt;p&gt;Next, the compiler pipeline generates the English architectural specification. Layer 2 applies domain-driven design concepts to model entities, invariants, post-conditions, and domain events:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Entity:&lt;/strong&gt; &lt;code&gt;Loan&lt;/code&gt; (attributes: &lt;code&gt;loanId&lt;/code&gt;, &lt;code&gt;memberId&lt;/code&gt;, &lt;code&gt;bookId&lt;/code&gt;, &lt;code&gt;loanDate&lt;/code&gt;, &lt;code&gt;dueDate&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Domain Invariant (&lt;code&gt;BorrowBookPolicy&lt;/code&gt;):&lt;/strong&gt; A loan can only be created if the member's current active loans count is less than 5, and the member's unpaid fine balance is zero.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Post-conditions:&lt;/strong&gt; A new &lt;code&gt;Loan&lt;/code&gt; is instantiated with &lt;code&gt;status: Active&lt;/code&gt; and &lt;code&gt;dueDate&lt;/code&gt; calculated as &lt;code&gt;loanDate + 21 days&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Domain Event:&lt;/strong&gt; The system emits a &lt;code&gt;BookBorrowed&lt;/code&gt; event containing &lt;code&gt;loanId&lt;/code&gt;, &lt;code&gt;memberId&lt;/code&gt;, &lt;code&gt;bookId&lt;/code&gt;, and &lt;code&gt;dueDate&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Layer 2 defines what the domain logic must enforce. It still contains zero HTTP headers or serialization details.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 3: Verifiable Contract
&lt;/h3&gt;

&lt;p&gt;Finally, the pipeline compiles Layer 2 into a verifiable API contract. Here is the OpenAPI definition for the borrow operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="s"&gt;/members/{memberId}/loans&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;post&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Borrow a book&lt;/span&gt;
      &lt;span class="na"&gt;operationId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;borrowBook&lt;/span&gt;
      &lt;span class="na"&gt;security&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;bearerAuth&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;loans:borrow"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;memberId&lt;/span&gt;
          &lt;span class="na"&gt;in&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;path&lt;/span&gt;
          &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
          &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
      &lt;span class="na"&gt;requestBody&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
        &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
              &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;bookId&lt;/span&gt;
              &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;bookId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;201"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Book borrowed successfully&lt;/span&gt;
          &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;$ref&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;#/components/schemas/Loan"&lt;/span&gt;
      &lt;span class="err"&gt;  &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;409"&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Business policy violation&lt;/span&gt;
          &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
                &lt;span class="na"&gt;required&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;error&lt;/span&gt;
                &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                    &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
                    &lt;span class="na"&gt;enum&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;loan-limit-reached&lt;/span&gt;
                      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;unpaid-fine&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice how this contract structures the response codes. The schema returns a 201 status code with the created loan on success, or a 409 Conflict with a machine-readable error when a domain rule fails.&lt;/p&gt;

&lt;h3&gt;
  
  
  Emphasizing Traceability
&lt;/h3&gt;

&lt;p&gt;Look closely at how every line in this Layer 3 contract traces directly back through Layer 2 to Layer 1:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The path &lt;code&gt;/members/{memberId}/loans&lt;/code&gt; and operation summary &lt;code&gt;Borrow a book&lt;/code&gt; trace directly to the plain-language user intent in Layer 1.&lt;/li&gt;
&lt;li&gt;The required permission &lt;code&gt;loans:borrow&lt;/code&gt; traces to the library policy that the caller must have borrowing privileges.&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;409&lt;/code&gt; error enum &lt;code&gt;loan-limit-reached&lt;/code&gt; traces through the Layer 2 &lt;code&gt;BorrowBookPolicy&lt;/code&gt; invariant back to Rule 1 (limit of 5 active loans).&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;409&lt;/code&gt; error enum &lt;code&gt;unpaid-fine&lt;/code&gt; traces through the Layer 2 invariant back to Rule 2 (no unpaid fines allowed).&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;dueDate&lt;/code&gt; in the returned &lt;code&gt;Loan&lt;/code&gt; object traces through the Layer 2 post-condition back to Rule 3 (21-day loan duration).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nothing in the contract appears out of thin air. If a model generates an unexpected error code or an extraneous query parameter, you can flag it instantly because it lacks an upstream parent in Layer 1. Every line in the contract traces directly back to an approved business rule.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Is This Three-Layer Scaffolding Temporary?
&lt;/h2&gt;

&lt;p&gt;We want to be clear about why this system exists today. These three separate stages are practical scaffolding built around the limits of current language models. Right now, models cannot reliably preserve nuance across multiple levels of abstraction in a single inference pass.&lt;/p&gt;

&lt;p&gt;We think of this three-layer pipeline as a deliberate engineering hedge. It stabilizes model outputs today without locking our architecture into rigid, permanent workflow machinery. We documented this architectural boundary choice in an &lt;a href="https://adr.github.io/" rel="noopener noreferrer"&gt;Architecture Decision Record (ADR)&lt;/a&gt; so our team understands why we enforce these boundaries and under what conditions we can simplify them.&lt;/p&gt;

&lt;p&gt;Eventually, models will be capable enough to jump from business discussions to validated wire schemas without intermediate steps. When that shift happens, explicit file-based handoffs between layers will disappear from our daily work. Even then, the underlying separation of concerns between business truth, functional design, and verifiable contracts will remain structurally sound.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does the Pipeline Stay Consistent When Models Change?
&lt;/h2&gt;

&lt;p&gt;Model behavior changes with every new release, but this pipeline keeps our specifications stable. Layer 1 functions as source code, while Layer 2 and Layer 3 act as compiled build artifacts. When business requirements shift, we update the business rules in Layer 1 and trigger a clean compilation rather than patching downstream files.&lt;/p&gt;

&lt;p&gt;In our team, we store domain rules, conventions, and architectural constraints inside version-controlled repository instructions and skills. Keeping guidance in git repositories ensures every engineer and continuous integration agent runs the exact same prompts. Ad-hoc chat sessions lose context quickly, but versioned skills keep that knowledge in the repository where everyone can use it.&lt;/p&gt;

&lt;p&gt;The architecture remains completely tool-agnostic. You can switch the underlying foundation model or migrate from TypeSpec to another interface definition language whenever you choose. Because your core functional design lives upstream in clean domain models, changing a code generator never forces a rewrite of your business rules. When tools improve, downstream layers are simply regenerated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendations for Your Team
&lt;/h2&gt;

&lt;p&gt;Setting up this pipeline requires discipline around layer boundaries and prompt management. If you want to set up this system in your own projects, here is our practical advice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Adopt a spec-first mindset:&lt;/strong&gt; Treat functional design as the single source of truth for user intent and business rules, not code stubs or handwritten schemas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat Layer 1 as the sole source of truth:&lt;/strong&gt; Update business rules in native prose and secure business sign-off before compiling downstream layers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ensure strict traceability:&lt;/strong&gt; Make sure every property, operation, and error status in Layer 3 is derived from and traceable to an approved Layer 1 rule.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;View the wire contract as the compiled layer:&lt;/strong&gt; Treat APIs and schemas as the final compiled representation of functional design, not the starting point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ground Layer 2 in domain-driven design:&lt;/strong&gt; Express domain invariants, entities, and events in clean English before worrying about transport protocols.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Regenerate downstream layers when tools improve:&lt;/strong&gt; Treat Layer 2 and Layer 3 as build artifacts. When you upgrade models or linters, recompile from Layer 1 instead of hand-patching files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automate downstream compilation:&lt;/strong&gt; Run models in strict one-way passes using automated scripts or continuous integration tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Version control your prompt instructions:&lt;/strong&gt; Commit architectural rules, ADRs, and skills to your git repository alongside the project code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate contracts with deterministic tools:&lt;/strong&gt; Use standard TypeSpec compilers or schema linters to catch syntax errors and contract flaws immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never edit generated layers by hand:&lt;/strong&gt; If you hand-tweak Layer 2 or Layer 3 outputs, subsequent compilations will wipe out your modifications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Avoid two-way sync:&lt;/strong&gt; Never attempt to back-propagate changes from a wire contract back into the domain model. Keep it strictly one-way.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep transport details out of business rules:&lt;/strong&gt; Serialization quirks, status codes, and HTTP headers do not belong in Layer 1 or Layer 2.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;A specification is not just an API contract. It is the functional design that captures user intent and business rules as your single source of truth. By treating functional design as a three-layer compiler pipeline, you stop hallucinations and keep your contracts consistent with what the business actually needs.&lt;/p&gt;

&lt;p&gt;Give the model one job at a time, and let the compiler do the rest.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is this three-layer system only for APIs?
&lt;/h3&gt;

&lt;p&gt;No. While the final layer often produces API schemas or interface definitions, the pipeline is about generating the functional design as a whole. The specification is the single source of truth for user intent and business rules, and the contract is simply the final compiled layer of that functional design.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you capture business rules before compiling?
&lt;/h3&gt;

&lt;p&gt;We capture business rules through collaborative workshops with domain experts, using Event Storming to discover events, Example Mapping to pin down rules and edge cases, and stakeholder interviews. We validate the rules as a coherent whole and secure formal business sign-off before running the AI compiler pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why write Layer 1 in native language instead of English?
&lt;/h3&gt;

&lt;p&gt;Writing Layer 1 in native business prose, such as Dutch in our projects in the Netherlands, captures policies directly from stakeholders without premature translation. This preserves regulatory nuances and policy details that domain experts care about, before Layer 2 translates and normalizes them into English domain concepts.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens when business requirements change or tools improve?
&lt;/h3&gt;

&lt;p&gt;You update the approved business rules in Layer 1 and recompile downstream layers through the automated pipeline. Because Layer 2 and Layer 3 are compiled build artifacts, you never hand-patch downstream files, and you can regenerate your entire contract whenever your foundation models or toolchain improve.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://berris.dev/nodes/the-three-layer-system-for-consistent-ai-specifications/" rel="noopener noreferrer"&gt;berris.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>softwareengineering</category>
      <category>documentation</category>
    </item>
    <item>
      <title>Start Treating AI Like a Human Colleague</title>
      <dc:creator>Roy Berris</dc:creator>
      <pubDate>Sun, 11 Oct 2026 13:36:47 +0000</pubDate>
      <link>https://dev.to/royberris/start-treating-ai-like-a-human-colleague-3dio</link>
      <guid>https://dev.to/royberris/start-treating-ai-like-a-human-colleague-3dio</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; The biggest mistake with AI in software engineering is component-level micromanagement: pre-designing every interface, class, and table yourself before handing off the pieces. When you treat AI like a capable engineering colleague, you define the desired business outcome and domain boundaries, then let it architect and build the whole system. Your job shifts from sketching components to steering outcomes, evaluating trade-offs, and mentoring.&lt;/p&gt;

&lt;p&gt;I'll be honest with you: as a software architect, my default instinct has always been to break everything down. For years, I treated system design like a puzzle where my job was to shape every single piece. When AI coding tools arrived, I brought that exact habit with me.&lt;/p&gt;

&lt;p&gt;I would spend hours mapping out every database table, defining every service interface, specifying every Data Transfer Object (DTO), and deciding which method called which helper. Once I had pre-chewed every component, I handed the pieces to the AI one by one, like an architect handing brick specifications to a bricklayer.&lt;/p&gt;

&lt;p&gt;It gave me a sense of control, but it was exhausting. I was doing ninety percent of the thinking myself, drawing boxes and dictating signatures. I was acting as the architect, the product manager, and the micromanager, using AI as a fast typist to fill in the blanks I had already solved. I was the biggest bottleneck on the project.&lt;/p&gt;

&lt;p&gt;The turning point came when I stopped and looked at how I work with senior engineering colleagues. When a talented engineer joins the team, you don't hand them a diagram of twelve classes and tell them what to name every variable. You explain the business problem, lay out the constraints and domain boundaries, and say: "We need this outcome. Design the system and build it." Once I gave an AI agent an outcome and let it design the whole system, the way I work changed completely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Component-Level Micromanagement Fail?
&lt;/h2&gt;

&lt;p&gt;When developers tell me that AI only produces shallow code, the problem is almost always how they divide the work. They do all the high-level thinking themselves, break the problem into tiny component-level tasks, and then ask the model to implement one isolated piece at a time. In my experience, this micromanagement breaks down in three ways.&lt;/p&gt;

&lt;p&gt;First, the architect stays the bottleneck. If you have to design every component before an agent can write a line of code, your delivery speed is capped by your own calendar. You spend your day writing micro-specifications, reviewing tiny pull requests, and answering questions about isolated classes. You aren't scaling your engineering output. You are just drowning in administrative overhead.&lt;/p&gt;

&lt;p&gt;Second, systems end up fragmented. When you feed an AI agent one component at a time, it has no visibility into how the whole system fits together. It writes a clean repository class, a nice validator, and an isolated controller. But when you wire them up, the pieces clash. The validator expects data the controller never extracts, the database transaction doesn't cover the event publisher, and error handling is inconsistent. Because the model was trapped inside a single component, it couldn't design for the real failure modes of the system.&lt;/p&gt;

&lt;p&gt;Third, you waste the model's actual reasoning strength. Modern frontier models can inspect an entire repository, understand cross-cutting concerns, and reason about architectural trade-offs across multiple layers. When you force the model to only fill in pre-defined method stubs, you discard that capability. You treat an engine capable of designing systems like a line-by-line code generator.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do You Hand Over an Outcome Instead of a Component?
&lt;/h2&gt;

&lt;p&gt;In my post on &lt;a href="https://berris.dev/nodes/ai-assisted-blogging/" rel="noopener noreferrer"&gt;AI-assisted blogging&lt;/a&gt;, I described how I treat AI as a writing partner rather than a ghostwriter. I bring the ideas, the real-world experiences, and the editorial judgment, while the model helps with structure and flow.&lt;/p&gt;

&lt;p&gt;Engineering delegation works the same way, but at the system level.&lt;/p&gt;

&lt;p&gt;The component-level approach looks like this: "Create an order validation service with an interface that checks inventory levels against the database and returns a validation result enum." You have already made every design decision. The model is just typing syntax.&lt;/p&gt;

&lt;p&gt;The outcome-level approach looks like this: "We need an order processing subsystem that handles checkout requests under high concurrency. It must guarantee that we never sell out-of-stock items, handle payment gateway timeouts cleanly without leaving orphaned orders, and publish domain events for downstream shipping. Respect our existing outbox pattern, write the database migrations, and verify the flow with integration tests."&lt;/p&gt;

&lt;p&gt;When you frame the task around an outcome, you let the AI do what engineers do: synthesize a solution. The agent evaluates the data model, chooses how components interact, designs the API contracts, and wires up the persistence layer.&lt;/p&gt;

&lt;p&gt;When you delegate at this macro level, your role shifts upward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Framing outcomes and success criteria:&lt;/strong&gt; You define what the system must accomplish, including performance targets, consistency requirements, and compliance rules.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guarding domain invariants:&lt;/strong&gt; You make sure illegal states are unrepresentable in your domain models, an idea I explored in &lt;a href="https://berris.dev/nodes/using-value-objects-in-net/" rel="noopener noreferrer"&gt;Using Value Objects in .NET&lt;/a&gt;. Strong domain invariants prevent the agent from inventing invalid business logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Setting system boundaries:&lt;/strong&gt; You establish the perimeter, including security policies, external contracts, and integration protocols (as discussed in &lt;a href="https://berris.dev/nodes/designing-apis-for-ai-agents/" rel="noopener noreferrer"&gt;Designing APIs for AI Agents&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reviewing architecture and mentoring:&lt;/strong&gt; You evaluate the agent's proposed system design and review the resulting pull request with the same rigor you would give a human peer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Does System-Level Delegation Look Like in Practice?
&lt;/h2&gt;

&lt;p&gt;Here's how this actually works in practice across our team's daily cadence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Aligning on Outcomes and Architecture
&lt;/h3&gt;

&lt;p&gt;Instead of spending hours drafting component specifications, I write a clear outcome brief for an entire subsystem. I describe the problem, the existing system topology, and our operational constraints.&lt;/p&gt;

&lt;p&gt;Before writing application code, the agent explores the repository and proposes an architectural plan or a draft Architecture Decision Record (ADR). As I wrote in &lt;a href="https://berris.dev/nodes/standardizing-api-conventions/" rel="noopener noreferrer"&gt;Standardizing API Conventions with ADRs&lt;/a&gt;, written records make architectural intent explicit. The agent outlines how it plans to structure components, what data models it needs, and what trade-offs it considered. For example, it might explain why it chose an outbox pattern with transactional messaging over synchronous HTTP calls to downstream services.&lt;/p&gt;

&lt;p&gt;I review that architecture proposal. If the design has a flaw, we discuss it and adjust the plan before any code is written.&lt;/p&gt;

&lt;h3&gt;
  
  
  System-Wide Implementation and Verification
&lt;/h3&gt;

&lt;p&gt;Once we align on the architectural direction, the agent builds the entire system across the codebase. It doesn't stop at one class. It creates the database migrations, implements domain logic, builds API endpoints, configures background workers, and writes integration test suites.&lt;/p&gt;

&lt;p&gt;Because the agent owns the full system, the contracts between components fit together naturally.&lt;/p&gt;

&lt;p&gt;The agent also runs the test suite locally. If an integration test fails because an outbox event didn't publish during a database rollback, the agent diagnoses the failure across both components and fixes the transaction boundary. It doesn't wait for me to find the bug in code review; it verifies its own work.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mentoring and Harness Feedback
&lt;/h3&gt;

&lt;p&gt;When I review the pull request, I look at the big picture: Does this architecture fit our long-term goals? Are our operational failure modes covered?&lt;/p&gt;

&lt;p&gt;If the agent made an architectural misstep, like introducing an unnecessary distributed lock or violating an API convention, I don't just quietly patch the code. I explain the issue and update our repository instructions, ADRs, or test fixtures. Just like mentoring a human colleague, you teach the agent how your team thinks so it doesn't repeat the mistake on the next system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Are the Boundaries Between Macro Delegation and Direct Control?
&lt;/h2&gt;

&lt;p&gt;Delegating entire systems to AI is powerful, but it does not mean stepping away from responsibility. You have to know where your boundaries lie.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Humans Must Own
&lt;/h3&gt;

&lt;p&gt;There are areas where human judgment is irreplaceable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Strategic business trade-offs:&lt;/strong&gt; AI cannot know the organizational politics of your company, which features can be compromised to hit a launch deadline, or how much operational risk your business can absorb. You own the strategy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security perimeters and blast radiuses:&lt;/strong&gt; Authentication mechanisms, tenant isolation rules, authorization boundaries, and secrets management must be strictly guarded. You define the perimeter; the agent builds inside it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Production accountability:&lt;/strong&gt; When a system breaks in production at 3 AM, the AI isn't on call. You are. That is why reviewing the macro architecture, inspecting failure modes, and enforcing automated verification remain your responsibility.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  When to Keep It Interactive
&lt;/h3&gt;

&lt;p&gt;There are two scenarios where autonomous system delegation doesn't work well:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Ambiguous or exploratory requirements:&lt;/strong&gt; If the business problem is still vague and you don't know what outcome you want, delegating an entire system will only produce the wrong architecture faster. In those moments, I sit with product managers and sketch ideas interactively before delegating anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Visual user interfaces:&lt;/strong&gt; Building web or mobile UIs relies heavily on visual feedback, brand feel, and subjective aesthetic judgment. Autonomous delegation often struggles here; interactive pairing with fast hot-reloading remains far more effective.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Save system-level delegation for substantive backend subsystems, data pipelines, integrations, and automated test coverage. Those are the areas where clear specifications, domain invariants, and automated verification shine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Treating AI like an engineering colleague isn't about pretending software has feelings. It is about understanding where human thinking creates the most value.&lt;/p&gt;

&lt;p&gt;When you spend your week micromanaging components, drafting method signatures, and pre-chewing classes, you become a bottleneck. You end up tired, and your systems end up fragmented.&lt;/p&gt;

&lt;p&gt;When you step back, define clear outcomes, and let AI architect and build the whole system, your day changes. You spend your time on what truly matters: defining domain boundaries, evaluating architectural trade-offs, and mentoring.&lt;/p&gt;

&lt;p&gt;If you want real impact from AI, stop pre-chewing every component and start handing over the whole problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What does it mean to delegate at the outcome level instead of the component level?
&lt;/h3&gt;

&lt;p&gt;It means asking AI to solve a complete business problem rather than write an isolated class or method. You provide the goals, domain boundaries, and constraints, and let the agent design the component architecture, data contracts, and implementation across the whole system.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you stop an AI agent from choosing the wrong architecture for a whole system?
&lt;/h3&gt;

&lt;p&gt;Have the agent produce an architectural plan or draft Architecture Decision Record (ADR) before it writes code. By reviewing the proposed component boundaries, data models, and trade-offs upfront, you catch misalignments before implementation starts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does letting AI design whole systems replace the software architect?
&lt;/h3&gt;

&lt;p&gt;No, it elevates the architect. Instead of spending your time specifying low-level interfaces and drawing component boxes, you focus on business strategy, domain invariants, security perimeters, and architectural evaluation.&lt;/p&gt;

&lt;h3&gt;
  
  
  When should you avoid delegating an entire system to an AI agent?
&lt;/h3&gt;

&lt;p&gt;Avoid it when the business requirements are still ambiguous, or when working on visual user interfaces that require subjective aesthetic judgment. In those cases, interactive pairing and rapid prototyping work much better.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://berris.dev/nodes/start-treating-ai-like-a-human-colleague/" rel="noopener noreferrer"&gt;berris.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>softwareengineering</category>
      <category>architecture</category>
      <category>leadership</category>
    </item>
    <item>
      <title>Managing AI Agents on a Remote Devbox</title>
      <dc:creator>Roy Berris</dc:creator>
      <pubDate>Sun, 11 Oct 2026 13:31:33 +0000</pubDate>
      <link>https://dev.to/royberris/managing-ai-agents-on-a-remote-devbox-5ad6</link>
      <guid>https://dev.to/royberris/managing-ai-agents-on-a-remote-devbox-5ad6</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Running autonomous AI coding agents across local terminal tabs and desktop apps drops running tasks whenever your laptop sleeps, and scatters mental focus across unmanaged branches. I moved my agent runtime to a dedicated Linux devbox orchestrated with git worktrees, tmux, and an open-source VS Code extension. This turns VS Code into a persistent control plane, keeps tasks isolated from my working tree, and retains access to private networks and multi-repo test suites.&lt;/p&gt;

&lt;p&gt;I'll be honest with you: when I first started running autonomous AI coding agents, I thought my local terminal was all I needed. I opened &lt;a href="https://docs.anthropic.com/en/docs/agents-and-tools/claude-code/overview" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt; in one tab, &lt;a href="https://openai.com/" rel="noopener noreferrer"&gt;OpenAI Codex&lt;/a&gt; in another, and kept the desktop app open on a second screen for quick questions. It felt fast at first. But within two weeks, that setup completely fell apart under heavy daily use.&lt;/p&gt;

&lt;p&gt;I was losing mental bandwidth trying to remember which terminal tab was refactoring which service. Worse, long-running agent tasks died the moment I shut my laptop lid to step into a meeting. If you run multiple coding agents concurrently, running them locally turns into a constant fight with interrupted sessions and dirty git trees.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Running AI Agents Locally Fall Apart?
&lt;/h2&gt;

&lt;p&gt;Running AI coding agents locally starts simple when you run a single prompt. But once you run multiple tasks concurrently, that setup quickly breaks down in three distinct ways.&lt;/p&gt;

&lt;p&gt;First, you run into open mental loops. Having four or five concurrent agent sessions scattered across local terminals, integrated editor panes, and desktop apps scatters your focus. You forget which task was running where, only to rediscover abandoned branches days later as open pull requests (PRs). Those unresolved loops linger in your head during evenings and weekends.&lt;/p&gt;

&lt;p&gt;Second, local sessions lack durability. The moment you close your laptop, walk into a meeting, or switch between Wi-Fi and ethernet, local terminal connections break. Long-running test suites, dependency builds, or multi-step agent refactors terminate abruptly halfway through execution. You have to inspect the partial state, clean up unstaged changes, and restart the prompt from scratch.&lt;/p&gt;

&lt;p&gt;Third, standalone desktop apps hit hard environment boundaries. The &lt;a href="https://claude.ai/download" rel="noopener noreferrer"&gt;Claude Desktop&lt;/a&gt; app handles isolated questions well, but it falls short in complex environments. It operates in isolation and cannot manage multi-repo workflows where a front-end monorepo must talk directly to a local services monorepo to run end-to-end integration tests. It also lacks access to private networks. It cannot authenticate against internal Virtual Private Clouds (VPCs) or query &lt;a href="https://learn.microsoft.com/azure/azure-monitor/app/app-insights-overview" rel="noopener noreferrer"&gt;Application Insights&lt;/a&gt; logs over a non-production Virtual Private Network (VPN) using a read-only &lt;a href="https://learn.microsoft.com/cli/azure/" rel="noopener noreferrer"&gt;Azure CLI&lt;/a&gt; session.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBzdWJncmFwaCBGcmFnbWVudGVkWyJUaGUgRnJhZ21lbnRlZCBTZXR1cCJdCiAgICAgICAgTG9jYWxUZXJtWyJMb2NhbCBUZXJtaW5hbHM8YnIvPihUZXJtaW5hdGVkIG9uIHNsZWVwKSJdCiAgICAgICAgQ2xhdWRlQXBwWyJEZXNrdG9wIEFwcHM8YnIvPihTaW5nbGUgcmVwbywgbm8gcHJpdmF0ZSBWUEMpIl0KICAgICAgICBWU0NvZGVUZXJtWyJWUyBDb2RlIFBhbmVzPGJyLz4oVW5tYW5hZ2VkIGJyYW5jaGVzIGFuZCBkaXJ0eSB0cmVlcykiXQogICAgZW5kCiAgICBMb2NhbFRlcm0gLS0-IENvbmZ1c2lvblsiTG9zdCBleGVjdXRpb24gY29udGV4dCw8YnIvPmZpbGUgY29sbGlzaW9ucyw8YnIvPm9wZW4gbWVudGFsIGxvb3BzIl0KICAgIENsYXVkZUFwcCAtLT4gQ29uZnVzaW9uCiAgICBWU0NvZGVUZXJtIC0tPiBDb25mdXNpb24KCiAgICBzdHlsZSBGcmFnbWVudGVkIGZpbGw6IzFlMWUyZSxzdHJva2U6IzZjNzA4Nixjb2xvcjojY2RkNmY0CiAgICBzdHlsZSBMb2NhbFRlcm0gZmlsbDojM2IxZjVjLHN0cm9rZTojYzA4NGZjLGNvbG9yOiNmM2U4ZmYKICAgIHN0eWxlIENsYXVkZUFwcCBmaWxsOiMzYjFmNWMsc3Ryb2tlOiNjMDg0ZmMsY29sb3I6I2YzZThmZgogICAgc3R5bGUgVlNDb2RlVGVybSBmaWxsOiMzYjFmNWMsc3Ryb2tlOiNjMDg0ZmMsY29sb3I6I2YzZThmZgogICAgc3R5bGUgQ29uZnVzaW9uIGZpbGw6IzRhMWUxZSxzdHJva2U6I2Y4NzE3MSxjb2xvcjojZmVlMmUy%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBzdWJncmFwaCBGcmFnbWVudGVkWyJUaGUgRnJhZ21lbnRlZCBTZXR1cCJdCiAgICAgICAgTG9jYWxUZXJtWyJMb2NhbCBUZXJtaW5hbHM8YnIvPihUZXJtaW5hdGVkIG9uIHNsZWVwKSJdCiAgICAgICAgQ2xhdWRlQXBwWyJEZXNrdG9wIEFwcHM8YnIvPihTaW5nbGUgcmVwbywgbm8gcHJpdmF0ZSBWUEMpIl0KICAgICAgICBWU0NvZGVUZXJtWyJWUyBDb2RlIFBhbmVzPGJyLz4oVW5tYW5hZ2VkIGJyYW5jaGVzIGFuZCBkaXJ0eSB0cmVlcykiXQogICAgZW5kCiAgICBMb2NhbFRlcm0gLS0-IENvbmZ1c2lvblsiTG9zdCBleGVjdXRpb24gY29udGV4dCw8YnIvPmZpbGUgY29sbGlzaW9ucyw8YnIvPm9wZW4gbWVudGFsIGxvb3BzIl0KICAgIENsYXVkZUFwcCAtLT4gQ29uZnVzaW9uCiAgICBWU0NvZGVUZXJtIC0tPiBDb25mdXNpb24KCiAgICBzdHlsZSBGcmFnbWVudGVkIGZpbGw6IzFlMWUyZSxzdHJva2U6IzZjNzA4Nixjb2xvcjojY2RkNmY0CiAgICBzdHlsZSBMb2NhbFRlcm0gZmlsbDojM2IxZjVjLHN0cm9rZTojYzA4NGZjLGNvbG9yOiNmM2U4ZmYKICAgIHN0eWxlIENsYXVkZUFwcCBmaWxsOiMzYjFmNWMsc3Ryb2tlOiNjMDg0ZmMsY29sb3I6I2YzZThmZgogICAgc3R5bGUgVlNDb2RlVGVybSBmaWxsOiMzYjFmNWMsc3Ryb2tlOiNjMDg0ZmMsY29sb3I6I2YzZThmZgogICAgc3R5bGUgQ29uZnVzaW9uIGZpbGw6IzRhMWUxZSxzdHJva2U6I2Y4NzE3MSxjb2xvcjojZmVlMmUy%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" alt="Diagram" width="921" height="320"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does a Remote Devbox Solve This?
&lt;/h2&gt;

&lt;p&gt;The alternative I moved to is treating my laptop as a thin display and moving the entire agent execution runtime to a dedicated Linux devbox. In my setup, that is an 8-core, 16 GB RAM Ubuntu server running in my home lab, but a cloud virtual machine or an office server works just as well. Instead of running agents directly inside my primary working tree, every agent session runs in its own persistent &lt;a href="https://github.com/tmux/tmux/wiki" rel="noopener noreferrer"&gt;tmux&lt;/a&gt; session attached to an isolated &lt;a href="https://git-scm.com/docs/git-worktree" rel="noopener noreferrer"&gt;git worktree&lt;/a&gt; (&lt;code&gt;agent/&amp;lt;task-name&amp;gt;&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;I considered using separate full repository clones or Docker containers for each agent. But full clones waste disk space on duplicate &lt;code&gt;.git&lt;/code&gt; directories and slow down dependency installs. Docker containers add overhead, slow down file watching, and make running native developer tooling clumsy. Git worktrees share the same underlying repository storage and object history, but check out a completely independent branch into a separate directory. This gives the agent total isolation without duplicating repository data, preventing file-lock collisions and unstaged diff pollution.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBDbGllbnRbIkxhcHRvcCBDbGllbnQ8YnIvPihWUyBDb2RlIFJlbW90ZSAtIFNTSCkiXSAtLT4gRGV2Ym94WyJSZW1vdGUgRGV2Ym94PGJyLz4oVWJ1bnR1IFNlcnZlcikiXQogICAgRGV2Ym94IC0tPiBFeHRlbnNpb25bImRldmJveC12c2NvZGUtZXh0ZW5zaW9uPGJyLz4oQ29udHJvbCBQbGFuZSkiXQogICAgRXh0ZW5zaW9uIC0tPiBUbXV4MVsidG11eCBzZXNzaW9uOjxici8-YWdlbnQvZnJvbnRlbmQtZmVhdHVyZSJdCiAgICBFeHRlbnNpb24gLS0-IFRtdXgyWyJ0bXV4IHNlc3Npb246PGJyLz5hZ2VudC9zZXJ2aWNlcy1hdXRoIl0KICAgIFRtdXgxIC0tPiBXVDFbIldvcmt0cmVlOiB-L3JlcG9zL2Zyb250ZW5kPGJyLz4oYnJhbmNoOiBhZ2VudC9mZWF0dXJlKSJdCiAgICBUbXV4MiAtLT4gV1QyWyJXb3JrdHJlZTogfi9yZXBvcy9zZXJ2aWNlczxici8-KGJyYW5jaDogYWdlbnQvYXV0aCkiXQogICAgV1QxICYgV1QyIC0tPiBTYW5kYm94WyJIb3N0IEVudmlyb25tZW50PGJyLz4oVlBOLCBSZWFkLU9ubHkgQ0xJLCBDb21tYW5kIEFsbG93bGlzdCkiXQoKICAgIHN0eWxlIENsaWVudCBmaWxsOiMwZTNhNGEsc3Ryb2tlOiM2N2U4ZjksY29sb3I6I2UwZjdmZgogICAgc3R5bGUgRGV2Ym94IGZpbGw6IzBlM2E0YSxzdHJva2U6IzY3ZThmOSxjb2xvcjojZTBmN2ZmCiAgICBzdHlsZSBFeHRlbnNpb24gZmlsbDojM2IxZjVjLHN0cm9rZTojYzA4NGZjLGNvbG9yOiNmM2U4ZmYKICAgIHN0eWxlIFRtdXgxIGZpbGw6IzE0NDMyYSxzdHJva2U6IzRhZGU4MCxjb2xvcjojZGNmY2U3CiAgICBzdHlsZSBUbXV4MiBmaWxsOiMxNDQzMmEsc3Ryb2tlOiM0YWRlODAsY29sb3I6I2RjZmNlNwogICAgc3R5bGUgV1QxIGZpbGw6IzRhMmUwZSxzdHJva2U6I2ZiOTIzYyxjb2xvcjojZmZlZGQ1CiAgICBzdHlsZSBXVDIgZmlsbDojNGEyZTBlLHN0cm9rZTojZmI5MjNjLGNvbG9yOiNmZmVkZDUKICAgIHN0eWxlIFNhbmRib3ggZmlsbDojMmQxYTQ3LHN0cm9rZTojYTg1NWY3LGNvbG9yOiNmM2U4ZmY%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZmxvd2NoYXJ0IFRECiAgICBDbGllbnRbIkxhcHRvcCBDbGllbnQ8YnIvPihWUyBDb2RlIFJlbW90ZSAtIFNTSCkiXSAtLT4gRGV2Ym94WyJSZW1vdGUgRGV2Ym94PGJyLz4oVWJ1bnR1IFNlcnZlcikiXQogICAgRGV2Ym94IC0tPiBFeHRlbnNpb25bImRldmJveC12c2NvZGUtZXh0ZW5zaW9uPGJyLz4oQ29udHJvbCBQbGFuZSkiXQogICAgRXh0ZW5zaW9uIC0tPiBUbXV4MVsidG11eCBzZXNzaW9uOjxici8-YWdlbnQvZnJvbnRlbmQtZmVhdHVyZSJdCiAgICBFeHRlbnNpb24gLS0-IFRtdXgyWyJ0bXV4IHNlc3Npb246PGJyLz5hZ2VudC9zZXJ2aWNlcy1hdXRoIl0KICAgIFRtdXgxIC0tPiBXVDFbIldvcmt0cmVlOiB-L3JlcG9zL2Zyb250ZW5kPGJyLz4oYnJhbmNoOiBhZ2VudC9mZWF0dXJlKSJdCiAgICBUbXV4MiAtLT4gV1QyWyJXb3JrdHJlZTogfi9yZXBvcy9zZXJ2aWNlczxici8-KGJyYW5jaDogYWdlbnQvYXV0aCkiXQogICAgV1QxICYgV1QyIC0tPiBTYW5kYm94WyJIb3N0IEVudmlyb25tZW50PGJyLz4oVlBOLCBSZWFkLU9ubHkgQ0xJLCBDb21tYW5kIEFsbG93bGlzdCkiXQoKICAgIHN0eWxlIENsaWVudCBmaWxsOiMwZTNhNGEsc3Ryb2tlOiM2N2U4ZjksY29sb3I6I2UwZjdmZgogICAgc3R5bGUgRGV2Ym94IGZpbGw6IzBlM2E0YSxzdHJva2U6IzY3ZThmOSxjb2xvcjojZTBmN2ZmCiAgICBzdHlsZSBFeHRlbnNpb24gZmlsbDojM2IxZjVjLHN0cm9rZTojYzA4NGZjLGNvbG9yOiNmM2U4ZmYKICAgIHN0eWxlIFRtdXgxIGZpbGw6IzE0NDMyYSxzdHJva2U6IzRhZGU4MCxjb2xvcjojZGNmY2U3CiAgICBzdHlsZSBUbXV4MiBmaWxsOiMxNDQzMmEsc3Ryb2tlOiM0YWRlODAsY29sb3I6I2RjZmNlNwogICAgc3R5bGUgV1QxIGZpbGw6IzRhMmUwZSxzdHJva2U6I2ZiOTIzYyxjb2xvcjojZmZlZGQ1CiAgICBzdHlsZSBXVDIgZmlsbDojNGEyZTBlLHN0cm9rZTojZmI5MjNjLGNvbG9yOiNmZmVkZDUKICAgIHN0eWxlIFNhbmRib3ggZmlsbDojMmQxYTQ3LHN0cm9rZTojYTg1NWY3LGNvbG9yOiNmM2U4ZmY%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" alt="Diagram" width="567" height="758"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This completely changes how I use Visual Studio Code (VS Code). Manual side-by-side programming is gone. VS Code becomes an orchestration dashboard and terminal control plane: one large terminal surface connected to the devbox, a sidebar tracking running sessions and processes, and the file explorer to inspect docs-as-code, &lt;a href="https://berris.dev/nodes/standardizing-api-conventions/" rel="noopener noreferrer"&gt;Architecture Decision Records (ADRs)&lt;/a&gt;, and generated diffs whenever verification is needed.&lt;/p&gt;

&lt;p&gt;A key practical advantage of using &lt;a href="https://code.visualstudio.com/docs/remote/ssh" rel="noopener noreferrer"&gt;VS Code Remote - SSH&lt;/a&gt; here is automatic port forwarding. When you or an agent run a dev server on the devbox (like &lt;code&gt;npm run dev&lt;/code&gt;), VS Code Remote detects the listening process and forwards the port automatically. You can open &lt;code&gt;localhost:3000&lt;/code&gt; directly in your local laptop browser without fiddling with manual SSH tunnel flags. The devbox handles the compute and file watching, but previewing the application feels like running it locally.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does the Extension Architecture Work?
&lt;/h2&gt;

&lt;p&gt;To tie this workflow together without switching windows, I built &lt;a href="https://github.com/royberris/devbox-vscode-extension" rel="noopener noreferrer"&gt;devbox-vscode-extension&lt;/a&gt; (documented in our &lt;a href="https://berris.dev/projects/devbox-vscode-extension/" rel="noopener noreferrer"&gt;Projects catalog&lt;/a&gt;). You can install it directly from the &lt;a href="https://marketplace.visualstudio.com/items?itemName=RoyBerris.devbox-agents" rel="noopener noreferrer"&gt;VS Code Marketplace&lt;/a&gt;. It runs directly inside VS Code over &lt;a href="https://code.visualstudio.com/docs/remote/ssh" rel="noopener noreferrer"&gt;VS Code Remote - SSH&lt;/a&gt; and acts as the interface layer over tmux, git worktrees, and running agent processes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Worktree Isolation and Session Management
&lt;/h3&gt;

&lt;p&gt;The extension eliminates manual git worktree plumbing. Launching an agent provisions a dedicated worktree and spawns a background tmux session tied to that directory.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Automated Worktree Lifecycle:&lt;/strong&gt; When an agent starts, a new branch is cut and mounted in an isolated worktree path under the repository root. When the work is merged or abandoned, tearing down the session removes the worktree cleanly without leaving stale directories behind.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Process Inspection and Cleanup:&lt;/strong&gt; The extension queries the host for running Claude Code and Codex process IDs. It identifies whether a process originated from tmux, an interactive SSH session, or became an orphaned daemon, exposing one-click kill controls right in the sidebar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repository Discovery:&lt;/strong&gt; It scans &lt;code&gt;~/repos&lt;/code&gt; on the devbox, showing ahead and behind commit counts, uncommitted working tree changes, and missing dependencies across every monorepo.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Reviewing Pull Requests Directly in the Editor
&lt;/h3&gt;

&lt;p&gt;Because the extension relies on native git worktrees rather than proprietary containers, it integrates cleanly with the official &lt;a href="https://marketplace.visualstudio.com/items?itemName=GitHub.vscode-pull-request-github" rel="noopener noreferrer"&gt;GitHub Pull Requests and Issues extension&lt;/a&gt;. When an agent completes a task and opens a PR, you check out and review the pull request directly inside VS Code over VS Code Remote - SSH on the existing worktree. You never leave the editor to open a browser tab or desktop application just to inspect code changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Host-Level Sandboxing and Guardrails
&lt;/h3&gt;

&lt;p&gt;Running agents on a dedicated Linux host makes sandboxing practical. Instead of granting blanket permissions on your personal workstation, the devbox runs agents with explicit command and network boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Command Allowlisting:&lt;/strong&gt; Critical deployment tools, production credentials, and destructive system commands are restricted or blocked entirely.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scoped Network Access:&lt;/strong&gt; The devbox maintains a non-production VPN tunnel and a read-only Azure CLI session. Agents can inspect telemetry and logs in &lt;a href="https://learn.microsoft.com/azure/azure-monitor/app/app-insights-overview" rel="noopener noreferrer"&gt;Application Insights&lt;/a&gt; to debug issues without having permissions to modify infrastructure or leak sensitive keys.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Happened in Practice: Gotchas and Trade-offs
&lt;/h2&gt;

&lt;p&gt;Adopting a remote devbox for agent workflows is not without friction. A few real-world trade-offs stand out from my experience:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Hardware Paradox:&lt;/strong&gt; Running agents on a remote 8-core Ubuntu devbox means my powerful local machine (an Apple Silicon M4 Max) sits mostly idle while the devbox compiles code. You trade raw local burst speed for persistence and isolation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate Limit Juggling:&lt;/strong&gt; Heavy agent usage burns through provider quotas quickly. When hitting Claude rate limits or weekly token caps, the devbox setup lets me pivot the existing worktree session to Codex without losing worktree state or branch progress.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Overkill for Casual Use:&lt;/strong&gt; If you run one or two simple prompts a day, maintaining a remote server, SSH keys, VPNs, and worktree extensions is unnecessary overhead. This architecture only pays for itself when managing multiple autonomous tasks concurrently across complex repositories.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Recommendations for You
&lt;/h2&gt;

&lt;p&gt;If you want to move your own agent workflows to a remote machine, start with simple boundaries. Here is what worked best in my setup:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Isolate every agent in a dedicated git worktree:&lt;/strong&gt; Never let an autonomous agent touch your main working copy. Worktrees isolate file locks, dependencies, and experimental commits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep architectural specifications inside the repository:&lt;/strong&gt; Store &lt;a href="https://berris.dev/nodes/standardizing-api-conventions/" rel="noopener noreferrer"&gt;Architecture Decision Records (ADRs)&lt;/a&gt; and specifications directly in the codebase so agents and human reviewers evaluate diffs against the same documented requirements.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run long tasks on persistent remote sessions:&lt;/strong&gt; Relying on local terminal sessions guarantees interrupted runs the moment your laptop sleeps or your Wi-Fi drops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sandbox permissions on a dedicated host:&lt;/strong&gt; Setting up a non-root user, firewall rules, and a command allowlist is straightforward on a dedicated Linux host, whereas locking down your primary laptop without breaking your daily tools is nearly impossible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Consolidate sessions into a single control plane:&lt;/strong&gt; Keep your terminal, session list, diff review, and pull request workflows anchored inside a single editor window instead of scattering them across multiple apps.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Moving AI coding agents from my local laptop to a remote Linux devbox solved the two biggest headaches I had: lost execution state and scattered mental focus. With git worktrees for isolation, tmux for session durability, and VS Code Remote - SSH as a control plane, I can kick off three agent refactors, shut my laptop, and come back to review clean pull requests when I'm ready.&lt;/p&gt;

&lt;p&gt;Autonomous coding agents are only as good as the environment you give them to run in.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Why use git worktrees instead of Docker containers or fresh git clones?
&lt;/h3&gt;

&lt;p&gt;Git worktrees share the same underlying repository storage and commit history without downloading gigabytes of duplicate files, while keeping the agent's changes completely isolated from your main working tree. They avoid the file-watching and tooling friction of Docker containers while preventing file-lock collisions.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens to a running agent session if my laptop loses Wi-Fi?
&lt;/h3&gt;

&lt;p&gt;Nothing stops running. Because the agent executes inside a tmux session on the remote devbox, tests keep running and the agent keeps editing code. When you reconnect over VS Code Remote - SSH, you attach right back to the active session.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you sandbox agent permissions on the remote devbox?
&lt;/h3&gt;

&lt;p&gt;The devbox runs agents with command allowlists that block critical deployment commands and destructive system tools. It uses a non-production VPN tunnel and a read-only Azure CLI login, letting agents inspect telemetry in &lt;a href="https://learn.microsoft.com/azure/azure-monitor/app/app-insights-overview" rel="noopener noreferrer"&gt;Application Insights&lt;/a&gt; without granting permission to change infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is a remote devbox setup worth it for casual AI coding?
&lt;/h3&gt;

&lt;p&gt;No. If you only prompt an AI assistant a couple of times a day for small snippets, setting up a remote server, SSH keys, and worktree tooling is unnecessary overhead. It only pays off when you run multiple autonomous tasks concurrently across complex or multi-repo projects.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://berris.dev/nodes/managing-ai-agents-on-a-remote-devbox/" rel="noopener noreferrer"&gt;berris.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>vscode</category>
      <category>linux</category>
    </item>
    <item>
      <title>Designing APIs for AI Agents: Schemas, Security and MCP</title>
      <dc:creator>Roy Berris</dc:creator>
      <pubDate>Sun, 11 Oct 2026 13:31:28 +0000</pubDate>
      <link>https://dev.to/royberris/designing-apis-for-ai-agents-schemas-security-and-mcp-59</link>
      <guid>https://dev.to/royberris/designing-apis-for-ai-agents-schemas-security-and-mcp-59</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; &lt;a href="https://www.postman.com/state-of-api/2025/" rel="noopener noreferrer"&gt;Postman's 2025 State of the API report&lt;/a&gt; shows that 89% of developers use AI in their daily work, but only 24% design APIs with AI agents in mind. My answer is to design for humans and AI agents at the same time: treat the schema as the shared language, put business context in it, keep every pattern consistent, rethink security for automated consumers and design endpoints as tools with clear contracts.&lt;/p&gt;

&lt;p&gt;The API development world changed a lot in 2024, and it caught many of us by surprise. While we were busy making APIs better for human developers, a new consumer appeared that works very fast: AI agents. &lt;a href="https://www.postman.com/state-of-api/2025/" rel="noopener noreferrer"&gt;Postman's 2025 State of the API report&lt;/a&gt; shows clearly that 89% of developers now use AI tools every day, but only 24% design APIs with AI agents in mind. This gap shows a big problem that needs new design patterns.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Do AI Agents Need Different API Design?
&lt;/h2&gt;

&lt;p&gt;When I design APIs today, I still think about the developer who will read my docs, understand my endpoint patterns, and write code to connect with it. But here's what Postman's research showed that really changed how I think about building APIs: AI agents are already using APIs at huge scale with a 40% increase from last year.&lt;/p&gt;

&lt;p&gt;The disconnect is there. While 89% of developers use AI tools for making code and solving problems, most of us keep designing APIs using patterns made for human use. Only 13% design equally for humans and AI agents, while just 7% mainly design for AI agents. This mismatch creates basic problems when AI agents meet APIs that don't have clear schemas, typed errors, and clear behavioral rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do You Design an API for Both Humans and AI Agents?
&lt;/h2&gt;

&lt;p&gt;The key thing I've learned is that AI agents are trained on human language, which means we shouldn't design only for machines. Instead, we need to design for both humans and machines at the same time through consistent, well-documented interfaces that show intent and purpose.&lt;/p&gt;

&lt;h3&gt;
  
  
  Schema as the Shared Language
&lt;/h3&gt;

&lt;p&gt;I think of the API schema as the shared language between humans and machines. The schema is not just a technical contract. It's a complete way to communicate that shows intent, purpose, and business context in ways both developers and AI agents can understand.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Semantic Metadata: Intent and Purpose&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The schema should tell a story about what the API does and why it exists. This means using clear field names that show business concepts, useful error codes that help fix problems, and response structures that show how things work together. When an AI agent sees a well-designed schema, it understands not just what data to send, but why that data matters and how it connects to real business work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Functional Documentation: Business Context&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Traditional &lt;a href="https://spec.openapis.org/oas/latest.html" rel="noopener noreferrer"&gt;OpenAPI&lt;/a&gt; specs focus on technical contracts, but AI agents need business context to make smart decisions. I learned this the hard way when building APIs that agents couldn't use effectively. Now I put functional context directly into schema definitions through custom properties (&lt;a href="https://spec.openapis.org/oas/latest.html#specification-extensions" rel="noopener noreferrer"&gt;OpenAPI specification extensions&lt;/a&gt;, the &lt;code&gt;x-&lt;/code&gt; fields) that give semantic meaning, links that connect technical operations to business outcomes, and workflow descriptions that show how endpoints work together to solve real problems.&lt;/p&gt;

&lt;p&gt;This approach changes the schema from a purely technical thing into a complete communication tool. Instead of keeping separate technical and functional documentation, the schema becomes one source of truth that shows both implementation details and business intent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Navigation Metadata: Easy Discovery&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The schema should help find related functionality through structured navigation data. This includes relationship patterns that mirror real-world workflows, endpoint hierarchies organized around functional use cases rather than technical resource structures, and built-in guidance that helps both humans and AI agents understand when and how to use specific operations.&lt;/p&gt;

&lt;p&gt;In my experience, organizing APIs around the problems they solve, rather than just data model relationships, makes them much easier for human developers while giving AI agents clear functional context about usage patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example: Traditional vs. Agent-Aware Schema Design&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's how the same endpoint looks when designed traditionally versus with AI agents in mind:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Traditional approach - technical focus&lt;/span&gt;
&lt;span class="s"&gt;/api/users/{id}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;200&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
                &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
                &lt;span class="na"&gt;email&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Agent-aware approach - semantic focus&lt;/span&gt;
&lt;span class="s"&gt;/api/users/{userId}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Retrieve&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;profile&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;account&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;management"&lt;/span&gt;
    &lt;span class="na"&gt;x-business-context&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;purpose&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Account&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;management&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;support"&lt;/span&gt;
      &lt;span class="na"&gt;workflows&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user-lookup"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing-inquiry"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;userId&lt;/span&gt;
        &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
          &lt;span class="na"&gt;pattern&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;^usr_[a-zA-Z0-9]{16}$"&lt;/span&gt;
        &lt;span class="na"&gt;x-semantic-meaning&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Primary&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;identifier"&lt;/span&gt;
    &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;200&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;application/json&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;userId&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;string&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
                &lt;span class="na"&gt;profile&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
                  &lt;span class="na"&gt;x-business-purpose&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Identity&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;verification"&lt;/span&gt;
                &lt;span class="na"&gt;accountStatus&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
                  &lt;span class="na"&gt;enum&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;active"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;suspended"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pending"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
                  &lt;span class="na"&gt;x-business-impact&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Determines&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;available&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;actions"&lt;/span&gt;
      &lt;span class="na"&gt;404&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;User&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;found"&lt;/span&gt;
        &lt;span class="na"&gt;x-remediation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Verify&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ID&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;format"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Consistency as a Design Language
&lt;/h3&gt;

&lt;p&gt;Consistency becomes the design language that both humans and machines can learn and use. This means using the same naming rules across all endpoints, standard HTTP status code usage, the same authentication methods, and error handling patterns that create a predictable interaction model.&lt;/p&gt;

&lt;p&gt;I use a pattern where every endpoint follows the same basic template: consistent parameter naming that shows business concepts, standard response envelopes that tell a complete story, the same error response formats that give useful guidance, and predictable resource relationship patterns that mirror real-world workflows. This consistency allows both human developers and AI agents to learn patterns once and apply them across the entire API surface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Unified Communication Architecture&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The diagram below shows how schema-driven design creates a unified communication layer that serves both human developers and AI agents:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZ3JhcGggVEIKICAgIHN1YmdyYXBoICJBUEkgU2NoZW1hIExheWVyIgogICAgICAgIFNjaGVtYVtBUEkgU2NoZW1hPGJyLz5PcGVuQVBJICsgRXh0ZW5zaW9uc10KICAgICAgICBTY2hlbWEgLS0-IFNNW1NlbWFudGljIE1ldGFkYXRhPGJyLz5CdXNpbmVzcyBJbnRlbnRdCiAgICAgICAgU2NoZW1hIC0tPiBGRFtGdW5jdGlvbmFsIERvY3VtZW50YXRpb248YnIvPlVzZSBDYXNlIENvbnRleHRdCiAgICAgICAgU2NoZW1hIC0tPiBOTVtOYXZpZ2F0aW9uIE1ldGFkYXRhPGJyLz5Xb3JrZmxvdyBSZWxhdGlvbnNoaXBzXQogICAgZW5kCiAgICAKICAgIHN1YmdyYXBoICJIdW1hbiBDb25zdW1lcnMiCiAgICAgICAgRGV2W0RldmVsb3Blcl0KICAgICAgICBEZXZUb29sc1tEZXZlbG9wbWVudCBUb29sczxici8-UG9zdG1hbiwgSW5zb21uaWFdCiAgICAgICAgRG9jc1tBUEkgRG9jdW1lbnRhdGlvbjxici8-U3dhZ2dlciBVSSwgUmVkb2NdCiAgICBlbmQKICAgIAogICAgc3ViZ3JhcGggIkFJIENvbnN1bWVycyIKICAgICAgICBBZ2VudFtBSSBBZ2VudF0KICAgICAgICBMTE1bTGFuZ3VhZ2UgTW9kZWw8YnIvPkdQVCwgQ2xhdWRlXQogICAgICAgIEF1dG9Ub29sc1tBdXRvbWF0aW9uIFRvb2xzPGJyLz5aYXBpZXIsIE1DUCBDbGllbnRzXQogICAgZW5kCiAgICAKICAgIHN1YmdyYXBoICJTaGFyZWQgVW5kZXJzdGFuZGluZyIKICAgICAgICBQYXR0ZXJuc1tDb25zaXN0ZW50IFBhdHRlcm5zPGJyLz5OYW1pbmcsIEVycm9ycywgQXV0aF0KICAgICAgICBDb250ZXh0W0J1c2luZXNzIENvbnRleHQ8YnIvPlB1cnBvc2UsIFdvcmtmbG93c10KICAgICAgICBEaXNjb3ZlcnlbRGlzY292ZXJhYmlsaXR5PGJyLz5SZWxhdGVkIE9wZXJhdGlvbnNdCiAgICBlbmQKICAgIAogICAgU2NoZW1hIC0tPiBQYXR0ZXJucwogICAgU2NoZW1hIC0tPiBDb250ZXh0CiAgICBTY2hlbWEgLS0-IERpc2NvdmVyeQogICAgCiAgICBQYXR0ZXJucyAtLT4gRGV2CiAgICBQYXR0ZXJucyAtLT4gQWdlbnQKICAgIAogICAgQ29udGV4dCAtLT4gRGV2CiAgICBDb250ZXh0IC0tPiBBZ2VudAogICAgCiAgICBEaXNjb3ZlcnkgLS0-IERldgogICAgRGlzY292ZXJ5IC0tPiBBZ2VudAogICAgCiAgICBEZXYgLS0-IERldlRvb2xzCiAgICBEZXYgLS0-IERvY3MKICAgIAogICAgQWdlbnQgLS0-IExMTQogICAgQWdlbnQgLS0-IEF1dG9Ub29scwogICAgCiAgICBzdHlsZSBTY2hlbWEgZmlsbDojMGUzYTRhLHN0cm9rZTojNjdlOGY5LGNvbG9yOiNlMGY3ZmYKICAgIHN0eWxlIFBhdHRlcm5zIGZpbGw6IzNiMWY1YyxzdHJva2U6I2MwODRmYyxjb2xvcjojZjNlOGZmCiAgICBzdHlsZSBDb250ZXh0IGZpbGw6IzE0NDMyYSxzdHJva2U6IzRhZGU4MCxjb2xvcjojZGNmY2U3CiAgICBzdHlsZSBEaXNjb3ZlcnkgZmlsbDojNGEyZTBlLHN0cm9rZTojZmI5MjNjLGNvbG9yOiNmZmVkZDU%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fmermaid.ink%2Fimg%2FZ3JhcGggVEIKICAgIHN1YmdyYXBoICJBUEkgU2NoZW1hIExheWVyIgogICAgICAgIFNjaGVtYVtBUEkgU2NoZW1hPGJyLz5PcGVuQVBJICsgRXh0ZW5zaW9uc10KICAgICAgICBTY2hlbWEgLS0-IFNNW1NlbWFudGljIE1ldGFkYXRhPGJyLz5CdXNpbmVzcyBJbnRlbnRdCiAgICAgICAgU2NoZW1hIC0tPiBGRFtGdW5jdGlvbmFsIERvY3VtZW50YXRpb248YnIvPlVzZSBDYXNlIENvbnRleHRdCiAgICAgICAgU2NoZW1hIC0tPiBOTVtOYXZpZ2F0aW9uIE1ldGFkYXRhPGJyLz5Xb3JrZmxvdyBSZWxhdGlvbnNoaXBzXQogICAgZW5kCiAgICAKICAgIHN1YmdyYXBoICJIdW1hbiBDb25zdW1lcnMiCiAgICAgICAgRGV2W0RldmVsb3Blcl0KICAgICAgICBEZXZUb29sc1tEZXZlbG9wbWVudCBUb29sczxici8-UG9zdG1hbiwgSW5zb21uaWFdCiAgICAgICAgRG9jc1tBUEkgRG9jdW1lbnRhdGlvbjxici8-U3dhZ2dlciBVSSwgUmVkb2NdCiAgICBlbmQKICAgIAogICAgc3ViZ3JhcGggIkFJIENvbnN1bWVycyIKICAgICAgICBBZ2VudFtBSSBBZ2VudF0KICAgICAgICBMTE1bTGFuZ3VhZ2UgTW9kZWw8YnIvPkdQVCwgQ2xhdWRlXQogICAgICAgIEF1dG9Ub29sc1tBdXRvbWF0aW9uIFRvb2xzPGJyLz5aYXBpZXIsIE1DUCBDbGllbnRzXQogICAgZW5kCiAgICAKICAgIHN1YmdyYXBoICJTaGFyZWQgVW5kZXJzdGFuZGluZyIKICAgICAgICBQYXR0ZXJuc1tDb25zaXN0ZW50IFBhdHRlcm5zPGJyLz5OYW1pbmcsIEVycm9ycywgQXV0aF0KICAgICAgICBDb250ZXh0W0J1c2luZXNzIENvbnRleHQ8YnIvPlB1cnBvc2UsIFdvcmtmbG93c10KICAgICAgICBEaXNjb3ZlcnlbRGlzY292ZXJhYmlsaXR5PGJyLz5SZWxhdGVkIE9wZXJhdGlvbnNdCiAgICBlbmQKICAgIAogICAgU2NoZW1hIC0tPiBQYXR0ZXJucwogICAgU2NoZW1hIC0tPiBDb250ZXh0CiAgICBTY2hlbWEgLS0-IERpc2NvdmVyeQogICAgCiAgICBQYXR0ZXJucyAtLT4gRGV2CiAgICBQYXR0ZXJucyAtLT4gQWdlbnQKICAgIAogICAgQ29udGV4dCAtLT4gRGV2CiAgICBDb250ZXh0IC0tPiBBZ2VudAogICAgCiAgICBEaXNjb3ZlcnkgLS0-IERldgogICAgRGlzY292ZXJ5IC0tPiBBZ2VudAogICAgCiAgICBEZXYgLS0-IERldlRvb2xzCiAgICBEZXYgLS0-IERvY3MKICAgIAogICAgQWdlbnQgLS0-IExMTQogICAgQWdlbnQgLS0-IEF1dG9Ub29scwogICAgCiAgICBzdHlsZSBTY2hlbWEgZmlsbDojMGUzYTRhLHN0cm9rZTojNjdlOGY5LGNvbG9yOiNlMGY3ZmYKICAgIHN0eWxlIFBhdHRlcm5zIGZpbGw6IzNiMWY1YyxzdHJva2U6I2MwODRmYyxjb2xvcjojZjNlOGZmCiAgICBzdHlsZSBDb250ZXh0IGZpbGw6IzE0NDMyYSxzdHJva2U6IzRhZGU4MCxjb2xvcjojZGNmY2U3CiAgICBzdHlsZSBEaXNjb3ZlcnkgZmlsbDojNGEyZTBlLHN0cm9rZTojZmI5MjNjLGNvbG9yOiNmZmVkZDU%3Ftype%3Dpng%26theme%3Ddark%26bgColor%3D111827" alt="Diagram" width="1904" height="571"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Should API Security Change for AI Agents?
&lt;/h2&gt;

&lt;p&gt;The security implications of AI agents as API consumers present a big challenge, but one that can be addressed with thoughtful design patterns. Postman's research shows that 51% of developers now cite unauthorized agent access as their top security concern, highlighting the need for better security approaches.&lt;/p&gt;

&lt;p&gt;Traditional API security designs assumed predictable human behavior, developers making dozens of calls per day, following documented patterns, operating within reasonable rate limits. AI agents challenge these assumptions by operating at high speeds, keeping persistent automated access, and potentially turning a single compromised API key into a gateway for extensive data extraction. However, these challenges create opportunities to build stronger security designs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Behavioral Security Architecture
&lt;/h3&gt;

&lt;p&gt;The unpredictable behavior of AI agents makes it hard to tell legitimate automation from attacks using traditional rule-based approaches. While I haven't needed to implement these patterns at scale yet, my team is exploring security designs that move beyond static rules to behavioral analysis systems that can recognize patterns in real-time and adjust responses dynamically.&lt;/p&gt;

&lt;p&gt;This approach needs dynamic rate limiting based on behavioral patterns, better monitoring for suspicious activity, and shorter-lived credentials with automatic rotation. The key is building systems that can tell the difference between legitimate automation and potential attacks through behavior rather than static rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does the Model Context Protocol (MCP) Change?
&lt;/h2&gt;

&lt;p&gt;The emergence of the &lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Model Context Protocol&lt;/a&gt; represents an interesting development in API architecture for AI use. While 70% of developers know about MCP according to Postman's research, only 10% use it regularly. This points to growing interest but limited readiness in the ecosystem.&lt;/p&gt;

&lt;p&gt;MCP introduces new patterns for structured interfaces between AI models and real-world systems. It addresses critical problems like unified agent access, standard security models, and structured tool definitions that agents can reliably understand. The protocol defines clear boundaries between what AI agents can discover, understand, and invoke.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Principles from MCP
&lt;/h3&gt;

&lt;p&gt;Regardless of whether MCP becomes the standard, the principles it embodies represent the direction we need to move. Structured interfaces with explicit capability declarations, clear tool definitions with typed parameters and responses, standard security models that work across different agent implementations, and discovery mechanisms that allow agents to understand available functionality.&lt;/p&gt;

&lt;p&gt;In my projects, implementing these principles improves API architecture even without full MCP adoption. Making APIs agent-consumable through structured interfaces provides immediate benefits and future flexibility, whether MCP succeeds or alternative standards emerge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example: Semantic Extensions in Practice&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's how I extend OpenAPI schemas with semantic metadata in my projects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"paths"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"/customers/{customerId}/support-tickets"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"post"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"summary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Create customer support ticket"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"x-business-context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"purpose"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Enable customers to report issues"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"workflow-stage"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"issue-reporting"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"related-operations"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"operation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"GET /customers/{customerId}/support-tickets"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"relationship"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"list-related"&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"requestBody"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"application/json"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"properties"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="nl"&gt;"subject"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="nl"&gt;"x-semantic-purpose"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Primary classification for routing"&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="nl"&gt;"priority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="nl"&gt;"enum"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"low"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"critical"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="nl"&gt;"x-business-rules"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                      &lt;/span&gt;&lt;span class="nl"&gt;"critical"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Service outage affecting multiple customers"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                      &lt;/span&gt;&lt;span class="nl"&gt;"high"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Feature broken for paying customer"&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"responses"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"201"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"x-business-outcome"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Customer issue is now tracked"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="nl"&gt;"application/json"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="nl"&gt;"schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="nl"&gt;"properties"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="nl"&gt;"ticketId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
                      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"string"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
                      &lt;/span&gt;&lt;span class="nl"&gt;"x-semantic-purpose"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Primary reference for follow-up"&lt;/span&gt;&lt;span class="w"&gt;
                    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
                  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
                &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
              &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
            &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This shows how semantic extensions (&lt;code&gt;x-business-context&lt;/code&gt;, &lt;code&gt;x-semantic-purpose&lt;/code&gt;) provide context for both developers and AI agents.&lt;/p&gt;

&lt;h3&gt;
  
  
  Preparing for Protocol Evolution
&lt;/h3&gt;

&lt;p&gt;The low adoption rate suggests practical barriers that the industry needs to address, but the design patterns remain valuable. Building APIs with explicit capability declarations, structured tool definitions, and clear security boundaries positions systems to adapt to whatever standards emerge.&lt;/p&gt;

&lt;p&gt;This approach means designing APIs as tool catalogs rather than simple data interfaces. Each endpoint becomes a tool with clear inputs, outputs, and behavioral contracts that both human developers and AI agents can understand and invoke reliably.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Design Priorities for the Future
&lt;/h2&gt;

&lt;p&gt;Based on what I've seen in recent projects and the trends Postman identified, there are big changes coming in API architecture. The shift from human-focused to machine-focused design needs new patterns, different security models, and better documentation strategies. Three design priorities emerge as critical for systems that need to support both human developers and AI agents well.&lt;/p&gt;

&lt;h3&gt;
  
  
  Machine-First Interface Design
&lt;/h3&gt;

&lt;p&gt;APIs must be built with AI agents as primary users rather than afterthoughts. This means designing interfaces that provide machine readable schemas, predictable patterns, and comprehensive behavioral specifications from the ground up. The architecture should assume automated use and optimize for reliability at high speeds.&lt;/p&gt;

&lt;h3&gt;
  
  
  Adaptive Security Architecture
&lt;/h3&gt;

&lt;p&gt;Security models must evolve beyond traditional patterns to handle high-speed exploitation, persistent automated attacks, and unpredictable behavior. This needs architectural approaches that can tell the difference between legitimate automation and attacks through behavioral analysis rather than static rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tool-Oriented Design Patterns
&lt;/h3&gt;

&lt;p&gt;API architecture must shift from simple data interfaces to structured tool catalogs where each endpoint represents a well-defined capability with clear inputs, outputs, and behavioral contracts. This design pattern enables both human developers and AI agents to understand and invoke functionality reliably.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The API landscape stands at a turning point where we must design for both human understanding and machine use at the same time. The key insight is that AI agents are trained on human language patterns, which creates an opportunity to build unified architectures that serve both audiences well.&lt;/p&gt;

&lt;p&gt;My experience shows that the practices needed for effective human-machine API use (consistent patterns, schema-driven functional documentation, and discoverable use case context) create better APIs overall. These aren't competing design goals but approaches that work together and strengthen each other.&lt;/p&gt;

&lt;p&gt;The future belongs to APIs that work as complete communication systems, where schemas serve as natural language interfaces and functional documentation is discoverable through the same mechanisms that AI agents use for technical discovery. The design patterns we implement today will determine whether our systems can communicate effectively with both human developers and AI agents.&lt;/p&gt;

&lt;p&gt;The time to start building these unified communication architectures is now.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How many developers design APIs for AI agents?
&lt;/h3&gt;

&lt;p&gt;According to &lt;a href="https://www.postman.com/state-of-api/2025/" rel="noopener noreferrer"&gt;Postman's 2025 State of the API report&lt;/a&gt;, 24% of developers design APIs with AI agents in mind. Only 13% design equally for humans and AI agents, and 7% mainly design for AI agents. At the same time, 89% of developers use AI in their daily work.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should an API schema include for AI agents?
&lt;/h3&gt;

&lt;p&gt;Clear field names that match business concepts, useful error codes, and the business context behind each operation. I add that context with custom &lt;code&gt;x-&lt;/code&gt; extensions such as &lt;code&gt;x-business-context&lt;/code&gt; and &lt;code&gt;x-semantic-purpose&lt;/code&gt;, including the workflow an endpoint belongs to and which operations relate to it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need MCP to make my API ready for AI agents?
&lt;/h3&gt;

&lt;p&gt;No. MCP adoption is still low, but its principles already help: explicit capability declarations, typed tool definitions, standard security models and discovery. In my projects, applying those principles improves the API even without full MCP adoption.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do AI agents change API security?
&lt;/h3&gt;

&lt;p&gt;AI agents work at high speed with persistent automated access, so a single compromised API key can open the door to large-scale data extraction. My team is exploring behavioral analysis, dynamic rate limiting and shorter-lived credentials with automatic rotation instead of static rules. I haven't needed these patterns at scale yet.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://berris.dev/nodes/designing-apis-for-ai-agents/" rel="noopener noreferrer"&gt;berris.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>architecture</category>
      <category>mcp</category>
    </item>
  </channel>
</rss>
