<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Paul Spread</title>
    <description>The latest articles on DEV Community by Paul Spread (@spread2009).</description>
    <link>https://dev.to/spread2009</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F291783%2F9845631e-84be-4923-ac15-143423dbf9c7.png</url>
      <title>DEV Community: Paul Spread</title>
      <link>https://dev.to/spread2009</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/spread2009"/>
    <language>en</language>
    <item>
      <title>Four Strategies Around the bStock Delta: Arbitrage, Market Making, Signals, Alerts</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:53:31 +0000</pubDate>
      <link>https://dev.to/spread2009/four-strategies-around-the-bstock-delta-arbitrage-market-making-signals-alerts-4hhn</link>
      <guid>https://dev.to/spread2009/four-strategies-around-the-bstock-delta-arbitrage-market-making-signals-alerts-4hhn</guid>
      <description>&lt;p&gt;A tokenized stock lives on two markets at once: the token on a crypto exchange, the share on a stock exchange. The gap between them — the delta — supports several working strategies. Here is each one, with its risks.&lt;/p&gt;

&lt;p&gt;A decade ago this toolkit belonged to institutions with dedicated market-data feeds and quant desks. Today an AI agent with a wallet can run the same loop — watch, decide, pay for its own data. That is what makes the tokenized-stocks market different from the equity market it mirrors: the tooling is open to anyone with USDC.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fopttaqdelhdfz7i0qw57.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fopttaqdelhdfz7i0qw57.png" alt="Four strategy icons orbiting a delta chart — the token tracks the share, the band stretches" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0opl4n1svlgg6gxiuqpl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0opl4n1svlgg6gxiuqpl.png" alt="Diagram: four strategies from one delta" width="800" height="180"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;One number, four strategies: arbitrage (buy the cheap token, wait for convergence), market making (earn the spread on thin books), signal trading (the delta leads Monday's gap), and alerts as a service (the signal finds you).&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Delta arbitrage
&lt;/h2&gt;

&lt;p&gt;When the token trades below the share by more than your costs (fees + spread + slippage), buy the token and wait for convergence. Classic setup: a −2% delta on a closed exchange → buy the token → exit at parity. The risks: the delta can widen further (cut it with a stop), convergence can take days, and pool liquidity caps your position size. Enter only when you understand &lt;em&gt;why&lt;/em&gt; the gap appeared.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzj4bskufw426aab4obv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzj4bskufw426aab4obv.png" alt="Delta arbitrage: the delta dips to −2%, a buy-the-token arrow at the dip, then convergence back to zero" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Market making
&lt;/h2&gt;

&lt;p&gt;bStock pools are thin — wide spreads mean you get paid for providing liquidity. A market maker earns the spread and fees, and the delta hints where quotes should move: token above the share — expect sellers; below — buyers.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Signal trading
&lt;/h2&gt;

&lt;p&gt;Delta is a leading indicator. The token reprices on weekend news before the equity market opens. A trader watching the delta knows about Monday's gap on Saturday — and positions accordingly, with the usual risk that the gap never comes.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Alerts as a service
&lt;/h2&gt;

&lt;p&gt;You do not have to trade the delta yourself — monitoring it is a product. Our tracker exposes it over MCP and Telegram: subscribe once, and a message arrives whenever any symbol's delta crosses the 0.5% threshold. The machine watches; you trade when it matters. The subscription settles on &lt;strong&gt;Arc&lt;/strong&gt; in USDC — one on-chain transaction, receipt-verified, 30 days of real-time. The same asset you trade bStocks with pays for the signal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The main rule
&lt;/h2&gt;

&lt;p&gt;Delta is not free money — it is the price of risk: liquidity, issuer counterparty, rebalancing delays. A strategy works while costs stay below the divergence. Count the costs before entry, not after.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Series finale: the risks of tokenized stocks — what can go wrong.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent guide (endpoints, limits, examples): &lt;a href="https://agentbadge.xyz/bstock-guide" rel="noopener noreferrer"&gt;agentbadge.xyz/bstock-guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;All articles in the series: &lt;a href="https://agentbadge.xyz/blog" rel="noopener noreferrer"&gt;agentbadge.xyz/blog&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP endpoint: &lt;code&gt;https://agentbadge.xyz/mcp/bstock/tools/get_delta&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Not financial advice. These strategies carry real risk — size positions accordingly.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>cryptocurrency</category>
      <category>trading</category>
      <category>arbitrage</category>
      <category>ai</category>
    </item>
    <item>
      <title>bStocks for Beginners: Tokenized Stocks Without the Magic</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Mon, 28 Sep 2026 07:14:58 +0000</pubDate>
      <link>https://dev.to/spread2009/bstocks-for-beginners-tokenized-stocks-without-the-magic-23jc</link>
      <guid>https://dev.to/spread2009/bstocks-for-beginners-tokenized-stocks-without-the-magic-23jc</guid>
      <description>&lt;p&gt;A tokenized stock is a blockchain token whose price follows a real share. AAPLB mirrors Apple, TSLAB mirrors Tesla. Buy the token — get the same price exposure as the share, but on a crypto exchange, around the clock.&lt;/p&gt;

&lt;p&gt;This is not a niche experiment — it is a bridge between two markets. For the financial ecosystem, tokenized stocks mean equities that trade when Wall Street sleeps, in sizes that fit any wallet, settled in stablecoins. For traders, they mean a new class of instruments — and a new spread to understand before trading it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj935yyngbrqqrqn4bx86.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj935yyngbrqqrqn4bx86.png" alt="A share coin and a token coin connected by a stretched rubber band — the token tracks the share, but the band can stretch" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9wu1jple470ud1igsqko.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9wu1jple470ud1igsqko.png" alt="Diagram: how a bStock works" width="800" height="1250"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The issuer holds the real shares and issues tokens (AAPLB = Apple × 1.0006) that trade on Binance 24/7, settled in USDC. While the two markets agree, the delta is near zero; when they diverge — nights, weekends, news — the gap becomes an opportunity for some and a risk for others.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works under the hood
&lt;/h2&gt;

&lt;p&gt;The issuer holds the real shares (or an equivalent) and issues tokens that mirror their value. The token trades on an exchange (for bStocks — Binance) 24/7, settles in stablecoins, and its price is anchored to the share through a multiplier — for AAPLB it is about 1.0006.&lt;/p&gt;

&lt;h2&gt;
  
  
  How a token differs from a share
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Trading hours.&lt;/strong&gt; The share trades during the exchange session; the token trades 24/7, weekends included.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Access.&lt;/strong&gt; The share needs a broker, KYC, minimum lots; the token needs an exchange account and USDC.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rights.&lt;/strong&gt; The token usually carries no voting rights and no direct dividends — it is a price tracker, not equity in the company. Read the issuer's terms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fractionality.&lt;/strong&gt; You can hold 0.01 of a share.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq1qn6puvugginjjbe0ep.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq1qn6puvugginjjbe0ep.png" alt="Comparison: share versus token — trading hours, access, rights, fractionality" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why prices diverge — and what the delta is
&lt;/h2&gt;

&lt;p&gt;Token and share are two different markets with different liquidity. At night the share sleeps while the token keeps trading — their prices drift apart. That gap is the &lt;strong&gt;delta&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A concrete picture: news breaks on Saturday. The stock cannot move until Monday's open. The token drops 3% within minutes. That is not the token "breaking" — that is the token market pricing the news before the stock market can. The delta makes this visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to start
&lt;/h2&gt;

&lt;p&gt;Exchange account → USDC → buy a bStock. From the first trade, watch the delta: it shows how far the token has drifted from the original and hints at both opportunity (convergence trades) and risk (the gap can widen). Monitoring dozens of tokens around the clock is a job for a machine — which is exactly what delta trackers, including ours for AI agents, are for.&lt;/p&gt;

&lt;p&gt;Our tracker exposes the delta to AI agents over MCP, and its paid tier runs on &lt;strong&gt;Arc&lt;/strong&gt; — Circle's USDC-native chain — so an agent pays for real-time data the same way it holds collateral: in USDC, on-chain.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next: strategies people actually run on the delta.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent guide (endpoints, limits, examples): &lt;a href="https://agentbadge.xyz/bstock-guide" rel="noopener noreferrer"&gt;agentbadge.xyz/bstock-guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;All articles in the series: &lt;a href="https://agentbadge.xyz/blog" rel="noopener noreferrer"&gt;agentbadge.xyz/blog&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP endpoint: &lt;code&gt;https://agentbadge.xyz/mcp/bstock/tools/get_delta&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>cryptocurrency</category>
      <category>beginners</category>
      <category>tutorial</category>
      <category>web3</category>
    </item>
    <item>
      <title>Free vs Real-Time: When the bStock Delta Actually Pays</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:16:57 +0000</pubDate>
      <link>https://dev.to/spread2009/free-vs-real-time-when-the-bstock-delta-actually-pays-3mn7</link>
      <guid>https://dev.to/spread2009/free-vs-real-time-when-the-bstock-delta-actually-pays-3mn7</guid>
      <description>&lt;p&gt;The bStock delta tracker has two tiers: free (1 request per minute) and paid (real-time, 5 USDC for 30 days). A fair question from any trader: why pay when free exists?&lt;/p&gt;

&lt;p&gt;The answer matters beyond one product. The delta is the health indicator of the whole tokenized-stocks market: the faster participants see divergence, the faster it corrects. A 24/7 market needs 24/7 monitoring — and that is a job for machines, not for humans with refresh buttons.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjlnjrcj9jgf1osgm054s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjlnjrcj9jgf1osgm054s.png" alt="Snapshot camera versus a real-time video stream — the free tier takes one picture a minute, the paid tier streams continuously" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fur1uooyflhrjqkl6fpms.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fur1uooyflhrjqkl6fpms.png" alt="Diagram: free vs paid" width="800" height="305"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The free tier is one request per minute — a snapshot for a one-off check or research. The paid tier (5 USDC for 30 days) is real-time: unlimited polling, Telegram alerts at the 0.5% threshold, and an hourly digest.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the free tier gives
&lt;/h2&gt;

&lt;p&gt;One request per minute is a &lt;strong&gt;snapshot&lt;/strong&gt;. Ask "what is AAPLB's delta" — get an answer. For research, a one-off check, a demo — enough. But between requests there is a minute of blindness. And delta lives exactly in those minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the paid tier gives
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unlimited requests&lt;/strong&gt; — poll every second if your strategy needs it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Telegram alerts&lt;/strong&gt; — the tracker messages you when a delta crosses the 0.5% threshold. No polling at all; the signal finds you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hourly digest&lt;/strong&gt; — a summary of everything that happened while you were away.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A 30-day ServicePass&lt;/strong&gt; — one payment, a month of access, proof on-chain.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  When the delta turns into money
&lt;/h2&gt;

&lt;p&gt;Three scenarios where a minute of blindness costs more than the subscription:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Weekends.&lt;/strong&gt; News drops on Saturday — the equity reacts on Monday, the token reacts now. A −3% delta on a closed exchange is advance information about Monday's gap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Earnings.&lt;/strong&gt; The report lands after the call — the stock is still in its auction, the token has already repriced.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Thin pools.&lt;/strong&gt; One large order moves the token an hour before liquidity rebalances — the window opens and closes by itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In all three, the edge is not "seeing it eventually" — it is seeing it &lt;strong&gt;first&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqyauzd4p8ljv0vi52bs5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqyauzd4p8ljv0vi52bs5.png" alt="Timeline: the delta crosses the −0.5% threshold and the alert fires within seconds — the free tier would still be blind for a whole minute" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Agent economics
&lt;/h2&gt;

&lt;p&gt;5 USDC for 30 days of real-time is ~0.17 USDC a day. One caught arbitrage window on a tokenized stock pays for years of subscription. For an agent that trades or alerts on deltas, the paid tier is not an expense — it is infrastructure, like a market data feed.&lt;/p&gt;

&lt;p&gt;And the payment itself is on-chain: 5 USDC on Arc, receipt verified in a block, ServicePass issued for 30 days. The agent's data budget is auditable on the same chain it trades on — no card statements, no billing portals, just transactions.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next — the educational block: what bStocks are, for people new to tokenized stocks.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent guide (endpoints, limits, examples): &lt;a href="https://agentbadge.xyz/bstock-guide" rel="noopener noreferrer"&gt;agentbadge.xyz/bstock-guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;All articles in the series: &lt;a href="https://agentbadge.xyz/blog" rel="noopener noreferrer"&gt;agentbadge.xyz/blog&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP endpoint: &lt;code&gt;https://agentbadge.xyz/mcp/bstock/tools/get_delta&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>cryptocurrency</category>
      <category>trading</category>
      <category>api</category>
    </item>
    <item>
      <title>x402 Payments on Arc Testnet: How an Agent Pays in USDC With No Facilitator</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Fri, 25 Sep 2026 11:15:07 +0000</pubDate>
      <link>https://dev.to/spread2009/x402-payments-on-arc-testnet-how-an-agent-pays-in-usdc-with-no-facilitator-4h4b</link>
      <guid>https://dev.to/spread2009/x402-payments-on-arc-testnet-how-an-agent-pays-in-usdc-with-no-facilitator-4h4b</guid>
      <description>&lt;p&gt;In classic x402, a &lt;strong&gt;facilitator&lt;/strong&gt; sits between the buyer and the seller: it takes the agent's signature and broadcasts the transaction on its behalf. Convenient — but an extra party to trust, an extra API to wait on, an extra point of failure.&lt;br&gt;
On Arc we skipped it. The scheme is called &lt;strong&gt;self-settle&lt;/strong&gt; — "settle it yourself".&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftten8t850x9tf0gogecs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftten8t850x9tf0gogecs.png" alt="A USDC coin flies from an agent's wallet to a treasury vault on Arc, with a trail of USDC coins paying for gas" width="800" height="447"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6pmr9x4btdyuiwz0csgg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6pmr9x4btdyuiwz0csgg.png" alt="Diagram: paying on Arc in 6 steps" width="800" height="1634"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The agent receives the 402 invoice, signs an EIP-3009 authorization (exactly 5 USDC to the treasury) and broadcasts on Arc itself — gas is paid in USDC. The server reads the receipt from the block: the Transfer event reached the treasury — access opens for 30 days.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Arc makes this possible
&lt;/h2&gt;

&lt;p&gt;Arc is a blockchain built by Circle — the company behind USDC. Its signature feature: &lt;strong&gt;gas is paid in USDC&lt;/strong&gt;, not in a separate token. A conventional agent would need to hold two assets: USDC for the payment and ETH for gas. On Arc one balance is enough — USDC covers both the payment and the fee.&lt;br&gt;
For the tokenized-stocks market this closes the loop: bStocks settle in USDC on Binance, the agent's gas is USDC, and the data subscription is USDC. One asset for trading, for fees, and for information — that is what a market built for machines looks like.&lt;/p&gt;

&lt;h2&gt;
  
  
  The flow in 6 steps
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;The agent requests data → gets a 402 with the invoice: scheme, network, amount, recipient.&lt;/li&gt;
&lt;li&gt;It signs an &lt;strong&gt;EIP-3009&lt;/strong&gt; authorization — the standard for "transfer with authorization": a signature that permits moving exactly 5 USDC from the agent's wallet to the recipient. Nothing more, no wallet access.&lt;/li&gt;
&lt;li&gt;It broadcasts the transaction itself (hence "client-broadcast").&lt;/li&gt;
&lt;li&gt;It waits for confirmation — seconds.&lt;/li&gt;
&lt;li&gt;It retries the request with the transaction hash attached.&lt;/li&gt;
&lt;li&gt;The server reads the blockchain: is the tx in a block, did USDC reach the treasury, is the amount right → access for 30 days.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frli6chug7ucs8lywa5li.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frli6chug7ucs8lywa5li.png" alt="Six-step pipeline: request, 402 invoice, EIP-3009 signature, broadcast, on-chain receipt, access for 30 days" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the server actually verifies
&lt;/h2&gt;

&lt;p&gt;Not a signature — a &lt;strong&gt;receipt&lt;/strong&gt;. The server asks the chain for &lt;code&gt;getTransactionReceipt(txHash)&lt;/code&gt; and checks: the transaction is really in a block, it contains a &lt;code&gt;Transfer&lt;/code&gt; event from the USDC contract to the treasury address, the amount covers the price. This cannot be forged: either the transaction is in a block or it does not exist.&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7k3w3f5jns3j7gjjk3lu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7k3w3f5jns3j7gjjk3lu.png" alt="A block on Arc with a highlighted Transfer log paying 5 USDC to the treasury, inspected by the server" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What this gives the ecosystem
&lt;/h2&gt;

&lt;p&gt;Removing the facilitator removes a point of failure and a trust assumption. Any wallet holding USDC on Arc becomes a payment client: one asset, one signature, one RPC call. For agents, buying data becomes as routine as calling an API.&lt;br&gt;
This is infrastructure for the tokenized-assets ecosystem, not a demo. Real-time delta data is what keeps bStock prices honest — and the agents that buy it settle on Arc in USDC, the same asset the tokens themselves settle in. Every payment is a public, auditable transaction: the market's information layer becomes as transparent as its trading layer.&lt;br&gt;
&lt;em&gt;Next: when the delta actually pays — free tier vs real-time.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent guide (endpoints, limits, examples): &lt;a href="https://agentbadge.xyz/bstock-guide" rel="noopener noreferrer"&gt;agentbadge.xyz/bstock-guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Originally published on the AgentBadge blog: &lt;a href="https://agentbadge.xyz/blog/bstock-arc-x402-payment" rel="noopener noreferrer"&gt;agentbadge.xyz/blog/bstock-arc-x402-payment&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP endpoint: &lt;code&gt;https://agentbadge.xyz/mcp/bstock/tools/get_delta&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>cryptocurrency</category>
      <category>web3</category>
      <category>api</category>
    </item>
    <item>
      <title>Freemium for AI Agents: Free vs Paid Tiers via HTTP 402 on Arc</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Thu, 24 Sep 2026 11:40:57 +0000</pubDate>
      <link>https://dev.to/spread2009/freemium-for-ai-agents-free-vs-paid-tiers-via-http-402-on-arc-23g4</link>
      <guid>https://dev.to/spread2009/freemium-for-ai-agents-free-vs-paid-tiers-via-http-402-on-arc-23g4</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published: &lt;a href="https://agentbadge.xyz/blog/bstock-freemium-402" rel="noopener noreferrer"&gt;https://agentbadge.xyz/blog/bstock-freemium-402&lt;/a&gt;&lt;br&gt;
Tags: ai, api, webdev, cryptocurrency&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every SaaS has a free tier and a paid tier. But how do you sell the paid tier when the customer is not a person but a program? An agent has no card, cannot fill a checkout form, cannot type a CVV.&lt;/p&gt;

&lt;p&gt;We solved it with &lt;strong&gt;HTTP 402 Payment Required&lt;/strong&gt; — a status code that waited three decades for its moment. It is not an error. It is an invoice.&lt;/p&gt;

&lt;p&gt;Why does this matter beyond our tracker? Tokenized stocks trade around the clock, and the traders who watch them increasingly delegate monitoring to AI agents. An agent that cannot pay for its own data is a crippled market participant. Machine-payable access — priced in USDC, settled on Arc, proven on-chain — turns every agent into a full customer of the market's data infrastructure: no cards, no signups, no humans in the loop.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffy8yvzm74uj32l582ig9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffy8yvzm74uj32l582ig9.png" alt="A robot inserts a glowing USDC coin into a turnstile marked 402" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxmt48616x3og734t81xi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxmt48616x3og734t81xi.png" alt="Diagram: the freemium gate" width="800" height="1794"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Every request passes three doors: a live ServicePass means instant real-time; otherwise the free bucket allows one request per minute (a snapshot); otherwise the server answers 402 with an invoice. One USDC transaction (5 USDC on Arc), an on-chain receipt check — and the agent holds a 30-day ServicePass.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How the gate works
&lt;/h2&gt;

&lt;p&gt;Every request to the tracker passes three checks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Already paid?&lt;/strong&gt; If the agent holds a live ServicePass (a 30-day access token) — data flows immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Free allowance?&lt;/strong&gt; One request per minute is free. A snapshot, not a stream.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Neither?&lt;/strong&gt; The server answers 402 and attaches an invoice to the response: what, to whom, how much.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Inside the invoice
&lt;/h2&gt;

&lt;p&gt;The 402 response is not just text. The &lt;code&gt;PAYMENT-REQUIRED&lt;/code&gt; header carries a machine-readable payment description:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"x402Version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"accepts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"scheme"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eip3009-client-broadcast"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"network"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eip155:5042002"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"asset"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"USDC"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"amount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"5000000"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"payTo"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0xcdd2...699d"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxTimeoutSeconds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;345600&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9aiag1q06xfqh4r469gh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9aiag1q06xfqh4r469gh.png" alt="A 402 Payment Required JSON response with the accepts block highlighted" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The agent reads it like a price list: the payment scheme (EIP-3009 — a standardized transfer signature), the network (Arc Testnet), the asset (USDC), the amount (5 USDC — the six zeros are decimals, the amount is in base units), the recipient, and the invoice expiry (4 days).&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not API keys and billing portals
&lt;/h2&gt;

&lt;p&gt;An API key needs signup, an email, a card, invoices — a human in the loop. 402 + x402 is a &lt;em&gt;programmable&lt;/em&gt; paywall: the agent sees the price → signs a transfer → pays → gets access. The whole cycle takes seconds. For a machine, this is the native way to buy things.&lt;/p&gt;

&lt;p&gt;And because it runs on &lt;strong&gt;Arc&lt;/strong&gt; — Circle's blockchain where gas itself is paid in USDC — the agent needs exactly one asset in its wallet. One balance covers the payment and the fee. For a market where bStocks themselves settle in USDC, the plumbing finally matches the asset.&lt;/p&gt;

&lt;h2&gt;
  
  
  One payment pays once
&lt;/h2&gt;

&lt;p&gt;After paying, the agent retries the request with the transaction hash attached. The server checks the blockchain — not a claimed signature, but the real receipt from a block — and opens access. The same hash cannot be presented twice: the server atomically claims it, and a second attempt gets &lt;code&gt;tx_replayed&lt;/code&gt;. Even ten parallel requests with the same hash — exactly one gets through.&lt;/p&gt;

&lt;p&gt;A paywall a program can read and pay turns data into a first-class on-chain service. More agents able to buy real-time data means more eyes on tokenized markets — tighter deltas, faster convergence, healthier price discovery for everyone who trades them.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next: the payment itself on Arc — and why gas there is paid in USDC.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent guide (endpoints, limits, examples): &lt;a href="https://agentbadge.xyz/bstock-guide" rel="noopener noreferrer"&gt;agentbadge.xyz/bstock-guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Originally published on the AgentBadge blog: &lt;a href="https://agentbadge.xyz/blog/bstock-freemium-402" rel="noopener noreferrer"&gt;agentbadge.xyz/blog/bstock-freemium-402&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP endpoint: &lt;code&gt;https://agentbadge.xyz/mcp/bstock/tools/get_delta&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>webdev</category>
      <category>cryptocurrency</category>
    </item>
    <item>
      <title>Tracking the Delta: How an AI Agent Watches Tokenized Stocks Drift From Their Underlyings</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Wed, 23 Sep 2026 19:13:29 +0000</pubDate>
      <link>https://dev.to/spread2009/tracking-the-delta-how-an-ai-agent-watches-tokenized-stocks-drift-from-their-underlyings-411m</link>
      <guid>https://dev.to/spread2009/tracking-the-delta-how-an-ai-agent-watches-tokenized-stocks-drift-from-their-underlyings-411m</guid>
      <description>&lt;p&gt;Picture this: Apple trades at $336.45 on Nasdaq. Tokenized Apple (AAPLB on Binance) trades at $336.44 at the same moment. A penny apart — noise. But on a Saturday night, with Nasdaq asleep, AAPLB can drift to −2%. That is not noise anymore. That is a signal.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The delta&lt;/strong&gt; — the gap between a tokenized stock's price and its underlying — is where the opportunities live. Humans cannot watch 20+ tokens around the clock. A machine can. So we built a tracker that computes the delta in real time and handed it to AI agents.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgraoaowb9levf2ek3tqz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgraoaowb9levf2ek3tqz.png" alt="Two price lines - AAPLB token vs AAPL underlying - with the -2% delta band highlighted" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsk5f70ebedmdzuaop1cf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsk5f70ebedmdzuaop1cf.png" alt="Diagram: how the delta tracker works" width="800" height="843"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Binance streams token prices around the clock; Finnhub/Alpaca supply the equity prices. DeltaEngine computes the delta, the market phase (O/C/P) and the stale flag (15s without updates). The MCP server exposes it to the agent through five tools, and Telegram receives an alert the moment a delta crosses the 0.5% threshold.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the tracker computes
&lt;/h2&gt;

&lt;p&gt;For every bStock the tracker keeps two prices:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;bStockPrice&lt;/strong&gt; — the token's price on Binance, where tokenized stocks trade 24/7;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;underlyingPrice&lt;/strong&gt; — the real stock's price from market data providers (Finnhub, with an Alpaca fallback).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Delta in percent: &lt;code&gt;(tokenPrice − stockPrice) / stockPrice × 100&lt;/code&gt;. Alongside it, three flags a trader actually needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;phase&lt;/strong&gt; — is the stock market open (&lt;code&gt;O&lt;/code&gt;), closed (&lt;code&gt;C&lt;/code&gt;), or in pre-market (&lt;code&gt;P&lt;/code&gt;);&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;stale&lt;/strong&gt; — the price feed went quiet for more than 15 seconds, so treat the number with care;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;inAlert&lt;/strong&gt; — the delta crossed the &lt;strong&gt;0.5% threshold&lt;/strong&gt; (configurable).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A live response looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"symbol"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"AAPLB"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"underlying"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"AAPL"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"multiplier"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;1.0006&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"bStockPrice"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;336.44&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"underlyingPrice"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;336.45&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"deltaPct"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;-0.064&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"phase"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"O"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"stale"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"inAlert"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgurvedvpk7m02a5wqeb4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgurvedvpk7m02a5wqeb4.png" alt="get_delta JSON response with deltaPct and inAlert highlighted" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How an agent uses it
&lt;/h2&gt;

&lt;p&gt;An AI agent is a program — Claude, a GPT-based bot, your own script — that calls services on its own. Agents talk to services over &lt;strong&gt;MCP&lt;/strong&gt;, an open protocol where a service lists its tools and the agent calls them like functions.&lt;/p&gt;

&lt;p&gt;The tracker is an MCP server with these tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;get_delta&lt;/code&gt; — delta for one symbol;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;list_deltas&lt;/code&gt; — all symbols at once;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;get_quote&lt;/code&gt; — current quote;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;get_events&lt;/code&gt; — event history (threshold crossings, stale feeds);&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;get_digest&lt;/code&gt; — the day's summary.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST https://agentbadge.xyz/mcp/bstock/tools/get_delta &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$TOKEN&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"symbol": "AAPLB"}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is also Telegram: &lt;code&gt;subscribe_telegram&lt;/code&gt; registers the agent's operator, and the tracker pushes a message whenever a delta crosses 0.5%. No polling needed — the signal finds you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frgpg4jghi5v9mx6yc4j7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frgpg4jghi5v9mx6yc4j7.png" alt="AI agent connected to the MCP tracker (Binance + Finnhub feeds, five tools) with a Telegram alert" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Freemium: a snapshot for free, the stream for a fee
&lt;/h2&gt;

&lt;p&gt;The first request each minute is free. After that the server answers with &lt;strong&gt;HTTP 402 Payment Required&lt;/strong&gt; — not an error, an invoice: "want real-time? pay". The x402 standard lets the agent pay automatically: 5 USDC on the Arc network buys 30 days of access. No signup, no card, one on-chain transaction — gas paid in USDC too, an Arc specialty. The server verifies the transaction on-chain and opens the door. The same payment cannot be replayed — replay protection is built in.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyyl9aczsdxezff36mo2e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyyl9aczsdxezff36mo2e.png" alt="402 paywall flow: request rejected, USDC payment to the Arc vault, 30-day access badge" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a trader should care
&lt;/h2&gt;

&lt;p&gt;A −0.06% delta is noise. A −2% delta on a closed exchange means "the token was sold off and the equity has not woken up yet". Whoever sees it first captures the convergence. A person cannot monitor every token every second — an agent can, and it pays for its own data feed without a human in the loop.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Next: inside the 402 paywall — how freemium works when the customer is a machine.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Links&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent guide (endpoints, limits, examples): &lt;a href="https://agentbadge.xyz/bstock-guide" rel="noopener noreferrer"&gt;agentbadge.xyz/bstock-guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Originally published on the AgentBadge blog: &lt;a href="https://agentbadge.xyz/blog/bstock-delta-tracker-case" rel="noopener noreferrer"&gt;agentbadge.xyz/blog/bstock-delta-tracker-case&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;MCP endpoint: &lt;code&gt;https://agentbadge.xyz/mcp/bstock/tools/get_delta&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>webdev</category>
      <category>cryptocurrency</category>
    </item>
    <item>
      <title>Inside an Agent Readiness Scanner: Rules, Evidence and Reproducibility</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Fri, 28 Aug 2026 12:09:49 +0000</pubDate>
      <link>https://dev.to/spread2009/inside-an-agent-readiness-scanner-rules-evidence-and-reproducibility-kco</link>
      <guid>https://dev.to/spread2009/inside-an-agent-readiness-scanner-rules-evidence-and-reproducibility-kco</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published at &lt;a href="https://agentbadge.xyz/blog/inside-an-agent-readiness-scanner" rel="noopener noreferrer"&gt;AgentBadge&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When someone tells you that an API has an &lt;strong&gt;87/100 Agent Readiness score&lt;/strong&gt;, the first question should not be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is 87 a good score?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Why is it 87?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the question after that is even more important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Can I reproduce the result myself?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the problem AgentBadge is designed to solve.&lt;/p&gt;

&lt;p&gt;Agent Readiness should not be an opinion generated by an LLM. It should be a measurable property of a service, calculated from explicit rules and supported by evidence.&lt;/p&gt;

&lt;p&gt;The core idea is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rules → Evidence → Assertions → Score → Report&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This article explains what happens inside that pipeline.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ugm36e7rdmwjgc0hxyw.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ugm36e7rdmwjgc0hxyw.webp" alt="Hero — Measurement pipeline diagram" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  1. A scanner should measure, not guess
&lt;/h2&gt;

&lt;p&gt;Imagine two tools scanning the same API.&lt;/p&gt;

&lt;p&gt;Tool A says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Your API appears to be highly suitable for AI agents."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Tool B says:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"AB-004 passed because &lt;code&gt;https://example.com/openapi.json&lt;/code&gt; returned HTTP 200 and contained a valid OpenAPI document."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Which result would you trust?&lt;/p&gt;

&lt;p&gt;The second one is much more useful.&lt;/p&gt;

&lt;p&gt;It tells you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what was checked;&lt;/li&gt;
&lt;li&gt;what rule was applied;&lt;/li&gt;
&lt;li&gt;what evidence was found;&lt;/li&gt;
&lt;li&gt;why the rule passed or failed;&lt;/li&gt;
&lt;li&gt;and where the evidence came from.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the fundamental design principle behind AgentBadge:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Every meaningful score should be explainable through evidence.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The scanner should not ask an AI model:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"How agent-ready does this API feel?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It should ask deterministic questions such as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Does this URL exist?"&lt;/p&gt;

&lt;p&gt;"Does it return the expected content type?"&lt;/p&gt;

&lt;p&gt;"Does the response contain an OpenAPI document?"&lt;/p&gt;

&lt;p&gt;"Does the declared authentication mechanism contain the required information?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The difference may look subtle, but architecturally it is enormous.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Rules are the measurement instrument
&lt;/h2&gt;

&lt;p&gt;A scanner is only as trustworthy as its rules.&lt;/p&gt;

&lt;p&gt;Instead of hiding the evaluation logic inside application code, AgentBadge treats rules as explicit, versioned measurement definitions.&lt;/p&gt;

&lt;p&gt;A simplified rule might look conceptually like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AB-001
Name: OpenAPI discoverability

Given:
  target = https://example.com

Check:
  GET /.well-known/openapi.json

Pass when:
  HTTP status = 200
  AND response is valid OpenAPI

Evidence:
  URL
  HTTP status
  content type
  content hash

Severity:
  medium
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important property is that another implementation should be able to understand the same rule.&lt;/p&gt;

&lt;p&gt;The rule is not:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The API looks well documented."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This specific machine-readable artifact was found and passed these specific checks."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That makes the scanner much easier to test, audit and reproduce.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Deterministic before intelligent
&lt;/h2&gt;

&lt;p&gt;This leads to one of the most important architectural principles of AgentBadge:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Deterministic before intelligent.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If something can be established deterministically, use deterministic logic.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Question&lt;/th&gt;
&lt;th&gt;Preferred method&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Does robots.txt exist?&lt;/td&gt;
&lt;td&gt;HTTP request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does sitemap exist?&lt;/td&gt;
&lt;td&gt;HTTP request + parser&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does OpenAPI exist?&lt;/td&gt;
&lt;td&gt;HTTP request + schema validation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Is JSON valid?&lt;/td&gt;
&lt;td&gt;JSON parser&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Does declared endpoint exist in another document?&lt;/td&gt;
&lt;td&gt;Exact matching&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What does an undocumented endpoint mean?&lt;/td&gt;
&lt;td&gt;AI-assisted inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What does an API capability actually mean?&lt;/td&gt;
&lt;td&gt;Human confirmation / assisted review&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;AI has a role, but it should not become the judge of facts that can be verified directly.&lt;/p&gt;

&lt;p&gt;An LLM can help interpret ambiguous documentation.&lt;/p&gt;

&lt;p&gt;It should not silently decide that an API supports refunds simply because a paragraph mentions the word "refund."&lt;/p&gt;

&lt;p&gt;This is why AgentBadge treats AI as a &lt;strong&gt;copilot&lt;/strong&gt;, not as the authority responsible for the score.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Evidence is the missing layer
&lt;/h2&gt;

&lt;p&gt;A score without evidence is difficult to trust.&lt;/p&gt;

&lt;p&gt;Consider this finding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Documentation: 18/25
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is almost nothing you can do with it.&lt;/p&gt;

&lt;p&gt;Now consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AB-007  OpenAPI discoverability

STATUS: VERIFIED

Evidence:
GET https://api.example.com/openapi.json
HTTP 200
Content-Type: application/json

OpenAPI version:
3.1.0

Confidence:
1.00
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the developer knows what happened.&lt;/p&gt;

&lt;p&gt;They can inspect the same resource themselves.&lt;/p&gt;

&lt;p&gt;This creates a chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTTP response → Evidence → Assertion → Rule result → Category score → Overall score
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The evidence is therefore not an optional explanation attached to the report.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evidence is part of the measurement itself.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Assertions connect evidence and scoring
&lt;/h2&gt;

&lt;p&gt;A useful internal abstraction is an &lt;strong&gt;assertion&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;An assertion answers one concrete question.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rule_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"AB-007"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"VERIFIED"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com/openapi.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"http_status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"content_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"application/json"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important states are intentionally explicit.&lt;/p&gt;

&lt;h3&gt;
  
  
  VERIFIED
&lt;/h3&gt;

&lt;p&gt;The scanner found direct evidence supporting the assertion.&lt;/p&gt;

&lt;h3&gt;
  
  
  INFERRED
&lt;/h3&gt;

&lt;p&gt;The scanner has a reasonable interpretation, but the evidence is not sufficient to treat it as fact.&lt;/p&gt;

&lt;h3&gt;
  
  
  CONFLICT
&lt;/h3&gt;

&lt;p&gt;Two sources disagree.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Guide says: POST /refund
OpenAPI says: POST /refund-request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  MISSING
&lt;/h3&gt;

&lt;p&gt;The expected capability or artifact could not be found.&lt;/p&gt;

&lt;p&gt;These states are more informative than a simple pass/fail system.&lt;/p&gt;

&lt;p&gt;They tell us not only &lt;strong&gt;what the scanner thinks&lt;/strong&gt;, but also &lt;strong&gt;how strongly it knows it&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Confidence is not the same as verification
&lt;/h2&gt;

&lt;p&gt;This distinction is important.&lt;/p&gt;

&lt;p&gt;A scanner may infer something with high confidence.&lt;/p&gt;

&lt;p&gt;That does not automatically make it verified.&lt;/p&gt;

&lt;p&gt;For example, an API documentation page might strongly suggest that a service supports refunds.&lt;/p&gt;

&lt;p&gt;An LLM may assign a confidence of 0.94 to that interpretation.&lt;/p&gt;

&lt;p&gt;But unless there is machine-readable evidence supporting the capability, the assertion should not magically become VERIFIED.&lt;/p&gt;

&lt;p&gt;Instead: INFERRED, confidence: 0.94&lt;/p&gt;

&lt;p&gt;The user can then: Confirm, Edit, Reject&lt;/p&gt;

&lt;p&gt;This is the boundary between &lt;strong&gt;automatic fixes&lt;/strong&gt; and &lt;strong&gt;assisted fixes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Safe, deterministic changes can be automated.&lt;/p&gt;

&lt;p&gt;Semantic claims require human confirmation.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Reproducibility matters
&lt;/h2&gt;

&lt;p&gt;Now we arrive at the second major property of the scanner.&lt;/p&gt;

&lt;p&gt;Suppose you scan &lt;code&gt;https://api.example.com&lt;/code&gt; today and receive 76/100.&lt;/p&gt;

&lt;p&gt;Someone else runs the same rules against the same captured state and should be able to understand how the result was produced.&lt;/p&gt;

&lt;p&gt;That requires more than storing the final number.&lt;/p&gt;

&lt;p&gt;The report needs to describe the measurement context.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ruleset"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent-readiness-v1.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"scanner_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.1.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"assertions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;76&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"categories"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"discovery"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"documentation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"authentication"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"machine_readability"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This allows the score to be understood as the output of a defined measurement process rather than a mysterious number.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb7uebfeiiqqyz0e2t974.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb7uebfeiiqqyz0e2t974.webp" alt="Reproducibility infographic" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Rules must be versioned
&lt;/h2&gt;

&lt;p&gt;Rules change. New standards appear. New machine-readable formats emerge.&lt;/p&gt;

&lt;p&gt;Some checks eventually turn out to be too strict or too weak.&lt;/p&gt;

&lt;p&gt;Therefore: Agent Readiness v1.0 must not silently become v1.1 while pretending the results are identical.&lt;/p&gt;

&lt;p&gt;Each ruleset should have an explicit version.&lt;/p&gt;

&lt;p&gt;Now a report can say: Score: 82/100, Ruleset: Agent Readiness v1.2&lt;/p&gt;

&lt;p&gt;This gives us an important property:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Same target + same measurement state + same ruleset = reproducible result.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  9. Why open rules do not destroy the product
&lt;/h2&gt;

&lt;p&gt;If AgentBadge publishes its rules, can't someone simply copy them?&lt;/p&gt;

&lt;p&gt;Yes. And that is intentional.&lt;/p&gt;

&lt;p&gt;The goal is not to create a secret scoring algorithm. The goal is to establish a useful measurement standard.&lt;/p&gt;

&lt;p&gt;The long-term value comes from the workflow around that standard: open specification, open scanner, GitHub Action, README badge, continuous monitoring, regression alerts, fix workflow, developer adoption.&lt;/p&gt;

&lt;p&gt;A competitor can copy AB-001, AB-002, AB-003.&lt;/p&gt;

&lt;p&gt;They cannot instantly copy: thousands of repositories displaying the badge, existing GitHub Actions, developer habits, historical scan data, integrations, workflow configuration, trust built around independently verifiable reports.&lt;/p&gt;

&lt;p&gt;The moat is not secret rules. It is a standard installed inside the developer workflow.&lt;/p&gt;




&lt;h2&gt;
  
  
  10. The score should explain itself
&lt;/h2&gt;

&lt;p&gt;A single number is useful for quick comparison. But it should never be the only information available.&lt;/p&gt;

&lt;p&gt;Suppose a developer's score changes: 76 to 72.&lt;/p&gt;

&lt;p&gt;The product should explain the delta: +8 OpenAPI documentation detected, -12 New authentication issue detected, +0 Discovery unchanged. Result: 76 to 72.&lt;/p&gt;

&lt;p&gt;This turns measurement into an improvement loop.&lt;/p&gt;




&lt;h2&gt;
  
  
  11. From Measure to Prove to Improve
&lt;/h2&gt;

&lt;p&gt;The architecture becomes a simple loop: Measure, Prove, Improve, Measure again.&lt;/p&gt;

&lt;p&gt;The score is the beginning of the workflow, not the end.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgzp7g5lfqabtz9zoj1x.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgzp7g5lfqabtz9zoj1x.webp" alt="Measure Prove Improve cycle diagram" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  12. What AgentBadge should never claim
&lt;/h2&gt;

&lt;p&gt;AgentBadge measures &lt;strong&gt;Agent Readiness&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It does not certify: API security, business correctness, service reliability, legal compliance, quality of business logic, whether an agent should trust the company.&lt;/p&gt;

&lt;p&gt;A high score does not mean "This API is safe." It means "This API satisfied these measurable Agent Readiness criteria."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't certify. Measure.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  13. What this enables
&lt;/h2&gt;

&lt;p&gt;Once the measurement layer exists, many higher-level products become possible.&lt;/p&gt;

&lt;p&gt;A developer can run &lt;code&gt;npx @agentbadge/cli scan https://api.example.com&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A CI pipeline can enforce a minimum score. A README can display the current measurement. A platform can query AgentBadge programmatically. An organization can compare vendors using the same ruleset.&lt;/p&gt;




&lt;h2&gt;
  
  
  14. The bigger idea
&lt;/h2&gt;

&lt;p&gt;The web has spent years developing tools for measuring websites. Performance has metrics. Accessibility has automated checks. Security has scanners. TLS has analyzers. SEO has crawlers and validators.&lt;/p&gt;

&lt;p&gt;The emerging agentic web needs something similar.&lt;/p&gt;

&lt;p&gt;AgentBadge's approach is deliberately conservative: Define the rules. Collect the evidence. Show the reasoning. Version the rules. Make the result reproducible.&lt;/p&gt;

&lt;p&gt;That is what makes Agent Readiness a measurement discipline rather than another AI-generated checklist.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Read more:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;What Is Agent Readiness?&lt;/a&gt; — Article 1&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-ai-agents-fail-to-use-apis" rel="noopener noreferrer"&gt;Why AI Agents Fail to Use APIs&lt;/a&gt; — Article 5&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-ai-agent-needs-to-understand-api" rel="noopener noreferrer"&gt;What Does an AI Agent Need to Understand an API?&lt;/a&gt; — Article 6&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-openapi-isnt-enough" rel="noopener noreferrer"&gt;Why Your OpenAPI Spec Isn't Enough&lt;/a&gt; — Article 7&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/how-do-you-measure-agent-readiness" rel="noopener noreferrer"&gt;How Do You Measure Agent Readiness?&lt;/a&gt; — Article 8&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Agent Knowledge Layer:&lt;/strong&gt; &lt;a href="https://agentbadge.xyz/agent-guide/" rel="noopener noreferrer"&gt;agentbadge.xyz/agent-guide&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>webdev</category>
      <category>testing</category>
    </item>
    <item>
      <title>Inside an Agent Readiness Scanner: Rules, Evidence and Reproducibility</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Fri, 28 Aug 2026 06:26:46 +0000</pubDate>
      <link>https://dev.to/spread2009/inside-an-agent-readiness-scanner-rules-evidence-and-reproducibility-3j9k</link>
      <guid>https://dev.to/spread2009/inside-an-agent-readiness-scanner-rules-evidence-and-reproducibility-3j9k</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;Originally published at &lt;a href="https://agentbadge.xyz/blog/inside-an-agent-readiness-scanner" rel="noopener noreferrer"&gt;AgentBadge&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When someone tells you that an API has an &lt;strong&gt;87/100 Agent Readiness score&lt;/strong&gt;, the first question should not be:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Is 87 a good score?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Why is it 87?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the question after that is even more important:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Can I reproduce the result myself?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the problem AgentBadge is designed to solve.&lt;/p&gt;

&lt;p&gt;Agent Readiness should not be an opinion generated by an LLM. It should be a measurable property of a service, calculated from explicit rules and supported by evidence.&lt;/p&gt;

&lt;p&gt;The core idea is simple:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rules → Evidence → Assertions → Score → Report&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  1. What the scanner actually does
&lt;/h2&gt;

&lt;p&gt;The scanner does not ask an LLM "is this API good?"&lt;/p&gt;

&lt;p&gt;It runs deterministic checks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does &lt;code&gt;robots.txt&lt;/code&gt; exist?&lt;/li&gt;
&lt;li&gt;Does &lt;code&gt;/.well-known/openapi.json&lt;/code&gt; return 200?&lt;/li&gt;
&lt;li&gt;Is the response valid OpenAPI?&lt;/li&gt;
&lt;li&gt;Does the homepage have machine-readable metadata?&lt;/li&gt;
&lt;li&gt;Is there an &lt;code&gt;agents.txt&lt;/code&gt; file?&lt;/li&gt;
&lt;li&gt;Are there &lt;code&gt;llms.txt&lt;/code&gt; or &lt;code&gt;llms-full.txt&lt;/code&gt; files?&lt;/li&gt;
&lt;li&gt;Does the API support content negotiation?&lt;/li&gt;
&lt;li&gt;Is there an A2A agent card?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each check produces evidence. Each piece of evidence becomes an assertion. Each assertion has a status.&lt;/p&gt;

&lt;p&gt;AI-assisted inference is used when deterministic checks are insufficient — but it is a copilot, not the judge.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Rule structure
&lt;/h2&gt;

&lt;p&gt;Every rule follows the same structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AB-001
Name: OpenAPI discoverability

Given:
  target = https://example.com

Check:
  GET /.well-known/openapi.json

Pass when:
  HTTP status = 200
  AND response is valid OpenAPI

Evidence:
  URL
  HTTP status
  content type
  content hash

Severity:
  medium
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not a prompt. It is a specification.&lt;/p&gt;

&lt;p&gt;The rule says what to check, what constitutes a pass, and what evidence to collect.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Assertion states
&lt;/h2&gt;

&lt;p&gt;Every assertion has exactly one of four states:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;VERIFIED&lt;/td&gt;
&lt;td&gt;Direct evidence found&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INFERRED&lt;/td&gt;
&lt;td&gt;Reasonable interpretation, insufficient evidence&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CONFLICT&lt;/td&gt;
&lt;td&gt;Two sources disagree&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MISSING&lt;/td&gt;
&lt;td&gt;Expected capability or artifact not found&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An LLM can be 94% confident that an API supports refunds. Without machine-readable evidence, the assertion stays INFERRED.&lt;/p&gt;

&lt;p&gt;Confidence is not the same as verification.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ugm36e7rdmwjgc0hxyw.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5ugm36e7rdmwjgc0hxyw.webp" alt="Assertion states infographic — four states from verified to missing" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Evidence is part of the measurement
&lt;/h2&gt;

&lt;p&gt;Evidence is not an optional explanation attached after the fact.&lt;/p&gt;

&lt;p&gt;Each finding includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The actual HTTP response (where applicable)&lt;/li&gt;
&lt;li&gt;The URL checked&lt;/li&gt;
&lt;li&gt;The content type received&lt;/li&gt;
&lt;li&gt;A content hash&lt;/li&gt;
&lt;li&gt;The timestamp of the check&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This means the assertion is not just "OpenAPI exists: yes" — it is "OpenAPI exists: yes, here is the proof."&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rule_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"AB-007"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"VERIFIED"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com/openapi.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"evidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"http_status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"content_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"application/json"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  5. Scoring has category floors
&lt;/h2&gt;

&lt;p&gt;The total score is not a simple average.&lt;/p&gt;

&lt;p&gt;Some categories are foundational. If &lt;code&gt;Discovery = 0&lt;/code&gt;, it does not matter how good the documentation is — agents cannot find the API.&lt;/p&gt;

&lt;p&gt;Category floors prevent a high score from hiding a critical zero.&lt;/p&gt;

&lt;p&gt;A score of 91/100 with &lt;code&gt;Discovery = 0&lt;/code&gt; is not a good score. It is a misleading one.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. What confidence means and what it does not
&lt;/h2&gt;

&lt;p&gt;An LLM can analyze an API documentation page and say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This API likely supports token-based authentication with rate limiting."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a reasonable inference. But without machine-readable evidence (an OpenAPI spec, a response header, a well-known endpoint), it remains INFERRED.&lt;/p&gt;

&lt;p&gt;The user can then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Confirm&lt;/strong&gt; — "Yes, this is correct, I checked manually"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Edit&lt;/strong&gt; — "Almost right, but the auth method is different"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reject&lt;/strong&gt; — "No, this is wrong"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This creates a feedback loop: automatic fixes for deterministic changes, human confirmation for semantic claims.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Reproducibility
&lt;/h2&gt;

&lt;p&gt;If the scanner gives a score of 76/100 for &lt;code&gt;https://api.example.com&lt;/code&gt; today, someone else running the same rules against the same captured state should be able to understand how the result was produced.&lt;/p&gt;

&lt;p&gt;The report needs to describe the measurement context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"target"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ruleset"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent-readiness-v1.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"scanner_version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.1.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"assertions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="err"&gt;...&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;76&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"categories"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"discovery"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"documentation"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"authentication"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;17&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"machine_readability"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;21&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb7uebfeiiqqyz0e2t974.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb7uebfeiiqqyz0e2t974.webp" alt="Reproducibility infographic — three identical inputs converging into the same result" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Rules must be versioned
&lt;/h2&gt;

&lt;p&gt;Rules change. New standards appear. Some checks eventually turn out to be too strict or too weak.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Agent Readiness v1.0&lt;/code&gt; must not silently become &lt;code&gt;v1.1&lt;/code&gt; while pretending the results are identical.&lt;/p&gt;

&lt;p&gt;Each ruleset has an explicit version. A report can say:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Score: 82/100&lt;br&gt;&lt;br&gt;
Ruleset: Agent Readiness v1.2&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Same target + same measurement state + same ruleset = reproducible result.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  9. Why open rules do not destroy the product
&lt;/h2&gt;

&lt;p&gt;If AgentBadge publishes its rules, can't someone simply copy them?&lt;/p&gt;

&lt;p&gt;Yes. And that is intentional.&lt;/p&gt;

&lt;p&gt;The goal is not a secret scoring algorithm. The goal is a useful measurement standard.&lt;/p&gt;

&lt;p&gt;The long-term value comes from the workflow around that standard:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Open specification → Open scanner → GitHub Action → README badge → Continuous monitoring → Regression alerts → Fix workflow → Developer adoption
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A competitor can copy rules. They cannot instantly copy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;thousands of repositories displaying the badge&lt;/li&gt;
&lt;li&gt;existing GitHub Actions&lt;/li&gt;
&lt;li&gt;developer habits&lt;/li&gt;
&lt;li&gt;historical scan data&lt;/li&gt;
&lt;li&gt;integrations&lt;/li&gt;
&lt;li&gt;trust built around independently verifiable reports&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The moat is not "our rules are secret." It is "our standard is installed inside the developer workflow."&lt;/p&gt;




&lt;h2&gt;
  
  
  10. The score should explain itself
&lt;/h2&gt;

&lt;p&gt;If a developer's score changes from 76 → 72, the product should explain the delta:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+8  OpenAPI documentation detected
-12  New authentication issue detected
+0   Discovery unchanged

Result: 76 → 72
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without a breakdown, a developer who fixed three problems and still sees the score decrease will think the product is broken.&lt;/p&gt;

&lt;p&gt;With a breakdown, it becomes: "The scanner found something new. Now I know what to fix."&lt;/p&gt;




&lt;h2&gt;
  
  
  11. From Measure to Prove to Improve
&lt;/h2&gt;

&lt;p&gt;The architecture is a loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MEASURE → PROVE → IMPROVE → Measure again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The score is the beginning of the workflow, not the end.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgzp7g5lfqabtz9zoj1x.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqgzp7g5lfqabtz9zoj1x.webp" alt="Measure → Prove → Improve cycle diagram" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  12. What AgentBadge should never claim
&lt;/h2&gt;

&lt;p&gt;AgentBadge measures &lt;strong&gt;Agent Readiness&lt;/strong&gt;. It does not certify:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API security&lt;/li&gt;
&lt;li&gt;business correctness&lt;/li&gt;
&lt;li&gt;service reliability&lt;/li&gt;
&lt;li&gt;legal compliance&lt;/li&gt;
&lt;li&gt;whether an agent should trust the company&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A high score means: "This API satisfied these measurable Agent Readiness criteria."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't certify. Measure.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  13. What this enables
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @agentbadge/cli scan https://api.example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A CI pipeline can enforce a minimum score. A README can display the current measurement. A platform can query AgentBadge programmatically. An organization can compare vendors using the same ruleset.&lt;/p&gt;




&lt;h2&gt;
  
  
  14. The bigger idea
&lt;/h2&gt;

&lt;p&gt;Performance has metrics. Accessibility has automated checks. Security has scanners. TLS has analyzers. SEO has crawlers and validators.&lt;/p&gt;

&lt;p&gt;The emerging agentic web needs something similar.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Don't ask an AI to invent a score.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Define the rules. Collect the evidence. Show the reasoning. Version the rules. Make the result reproducible.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Then let developers improve their systems.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Read more:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;What Is Agent Readiness?&lt;/a&gt; — Article 1&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-ai-agents-fail-to-use-apis" rel="noopener noreferrer"&gt;Why AI Agents Fail to Use APIs&lt;/a&gt; — Article 5&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-ai-agent-needs-to-understand-api" rel="noopener noreferrer"&gt;What Does an AI Agent Need to Understand an API?&lt;/a&gt; — Article 6&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-openapi-isnt-enough" rel="noopener noreferrer"&gt;Why Your OpenAPI Spec Isn't Enough&lt;/a&gt; — Article 7&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/how-do-you-measure-agent-readiness" rel="noopener noreferrer"&gt;How Do You Measure Agent Readiness?&lt;/a&gt; — Article 8&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Agent Knowledge Layer:&lt;/strong&gt; &lt;a href="https://agentbadge.xyz/agent-guide/" rel="noopener noreferrer"&gt;agentbadge.xyz/agent-guide&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>webdev</category>
      <category>testing</category>
    </item>
    <item>
      <title>How Do You Measure Agent Readiness?</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Wed, 26 Aug 2026 13:11:26 +0000</pubDate>
      <link>https://dev.to/spread2009/how-do-you-measure-agent-readiness-3328</link>
      <guid>https://dev.to/spread2009/how-do-you-measure-agent-readiness-3328</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;If Agent Readiness is real, it should be measurable. And the measurement should be reproducible.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You've read about &lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;what Agent Readiness is&lt;/a&gt;. You've seen &lt;a href="https://agentbadge.xyz/blog/why-ai-agents-fail-to-use-apis" rel="noopener noreferrer"&gt;why AI agents fail to use APIs&lt;/a&gt; and &lt;a href="https://agentbadge.xyz/blog/what-ai-agent-needs-to-understand-api" rel="noopener noreferrer"&gt;what an agent needs to understand&lt;/a&gt;. You know &lt;a href="https://agentbadge.xyz/blog/why-openapi-isnt-enough" rel="noopener noreferrer"&gt;why OpenAPI alone isn't enough&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Now the question shifts from "what" to "how":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do you objectively determine whether an API is ready for AI agents?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This article introduces a measurement framework for Agent Readiness — one built on deterministic checks, evidence, and reproducibility. Not opinions. Not LLM scores. Measurable properties that any scanner can verify.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qxzawufp8rpo3gfxkwf.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5qxzawufp8rpo3gfxkwf.webp" alt="Hero — Subjective labels on the left (" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Measurement Problem
&lt;/h2&gt;

&lt;p&gt;Labels like "AI-friendly API", "Agent-ready", and "Optimized for AI" are everywhere. They sound useful. They aren't.&lt;/p&gt;

&lt;p&gt;Two auditors can look at the same API and disagree on whether it's "agent-friendly." An LLM can score the same API differently on different runs. A marketing page can claim "AI-optimized" without any way to verify what that means.&lt;/p&gt;

&lt;p&gt;The problem isn't that these labels are wrong. The problem is that they're &lt;strong&gt;not reproducible&lt;/strong&gt;. If two people can look at the same API and reach different conclusions, the measurement isn't real — it's an opinion.&lt;/p&gt;

&lt;p&gt;If Agent Readiness is a real property of an API, it should be measurable. And the measurement should satisfy a simple requirement:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;same URL + same ruleset + same point in time = same result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the reproducibility requirement. It's what separates measurement from opinion.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Should We Measure?
&lt;/h2&gt;

&lt;p&gt;Agent Readiness isn't a single number. It's a set of properties across four categories:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Discovery&lt;/strong&gt; — Can an agent find the API?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Documentation&lt;/strong&gt; — Can an agent understand the API?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication&lt;/strong&gt; — Can an agent authenticate autonomously?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Machine Readability&lt;/strong&gt; — Can an agent interact machine-to-machine?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But these aren't just checkboxes. Each category contains specific, testable assertions — properties that can be verified with HTTP requests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Discovery
  ✓ OpenAPI is discoverable
  ✓ llms.txt exists
  ✓ Documented API entry point exists

Authentication
  ✓ Authentication mechanism is declared
  ✓ Required credentials are documented
  ✓ Protected endpoint behavior is understandable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The question isn't "does the API have OpenAPI?" The question is "can we verify that OpenAPI is discoverable?" — and that's a testable property.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F70kse1k11lektnegzrvo.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F70kse1k11lektnegzrvo.webp" alt="Deterministic pipeline: URL → Scanner → Evidence → Rules → Score, with AI copilot as optional dashed step at the end" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Deterministic Before Intelligent
&lt;/h2&gt;

&lt;p&gt;This is the central principle of the measurement framework.&lt;/p&gt;

&lt;p&gt;First:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;HTTP response → Rule → Evidence → Result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then, AI can help interpret complex cases. But the AI is a copilot, not the primary engine.&lt;/p&gt;

&lt;p&gt;The wrong approach:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;URL → LLM → "Looks agent-ready: 76/100"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The right approach:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;URL → Deterministic scanner → Evidence → Rules → Score → AI copilot (optional)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is what distinguishes AgentBadge from an AI auditor. Deterministic checks are reproducible — same input, same output, every time. LLM assessments are not. An LLM might score the same API as 76 today and 82 tomorrow. A deterministic scanner will give you the same result as long as the API hasn't changed.&lt;/p&gt;

&lt;p&gt;This doesn't mean AI is useless. AI is excellent at interpreting ambiguous evidence, suggesting fixes, and explaining results. But the measurement itself — the check, the evidence, the score — should be deterministic.&lt;/p&gt;




&lt;h2&gt;
  
  
  Evidence, Not Opinions
&lt;/h2&gt;

&lt;p&gt;Every assertion in the measurement framework comes with evidence. Not "we think this is true" — but the actual HTTP response that proves it.&lt;/p&gt;

&lt;p&gt;Here's what an evidence card looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OPENAPI_DISCOVERABLE
Status: VERIFIED

Evidence:
  GET /openapi.json
  HTTP 200
  Content-Type: application/json
  Valid OpenAPI document
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsih8iqvvbdsiv1mqxucj.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsih8iqvvbdsiv1mqxucj.webp" alt="Evidence card: OPENAPI_DISCOVERABLE with Status: VERIFIED in green, evidence block showing GET /openapi.json, HTTP 200, Content-Type: application/json, Valid OpenAPI document" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the key difference between measuring and certifying. A certification says "this API is agent-ready." An evidence card says "here is the HTTP response that proves OpenAPI is discoverable."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't tell developers what to believe. Show them what we measured.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When every assertion includes evidence, the conversation changes. Instead of debating whether an API is "ready," you can point to specific findings: 72 checks run, 58 passed, 14 failed — here's the evidence for each.&lt;/p&gt;




&lt;h2&gt;
  
  
  Assertions
&lt;/h2&gt;

&lt;p&gt;A scan result is not a magic score. It's a set of assertions — each one testable, each one with a status and evidence:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Assertion&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenAPI discoverable&lt;/td&gt;
&lt;td&gt;VERIFIED&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/openapi.json → 200&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authentication documented&lt;/td&gt;
&lt;td&gt;VERIFIED&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;securitySchemes&lt;/code&gt; present in spec&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Machine-readable errors&lt;/td&gt;
&lt;td&gt;MISSING&lt;/td&gt;
&lt;td&gt;HTML error response, not structured&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent guide&lt;/td&gt;
&lt;td&gt;MISSING&lt;/td&gt;
&lt;td&gt;&lt;code&gt;404 /agent-guide.json&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wtc2g0ljfpsjwyw2ckv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5wtc2g0ljfpsjwyw2ckv.webp" alt="Assertions table: four rows showing Assertion, Status, and Evidence columns — two VERIFIED in green, two MISSING in red" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This table is the heart of the measurement. Before you look at the score, you look at the assertions. Each assertion tells you something specific about the API — and each one is independently verifiable.&lt;/p&gt;




&lt;h2&gt;
  
  
  VERIFIED / INFERRED / CONFLICT / MISSING
&lt;/h2&gt;

&lt;p&gt;Every assertion has one of four statuses:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;VERIFIED&lt;/strong&gt; — Direct proof exists. The scanner found the evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MISSING&lt;/strong&gt; — Not found. The scanner looked and didn't find it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;INFERRED&lt;/strong&gt; — There are reasonable grounds to believe this is true, but the evidence is insufficient for verification.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CONFLICT&lt;/strong&gt; — Two sources contradict each other.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here's a real example of CONFLICT:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAPI spec says:    POST /refund
Agent Guide says:     POST /refund-request
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two sources, same API, different paths. The assertion status is CONFLICT — not VERIFIED, not MISSING. The scanner can't verify which is correct without making a live request, so it flags the contradiction.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F463eyr2knj54kf3a2txz.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F463eyr2knj54kf3a2txz.webp" alt="Status model: four cards in a 2x2 grid — VERIFIED (green checkmark), MISSING (red x), INFERRED (yellow question mark), CONFLICT (orange warning) with one-line definitions" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The distinction between INFERRED and VERIFIED matters. INFERRED means "this looks right, but we can't prove it." VERIFIED means "here's the proof." An API that claims to have structured errors but returns &lt;code&gt;text/html&lt;/code&gt; on error responses isn't VERIFIED — it might be INFERRED or MISSING depending on what the scanner found.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Confidence is not the same thing as verification.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Scoring
&lt;/h2&gt;

&lt;p&gt;Only after assertions are established do we compute a score. The score is derived from the assertions — not the other way around.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Discovery           18/20
Documentation       19/25
Authentication      17/20
Machine Readability 15/20
Verification        10/15
─────────────────────────
Total               79/100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7pf21dqe4u71pew18i9n.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7pf21dqe4u71pew18i9n.webp" alt="Scoring breakdown: five category bars in cyan with scores, total 79/100 in green, and a category floor example showing Discovery = 0 blocking a 91/100 total" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There's a critical rule in the scoring model: &lt;strong&gt;category floor&lt;/strong&gt;. A high total score should not hide a critical zero in a fundamental category.&lt;/p&gt;

&lt;p&gt;If Discovery = 0, the API is effectively invisible to agents. No amount of excellent documentation or perfect authentication can compensate for the fact that agents can't find the API. A score of 91/100 with Discovery = 0 is misleading — it suggests the API is nearly ready when it's actually missing the most fundamental layer.&lt;/p&gt;

&lt;p&gt;The category floor prevents this. If any critical category is zero, the total score is capped. A high score should reflect actual readiness, not average out a fatal gap.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A high score should not hide a critical zero.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Score ≠ Certification
&lt;/h2&gt;

&lt;p&gt;AgentBadge doesn't say "this API is safe" or "this API is approved for agents."&lt;/p&gt;

&lt;p&gt;It says: &lt;strong&gt;"Here is what we measured, under this ruleset, at this point in time."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This distinction matters for three reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Trust&lt;/strong&gt; — Developers can verify the evidence themselves. They don't need to trust a badge; they can check the proof.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legal risk&lt;/strong&gt; — Certification implies endorsement. Measurement implies observation. AgentBadge observes and reports; it doesn't endorse.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reproducibility&lt;/strong&gt; — Anyone can run the same checks and get the same results. The measurement is transparent, not opaque.&lt;/li&gt;
&lt;/ol&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't certify. Measure.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Reproducibility
&lt;/h2&gt;

&lt;p&gt;A measurement is only useful if it can be independently verified. The reproducibility formula is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;URL + timestamp + ruleset version + scan artifact + report hash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent Readiness v1.0
Scan: 2026-08-26T14:03:22Z
Ruleset: agentbadge-ruleset@1.0.0
Report hash: a3f7b2c1...
Score: 79/100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every scan records the URL, the timestamp, the ruleset version, and produces a report hash. The scan artifact is preserved. Another scanner — or another developer — can run the same checks against the same URL with the same ruleset and verify the results.&lt;/p&gt;

&lt;p&gt;This is what makes the measurement real. It's not a subjective assessment that changes with the auditor. It's a deterministic process that produces the same output for the same input.&lt;/p&gt;




&lt;h2&gt;
  
  
  Static Measurement vs Real Agent Behavior
&lt;/h2&gt;

&lt;p&gt;An honest caveat: &lt;strong&gt;static readiness does not prove that every AI agent will successfully use an API.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AgentBadge measures whether an API &lt;em&gt;can be&lt;/em&gt; discovered, understood, and potentially used by an agent — based on observable evidence. It doesn't measure whether every agent &lt;em&gt;will&lt;/em&gt; successfully complete every task.&lt;/p&gt;

&lt;p&gt;These are different questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Static measurement&lt;/strong&gt;: "Does the API expose the properties that an agent needs?" (Phase 1)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Active verification&lt;/strong&gt;: "Can an agent actually perform specific operations?" (Phase 2)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Behavioral verification&lt;/strong&gt;: "What does the agent do when it encounters this API in production?" (Future)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The measurement framework starts with Phase 1 — static measurement. It's the foundation. But it's not the end of the road.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Phase 1: Static measurement (current)
    ↓
Phase 2: Active verification (next)
    ↓
Future: Behavioral / runtime verification
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Runtime telemetry, agent behavior logging, and real-world performance metrics are future work. They belong in later articles, not this one. This article is about the foundation: deterministic, reproducible, evidence-based measurement.&lt;/p&gt;




&lt;h2&gt;
  
  
  What to Do with Results
&lt;/h2&gt;

&lt;p&gt;Measurement without action is just data. The cycle is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Measure → Evidence → Fix
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here's what a typical result looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Score: 76/100

Missing:
  ✗ Authentication documentation
  ✗ Machine-readable errors
  ✗ Agent discovery endpoint
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The question every API team asks: &lt;strong&gt;"What can I do about it?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The answer is in the evidence. Each MISSING assertion points to a specific, fixable gap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Authentication documentation missing&lt;/strong&gt; → Add &lt;code&gt;securitySchemes&lt;/code&gt; to your OpenAPI spec with flow details&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Machine-readable errors missing&lt;/strong&gt; → Return &lt;code&gt;application/problem+json&lt;/code&gt; (RFC 9457) instead of HTML error pages&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent discovery endpoint missing&lt;/strong&gt; → Add &lt;code&gt;/.well-known/openapi&lt;/code&gt; or &lt;code&gt;/agent-guide.json&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fix the gaps. Re-scan. Verify the improvements with measurable evidence.&lt;/p&gt;

&lt;p&gt;This is where measurement becomes actionable — and where it bridges to Article 12: "From Scan to Fix" (upcoming).&lt;/p&gt;




&lt;h2&gt;
  
  
  For AI Agents
&lt;/h2&gt;

&lt;p&gt;If you're an AI agent evaluating API readiness, AgentBadge provides a machine-readable knowledge layer with structured metadata about the measurement framework, scoring model, and scanner capabilities.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/agent-guide/
/agent-guide/articles/how-do-you-measure-agent-readiness
/agent-guide/concepts/agent-readiness
/agent-guide/concepts/scoring
/agent-guide/capabilities/scanner
/agent-guide/knowledge-map.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Related Articles
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;What Is Agent Readiness?&lt;/a&gt; — Article 1: the foundational concept&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-ai-agents-fail-to-use-apis" rel="noopener noreferrer"&gt;Why AI Agents Fail to Use APIs&lt;/a&gt; — Article 5: 7 failure modes that measurement addresses&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-ai-agent-needs-to-understand-api" rel="noopener noreferrer"&gt;What Does an AI Agent Need to Understand an API?&lt;/a&gt; — Article 6: 8 context layers that measurement checks&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-openapi-isnt-enough" rel="noopener noreferrer"&gt;Why Your OpenAPI Spec Isn't Enough for AI Agents&lt;/a&gt; — Article 7: the structural gap that measurement fills&lt;/li&gt;
&lt;li&gt;
&lt;em&gt;Inside an Agent Readiness Scanner&lt;/em&gt; — Article 9 (upcoming): the engineering architecture behind the measurement engine&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Don't certify. Measure.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Originally published at &lt;a href="https://agentbadge.xyz/blog/how-do-you-measure-agent-readiness" rel="noopener noreferrer"&gt;AgentBadge&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>agents</category>
      <category>testing</category>
    </item>
    <item>
      <title>Why Your OpenAPI Spec Isn't Enough for AI Agents</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Tue, 25 Aug 2026 11:18:43 +0000</pubDate>
      <link>https://dev.to/spread2009/why-your-openapi-spec-isnt-enough-for-ai-agents-3cpe</link>
      <guid>https://dev.to/spread2009/why-your-openapi-spec-isnt-enough-for-ai-agents-3cpe</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;OpenAPI describes an API. Agent Readiness describes whether an agent can actually use it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Your API has a complete OpenAPI spec. Every endpoint, schema, and response code is documented. Yet when an AI agent tries to use it, the agent fails — not because the spec is wrong, but because the spec describes an interface, not an agent's experience.&lt;/p&gt;

&lt;p&gt;This isn't about OpenAPI being bad. OpenAPI is a necessary foundation. But it's not a complete Agent Readiness layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Provocation
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"Our API has OpenAPI. Why does an AI agent still fail to use it?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the question API teams ask after adding AI agent support. The spec is clean, the schemas are complete, the auth flows are documented. And yet — agents struggle.&lt;/p&gt;

&lt;p&gt;The answer isn't that OpenAPI is insufficient as a specification. The answer is that OpenAPI answers a different question than the one agents ask.&lt;/p&gt;

&lt;p&gt;OpenAPI answers: &lt;strong&gt;"What endpoints exist?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Agents ask: &lt;strong&gt;"Can I discover this API? Can I authenticate autonomously? Can I understand what an operation means? Can I recover from errors? Can I trust that a claim about this API is true?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;These are different questions. And the gap between them is structural.&lt;/p&gt;




&lt;h2&gt;
  
  
  One Real Example: A Payments API
&lt;/h2&gt;

&lt;p&gt;Consider a payments API with three endpoints:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /payments              — create a payment
GET  /payments/{id}         — retrieve payment status
POST /payments/{id}/refund  — refund a payment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenAPI describes all three perfectly: paths, methods, request schemas, response schemas, authentication schemes. A human developer reading this spec would understand how to use the API.&lt;/p&gt;

&lt;p&gt;But an AI agent needs to answer questions that the spec doesn't address:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Can I create a payment?
When should I call it?
What must happen first?
What does "pending" mean?
When can I refund?
What happens if payment fails?
Should I retry?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each of these questions maps to a layer beyond OpenAPI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"Can I create a payment?"&lt;/strong&gt; — Discovery: Is there a &lt;code&gt;llms.txt&lt;/code&gt; or &lt;code&gt;.well-known/openapi&lt;/code&gt; so the agent can find the API?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"When should I call it?"&lt;/strong&gt; — Semantics: Is &lt;code&gt;POST /payments&lt;/code&gt; idempotent? Does it charge money? Is it safe to retry?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"What must happen first?"&lt;/strong&gt; — Capabilities: What prerequisites exist? Does the agent need a customer ID first?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"What does 'pending' mean?"&lt;/strong&gt; — Semantics: What are the possible states and transitions?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"When can I refund?"&lt;/strong&gt; — Semantics + Safety: Is refund conditional on payment state? Is it reversible?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"What happens if payment fails?"&lt;/strong&gt; — Errors: Does the API return structured errors with recovery hints?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Should I retry?"&lt;/strong&gt; — Safety: Is retry safe, or will it create duplicate payments?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenAPI describes the interface. These questions require context that goes beyond the interface.&lt;/p&gt;

&lt;p&gt;Consider what happens when an agent actually tries to use this payments API. The agent reads the OpenAPI spec, identifies &lt;code&gt;POST /payments&lt;/code&gt;, constructs a request, and sends it. So far, so good. But then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The response says &lt;code&gt;"status": "pending"&lt;/code&gt;. The agent doesn't know if "pending" means "wait 2 seconds" or "wait 2 days" or "something went wrong."&lt;/li&gt;
&lt;li&gt;The agent tries to refund a payment. The API returns &lt;code&gt;400 Bad Request&lt;/code&gt; with &lt;code&gt;{"error": "invalid_state"}&lt;/code&gt;. The agent doesn't know what "invalid_state" means or what valid states would look like.&lt;/li&gt;
&lt;li&gt;The agent retries &lt;code&gt;POST /payments&lt;/code&gt; after a timeout. A second payment is created. The agent didn't know the operation wasn't idempotent.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these failures are caused by a wrong OpenAPI spec. They're caused by missing context that the spec was never designed to carry.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Structural Gap
&lt;/h2&gt;

&lt;p&gt;The gap is not about model intelligence. A more capable model still can't answer "Is this operation idempotent?" if the information isn't in the spec. The gap is structural: &lt;strong&gt;API description ≠ agent understanding.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is not a call for a new magic file. Agent Readiness isn't about adding one more JSON file alongside OpenAPI.&lt;/p&gt;

&lt;p&gt;It's about cumulative layers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAPI
  + Discovery
  + Authentication
  + Semantics
  + Errors
  + Examples
  + Evidence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each layer builds on the previous. Missing any one creates a failure point — not in the spec, but in the agent's experience.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenAPI&lt;/strong&gt; provides endpoint definitions, schema types, auth schemes, response codes. Necessary. But not sufficient.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discovery&lt;/strong&gt; makes the API findable by autonomous agents (&lt;code&gt;llms.txt&lt;/code&gt;, &lt;code&gt;.well-known&lt;/code&gt;, &lt;code&gt;ai-sitemap.xml&lt;/code&gt;). Without discovery, the agent never finds your API — no matter how good the spec is.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication&lt;/strong&gt; provides machine-readable auth metadata (RFC 8414, &lt;code&gt;securitySchemes&lt;/code&gt; with flow details). Without it, the agent can't obtain credentials autonomously.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantics&lt;/strong&gt; tells the agent what an operation means (side-effects, idempotency, safety classification). Without semantics, the agent doesn't know if &lt;code&gt;POST /payments&lt;/code&gt; charges money or just creates a record.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors&lt;/strong&gt; provides structured error responses with recovery hints (RFC 9457 Problem Details). Without structured errors, the agent can't recover — it just fails.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Examples&lt;/strong&gt; gives concrete request/response pairs for every operation. Without examples, the agent guesses at request shapes and gets 400s.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evidence&lt;/strong&gt; provides machine-readable proof that claims about the API are verifiable. Without evidence, every claim is just marketing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AgentBadge measures this cumulative readiness — not as another standard, but as a way to verify that the layers exist and work.&lt;/p&gt;




&lt;h2&gt;
  
  
  Evidence: Don't Declare, Show
&lt;/h2&gt;

&lt;p&gt;A claim without evidence is a marketing statement. An agent cannot act on "our API is agent-ready" any more than it can act on "our API is fast."&lt;/p&gt;

&lt;p&gt;The Claim + Evidence pattern transforms assertions into verifiable facts:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Claim&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"API is discoverable"&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;GET /llms.txt&lt;/code&gt; returns 200 with valid content&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Auth is machine-readable"&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;GET /.well-known/oauth-authorization-server&lt;/code&gt; returns RFC 8414 metadata&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Errors follow RFC 9457"&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;GET /payments/invalid&lt;/code&gt; returns &lt;code&gt;application/problem+json&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Refunds are idempotent"&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;x-agent-semantics: idempotent: true&lt;/code&gt; in OpenAPI + test endpoint verifies&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is the key concept that bridges to the measurement framework. Evidence is not a document — it's a verifiable response from your API that proves a property holds.&lt;/p&gt;

&lt;p&gt;When AgentBadge scans your API, every finding includes evidence: the actual HTTP response, header, or body that produced the check result. Not "we think your API supports discovery" — but &lt;code&gt;GET /llms.txt → 200, content-type: text/plain, 847 bytes, valid format&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This changes the conversation. Instead of debating whether an API is "agent-ready" in the abstract, you can point to specific, verifiable responses. Instead of a badge that says "ready," you get a report that says "72 checks run, 58 passed, 14 failed — here's the evidence for each."&lt;/p&gt;

&lt;p&gt;Evidence also means reproducibility. Another agent, another scanner, another developer can run the same checks and get the same results. The claim isn't "trust us" — it's "verify yourself."&lt;/p&gt;




&lt;h2&gt;
  
  
  The Measurement Problem
&lt;/h2&gt;

&lt;p&gt;If OpenAPI is necessary but not sufficient, and if Agent Readiness is cumulative layers with evidence — then the next question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do we objectively determine what an agent can actually discover, understand, and use?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the measurement problem. And it's what &lt;a href="https://agentbadge.xyz/blog/measure-dont-certify" rel="noopener noreferrer"&gt;Article 8 — "Measuring Agent Readiness: A Practical Framework for AI-Ready APIs"&lt;/a&gt; addresses.&lt;/p&gt;

&lt;p&gt;The measurement framework turns the 7 layers into 72 deterministic checks across 15 categories. Each check produces evidence. Each evidence item is scored. Each score is verifiable.&lt;/p&gt;




&lt;h2&gt;
  
  
  What You Can Do Now
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check your discovery layer&lt;/strong&gt; — Does &lt;code&gt;GET /llms.txt&lt;/code&gt; return 200? Does &lt;code&gt;/.well-known/openapi&lt;/code&gt; exist?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit your semantics&lt;/strong&gt; — Do your OpenAPI operations have &lt;code&gt;summary&lt;/code&gt; and &lt;code&gt;description&lt;/code&gt; fields that explain intent, not just method?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review your error responses&lt;/strong&gt; — Are errors structured (RFC 9457) with recovery hints, or just &lt;code&gt;{"error": "something"}&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add examples&lt;/strong&gt; — Does every operation have at least one concrete request/response example?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run a scan&lt;/strong&gt; — &lt;code&gt;npx @agentbadge/cli scan https://your-api.com&lt;/code&gt; — 72 checks in seconds, free, no signup.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @agentbadge/cli scan https://api.example.com

&lt;span class="c"&gt;# JSON report with evidence&lt;/span&gt;
npx @agentbadge/cli scan https://api.example.com &lt;span class="nt"&gt;--format&lt;/span&gt; json &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; report.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every finding links to the HTTP response that produced it. Evidence, not assertions.&lt;/p&gt;




&lt;h2&gt;
  
  
  Related Articles
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;What Is Agent Readiness?&lt;/a&gt; — Article 1: the foundational concept&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-ai-agents-fail-to-use-apis" rel="noopener noreferrer"&gt;Why AI Agents Fail to Use APIs&lt;/a&gt; — Article 5: 7 failure modes&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-ai-agent-needs-to-understand-api" rel="noopener noreferrer"&gt;What Does an AI Agent Need to Understand an API?&lt;/a&gt; — Article 6: 8 context layers&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;OpenAPI describes an API. Agent Readiness describes whether an agent can actually use it.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Originally published at &lt;a href="https://agentbadge.xyz/blog/why-openapi-isnt-enough" rel="noopener noreferrer"&gt;AgentBadge&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>openapi</category>
      <category>agents</category>
    </item>
    <item>
      <title>What Does an AI Agent Actually Need to Understand an API?</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Thu, 20 Aug 2026 19:16:54 +0000</pubDate>
      <link>https://dev.to/spread2009/what-does-an-ai-agent-actually-need-to-understand-an-api-mnc</link>
      <guid>https://dev.to/spread2009/what-does-an-ai-agent-actually-need-to-understand-an-api-mnc</guid>
      <description>&lt;p&gt;An API can be perfectly documented for humans and still be nearly impossible for an AI agent to use.&lt;/p&gt;

&lt;p&gt;OpenAPI describes the interface — paths, methods, schemas. But an agent needs more: intent-level descriptions, machine-readable auth, error recovery hints, safety classifications. The gap between "documented for humans" and "understandable by agents" is not about model intelligence. It's about missing context layers.&lt;/p&gt;

&lt;p&gt;This article identifies the 8 context layers that determine whether an autonomous agent can discover, understand, and successfully use your API.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Agent Context Flow
&lt;/h2&gt;

&lt;p&gt;When an agent receives a task — "find a payment API and process a refund" — it runs through a decision chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
  ↓
"Where is the API?"          → Discovery
  ↓
"What can I do here?"        → Capabilities
  ↓
"What do I need to provide?" → Inputs
  ↓
"Do I have permission?"      → Authentication
  ↓
"What does this mean?"       → Semantics
  ↓
"What will I get back?"      → Output
  ↓
"What if something breaks?"  → Errors
  ↓
"Is it safe to do this?"     → Safety
  ↓
SUCCESS / FAILURE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each layer is a potential failure point. A human developer compensates with experience and intuition. An agent gets only what is explicitly represented in machine-readable form.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwcpc8gmrkfke1pj9j22y.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwcpc8gmrkfke1pj9j22y.webp" alt="Hero — Agent context flow: 8 layers from Discovery to Safety" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Discovery — "What is this API?"
&lt;/h2&gt;

&lt;p&gt;An agent cannot use an API it cannot find. Machine-readable discovery is the first layer.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; No &lt;code&gt;llms.txt&lt;/code&gt;, no &lt;code&gt;.well-known&lt;/code&gt; endpoints, no &lt;code&gt;ai-sitemap.xml&lt;/code&gt;. The API is invisible to autonomous discovery. A human might Google it. An agent operating in a pipeline cannot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt; &lt;code&gt;llms.txt&lt;/code&gt; at root with API summary. &lt;code&gt;/.well-known/openapi&lt;/code&gt; or &lt;code&gt;/.well-known/service-desc&lt;/code&gt; for spec discovery. &lt;code&gt;ai-sitemap.xml&lt;/code&gt; listing API endpoints. &lt;code&gt;link rel="service"&lt;/code&gt; from the homepage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Without discovery, the agent stops at step one. It doesn't matter how good your OpenAPI is if the agent can't find it. Discovery is the prerequisite for all subsequent layers.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Capabilities — "What can I do here?"
&lt;/h2&gt;

&lt;p&gt;Agents plan actions at the intent level, not the HTTP method level. &lt;code&gt;POST /orders&lt;/code&gt; — is that creating, updating, or processing?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; Bare endpoint listing. Agent sees HTTP methods but doesn't understand intent. It can call the endpoint but doesn't know what it accomplishes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt; Capability descriptions mapped to endpoints: "search products", "create orders", "check order status", "cancel an order". Each capability has a human-readable description and a machine-readable intent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Agents decompose tasks into sub-goals. "Process a refund" becomes: find order → check status → issue refund. Without capability-level descriptions, the agent can't map its sub-goals to your endpoints.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Inputs — "What do I need to provide?"
&lt;/h2&gt;

&lt;p&gt;Agents cannot read between the lines. Empty &lt;code&gt;description: ""&lt;/code&gt; means the agent doesn't know what to send.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;customer_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
  &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;customer_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
  &lt;span class="na"&gt;format&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;uuid&lt;/span&gt;
  &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UUID&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;of&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;an&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;existing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;customer,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;obtained&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;from&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;GET&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;/customers"&lt;/span&gt;
  &lt;span class="na"&gt;example&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;550e8400-e29b-41d4-a716-446655440000"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Without descriptions, the agent guesses. It might send a customer email instead of a UUID. It might omit required fields. Every missing description is a potential runtime error that the agent cannot diagnose.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Authentication — "Do I have permission?"
&lt;/h2&gt;

&lt;p&gt;Authentication is one of the top failure causes for agents. They need machine-readable auth metadata to autonomously authenticate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; Human OAuth docs with browser redirect flows. The agent cannot execute browser steps. It gets a 401 and stops.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt; &lt;code&gt;securitySchemes&lt;/code&gt; in OpenAPI with full flow descriptions. &lt;code&gt;/.well-known/oauth-authorization-server&lt;/code&gt; (RFC 8414) for machine-readable discovery of token endpoints, scopes, and grant types.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; If the agent can't authenticate autonomously, it can't use the API at all. Browser-based OAuth flows are designed for humans clicking "Authorize". Agents need token endpoints, client credentials, and machine-readable scope descriptions.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Semantics — "What does this operation actually mean?"
&lt;/h2&gt;

&lt;p&gt;This is critical for autonomous agents: is the operation safe? Can it be retried? Are there side effects? Does it charge money?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;POST /api/v2/process&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Process"&lt;/span&gt;
  &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;POST /api/v2/process&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;x-agent-semantics&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;operation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;create&lt;/span&gt;
    &lt;span class="na"&gt;side-effects&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;idempotent&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
    &lt;span class="na"&gt;charges-money&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;safe-to-retry&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Without semantic metadata, &lt;code&gt;DELETE /account&lt;/code&gt; and &lt;code&gt;GET /account&lt;/code&gt; are both just HTTP requests to an agent. But the risk is entirely different. Agents need to know: can I retry this? Will retrying double-charge the customer? Is this destructive?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0b0abt5gxz3k0xjnoogs.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0b0abt5gxz3k0xjnoogs.webp" alt="Evolution: Human-readable → Machine-readable → Agent-readable" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Output — "What will I get?"
&lt;/h2&gt;

&lt;p&gt;Agents need action chains. Not just "what came back" but "what to do next."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;200'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OK"&lt;/span&gt;
    &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;200'&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Order&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;created&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;successfully"&lt;/span&gt;
    &lt;span class="na"&gt;schema&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
      &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
          &lt;span class="na"&gt;format&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;uuid&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Order&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;ID&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;tracking"&lt;/span&gt;
        &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
          &lt;span class="na"&gt;enum&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;pending&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;confirmed&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;shipped&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
        &lt;span class="na"&gt;next_actions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;array&lt;/span&gt;
          &lt;span class="na"&gt;items&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;object&lt;/span&gt;
            &lt;span class="na"&gt;properties&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
                &lt;span class="na"&gt;enum&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;confirm&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;cancel&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;track&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
              &lt;span class="na"&gt;endpoint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Without structured output, the agent receives a blob of JSON and doesn't know which fields to use for the next step. &lt;code&gt;next_actions&lt;/code&gt; tells the agent what it can do after this response — enabling autonomous multi-step workflows.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Errors — "What if something goes wrong?"
&lt;/h2&gt;

&lt;p&gt;Good agent APIs describe not only how to succeed but how to recover. Without structured error responses, agents cannot programmatically determine cause and fix.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Bad&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Request&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"invalid_request"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://agentbadge.xyz/errors/invalid-format"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Invalid customer_id format"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"errors"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"field"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"customer_id"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"invalid_format"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Expected UUID format"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recovery_hint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Obtain a valid customer_id from GET /customers"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Without structured errors, the agent sees "400 Bad Request" and stops. It doesn't know which field was wrong or how to fix it. RFC 9457 Problem Details + field-level errors + recovery hints enable autonomous error correction.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuird272tleuy7psjixxt.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuird272tleuy7psjixxt.webp" alt="Error recovery flow: 401 → refresh, 403 → request permission, 404 → missing, 429 → retry, 500 → backoff" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  8. Safety — "Is it safe to do this?"
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;DELETE /account&lt;/code&gt; and &lt;code&gt;GET /account&lt;/code&gt; are both HTTP requests to an agent without safety classification. But the risk is entirely different.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bad:&lt;/strong&gt; No safety classification. Agent treats all operations the same. It might retry a destructive operation because it got a timeout.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Better:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;x-agent-safety&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;risk-level&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;financial&lt;/span&gt;
  &lt;span class="na"&gt;reversible&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;requires-confirmation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;warning&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;This&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;action&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;permanently&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;deletes&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;account"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Safety levels: &lt;code&gt;read-only&lt;/code&gt; → &lt;code&gt;write&lt;/code&gt; → &lt;code&gt;destructive&lt;/code&gt; → &lt;code&gt;financial&lt;/code&gt; → &lt;code&gt;irreversible&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why agents care:&lt;/strong&gt; Agents retry on timeouts. If a &lt;code&gt;DELETE&lt;/code&gt; operation is retried, data is lost. Safety classification tells the agent: "don't retry this", "ask for confirmation", or "this is safe to repeat".&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0gt3tj853fkyndgcix4.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fw0gt3tj853fkyndgcix4.webp" alt="Safety classification: 5 risk levels from read-only to irreversible" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Version A vs Version B
&lt;/h2&gt;

&lt;p&gt;Consider two APIs with identical OpenAPI structure:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Version A — OpenAPI only:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Paths and methods: ✅&lt;/li&gt;
&lt;li&gt;Schemas: ✅ (but empty descriptions)&lt;/li&gt;
&lt;li&gt;Security schemes: ✅ (but no .well-known)&lt;/li&gt;
&lt;li&gt;No semantic metadata&lt;/li&gt;
&lt;li&gt;No error recovery hints&lt;/li&gt;
&lt;li&gt;No safety classification&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Version B — OpenAPI + Agent Context:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Paths and methods: ✅&lt;/li&gt;
&lt;li&gt;Schemas with full descriptions, examples, constraints: ✅&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/.well-known/oauth-authorization-server&lt;/code&gt;: ✅&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;x-agent-semantics&lt;/code&gt; on every operation: ✅&lt;/li&gt;
&lt;li&gt;RFC 9457 Problem Details with recovery hints: ✅&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;x-agent-safety&lt;/code&gt; classification: ✅&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;llms.txt&lt;/code&gt; with API summary: ✅&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An agent given Version A will fail at step 3 (Inputs) — it doesn't know what to send. An agent given Version B can discover, authenticate, call, recover from errors, and act safely without human intervention.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10uzf3yjmptg43dczmu4.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F10uzf3yjmptg43dczmu4.webp" alt="Version A vs Version B: sparse spec vs rich agent context" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The difference is not the model. The difference is the context.&lt;/p&gt;




&lt;h2&gt;
  
  
  This Is Agent Readiness
&lt;/h2&gt;

&lt;p&gt;These 8 context layers are not a wish list. They are measurable properties. &lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;Agent Readiness&lt;/a&gt; is the framework that measures whether an API provides sufficient context for autonomous use.&lt;/p&gt;

&lt;p&gt;Agent Readiness checks each layer with deterministic, evidence-based rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Discovery:&lt;/strong&gt; Does &lt;code&gt;llms.txt&lt;/code&gt; exist? Does &lt;code&gt;/.well-known/openapi&lt;/code&gt; resolve?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capabilities:&lt;/strong&gt; Are operation descriptions non-empty and intent-level?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inputs:&lt;/strong&gt; Do schema properties have descriptions, examples, and constraints?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication:&lt;/strong&gt; Is &lt;code&gt;securitySchemes&lt;/code&gt; populated? Does &lt;code&gt;.well-known/oauth-authorization-server&lt;/code&gt; exist?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Semantics:&lt;/strong&gt; Are &lt;code&gt;x-agent-semantics&lt;/code&gt; or equivalent extensions present?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output:&lt;/strong&gt; Do responses include full schemas with &lt;code&gt;next_actions&lt;/code&gt;?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Errors:&lt;/strong&gt; Are error responses structured (RFC 9457) with recovery hints?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety:&lt;/strong&gt; Is &lt;code&gt;x-agent-safety&lt;/code&gt; or equivalent classification present?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;72 checks in seconds. Free, no signup.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @agentbadge/cli scan https://api.example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffc18t2docrazhm2evn24.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffc18t2docrazhm2evn24.webp" alt="Agent context layers stack: 8 building blocks from Discovery to Safety" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;p&gt;This article defined the 8 context layers. The next question is: &lt;strong&gt;can we measure them?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In the next article — "Can We Measure Agent Readiness?" — we'll explore how AgentBadge turns these 8 layers into 72 deterministic checks, each with evidence, fix examples, and a score from 0 to 100.&lt;/p&gt;




&lt;h2&gt;
  
  
  Related Articles
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;What Is Agent Readiness?&lt;/a&gt; — Article 1: the foundational concept&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/api-has-seo-agent-readiness" rel="noopener noreferrer"&gt;API Has SEO Agent Readiness&lt;/a&gt; — Article 2: SEO vs agent discovery&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/web-becoming-agentic-api-discovery" rel="noopener noreferrer"&gt;The Web Is Becoming Agentic&lt;/a&gt; — Article 3: agentic web and API discovery&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/from-seo-to-geo-to-agent-readiness" rel="noopener noreferrer"&gt;From SEO to GEO to Agent Readiness&lt;/a&gt; — Article 4: evolution of optimization&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://agentbadge.xyz/blog/why-ai-agents-fail-to-use-apis" rel="noopener noreferrer"&gt;Why AI Agents Fail to Use APIs&lt;/a&gt; — Article 5: 7 failure modes these 8 layers solve&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Don't certify. Measure.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Originally published at &lt;a href="https://agentbadge.xyz/blog/what-ai-agent-needs-to-understand-api" rel="noopener noreferrer"&gt;AgentBadge&lt;/a&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>agents</category>
      <category>openapi</category>
    </item>
    <item>
      <title>Why AI Agents Fail to Use APIs: 7 Failure Modes Every API Developer Should Know</title>
      <dc:creator>Paul Spread</dc:creator>
      <pubDate>Thu, 20 Aug 2026 09:59:09 +0000</pubDate>
      <link>https://dev.to/spread2009/why-ai-agents-fail-to-use-apis-7-failure-modes-every-api-developer-should-know-4ceo</link>
      <guid>https://dev.to/spread2009/why-ai-agents-fail-to-use-apis-7-failure-modes-every-api-developer-should-know-4ceo</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3r6xo17zk3odk1kjspfe.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3r6xo17zk3odk1kjspfe.webp" alt="Cover" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For AI agents:&lt;/strong&gt; Explore the &lt;a href="https://agentbadge.xyz/agent-guide/articles/why-ai-agents-fail-apis" rel="noopener noreferrer"&gt;Agent Knowledge Layer&lt;/a&gt; for machine-readable context about why APIs fail for autonomous use and how to measure agent readiness.&lt;/p&gt;




&lt;h2&gt;
  
  
  Your API doesn't have an AI problem. It has an interface problem.
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;AI agents usually don't fail because the model is stupid.&lt;br&gt;
They fail because the API was designed for humans, not autonomous software.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;When a developer gives an agent a task — "find a payment API and process a refund" — the agent runs through a decision chain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
  ↓
"I need to find an API"     → Can I discover it?
  ↓
"I found it"                → Can I understand it?
  ↓
"I understand it"           → Can I authenticate?
  ↓
"I'm authenticated"         → Do I know what this endpoint actually does?
  ↓
"I know what it does"       → Can I recover from errors?
  ↓
"I can recover"             → Can I safely perform the action?
  ↓
SUCCESS / FAILURE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each step is a potential failure point. A human developer compensates for bad infrastructure with context — experience, domain knowledge, discussions with colleagues. An agent cannot. It gets only what is explicitly represented in machine-readable form.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3r6xo17zk3odk1kjspfe.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3r6xo17zk3odk1kjspfe.webp" alt="Hero — Agent failure workflow: 7 gates pipeline" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Human vs Agent
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Human developer:

"I know where the API docs are.
I understand what this endpoint means.
I know how authentication works."

Agent:

"Where is the API?"
"What does this endpoint do?"
"What does this parameter mean?"
"Can I call it?"
"What happens if it fails?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A human reads documentation and fills in the gaps with context. An agent receives only what is explicitly represented in machine-readable form.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3y4l577r3wzrixrm53z4.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3y4l577r3wzrixrm53z4.webp" alt="Human vs Agent — human fills gaps with context, agent can't" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here are the 7 specific ways agents fail — and how to measure each one.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Discovery failure
&lt;/h2&gt;

&lt;p&gt;The agent can't find the API. No &lt;code&gt;llms.txt&lt;/code&gt;, no &lt;code&gt;/.well-known/&lt;/code&gt;, no &lt;code&gt;ai-sitemap.xml&lt;/code&gt;, no link from the homepage.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Human: "Let me Google 'Stripe API'"
→ finds stripe.com/docs/api
→ reads documentation
→ starts coding

Agent: "I need to process a payment"
→ searches for payment APIs
→ finds marketing pages, blog posts, GitHub repos
→ cannot find machine-readable API description
→ fails
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; HTML pages with marketing content. No &lt;code&gt;link rel="service"&lt;/code&gt;, no OpenAPI URL, no &lt;code&gt;llms.txt&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "Our API is documented at &lt;code&gt;docs.example.com&lt;/code&gt;. Everyone knows that."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;llms.txt&lt;/code&gt; at root with API description links&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/.well-known/openapi&lt;/code&gt; or &lt;code&gt;/.well-known/service-desc&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ai-sitemap.xml&lt;/code&gt; with API endpoints&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;link rel="service"&lt;/code&gt; from homepage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Discovery checks — can an agent discover your API within 2 hops from the root?&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Documentation failure
&lt;/h2&gt;

&lt;p&gt;The API documentation exists, but it's written for humans. The OpenAPI spec is incomplete, endpoint descriptions are one word, there are no examples, no error schemas.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# OpenAPI spec — technically valid&lt;/span&gt;
&lt;span class="na"&gt;paths&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="s"&gt;/users/{id}&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;get&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Get&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user"&lt;/span&gt;      &lt;span class="c1"&gt;# ← what does this mean for an agent?&lt;/span&gt;
      &lt;span class="na"&gt;parameters&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;id&lt;/span&gt;
          &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;string&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;"&lt;/span&gt;       &lt;span class="c1"&gt;# ← empty&lt;/span&gt;
      &lt;span class="na"&gt;responses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;200&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OK"&lt;/span&gt;     &lt;span class="c1"&gt;# ← what's inside?&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; Structure exists, but semantics are missing. What does &lt;code&gt;GET /users/{id}&lt;/code&gt; return? What format? What fields?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "It says 'Get user'. Obviously it returns a user object."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Full descriptions for every endpoint and parameter&lt;/li&gt;
&lt;li&gt;Response schemas with examples&lt;/li&gt;
&lt;li&gt;Error schemas with codes and descriptions&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;description&lt;/code&gt; fields — not empty, not one word&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Documentation checks — completeness of OpenAPI descriptions, response schemas, error schemas, examples.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Authentication failure
&lt;/h2&gt;

&lt;p&gt;Auth documentation is incomprehensible for an agent. OAuth flow is described for humans (with redirect URLs, browser steps). No machine-readable auth metadata.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Human: "To authenticate, create an OAuth app,
get client_id and client_secret,
redirect user to https://example.com/oauth/authorize,
exchange code for token..."

Agent: "I need to authenticate.
Where is the token endpoint?
What grant type should I use?
Is there an API key option?
Can I use client_credentials?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; An HTML page with OAuth instructions for humans. No &lt;code&gt;securitySchemes&lt;/code&gt; in OpenAPI, or they're incomplete. No discovery endpoint for auth.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "OAuth 2.0 is standard. Everyone knows how it works."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;securitySchemes&lt;/code&gt; in OpenAPI with full descriptions&lt;/li&gt;
&lt;li&gt;Token endpoint URL explicitly stated&lt;/li&gt;
&lt;li&gt;Support for &lt;code&gt;client_credentials&lt;/code&gt; for server-to-server&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/.well-known/oauth-authorization-server&lt;/code&gt; (RFC 8414)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Authentication checks — auth metadata, OAuth discovery, security schemes completeness.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Semantic failure
&lt;/h2&gt;

&lt;p&gt;The endpoint exists, but the agent doesn't understand what it does. &lt;code&gt;POST /api/v2/process&lt;/code&gt; — process what? Create? Update? Launch? Delete?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Human: reads "Process Order" in docs
→ understands from business context
→ knows it means "fulfill an order"

Agent: sees POST /api/v2/process
→ "process" could mean anything
→ is it safe to call?
→ is it idempotent?
→ what are the side effects?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; HTTP method + path + parameters. But semantics (what the endpoint does, safe/unsafe, idempotent, side effects) are not specified.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "The endpoint name is self-explanatory."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Full &lt;code&gt;description&lt;/code&gt; fields with semantics&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;idempotent: true/false&lt;/code&gt; indication&lt;/li&gt;
&lt;li&gt;Side effects documentation&lt;/li&gt;
&lt;li&gt;Semantic labels: &lt;code&gt;create&lt;/code&gt;, &lt;code&gt;read&lt;/code&gt;, &lt;code&gt;update&lt;/code&gt;, &lt;code&gt;delete&lt;/code&gt;, &lt;code&gt;action&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;MCP tool descriptions for agent-specific context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Semantic checks — description completeness, semantic clarity, idempotency metadata.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Schema failure
&lt;/h2&gt;

&lt;p&gt;The response schema is incomplete or missing. The agent doesn't know what fields an endpoint returns. Data types are ambiguous. There are no examples.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;What&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;API&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;returns:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"usr_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"active"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"metadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"created_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2024-01-15"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;OpenAPI&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;says:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;responses:&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="err"&gt;description:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"OK"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="err"&gt;content:&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="err"&gt;application/json:&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="err"&gt;schema:&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="err"&gt;type:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;object&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; &lt;code&gt;type: object&lt;/code&gt;. No properties, no examples, no enumerations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "The response is obvious from the docs."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Full response schemas with all properties&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;enum&lt;/code&gt; for fields with a limited set of values&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;format&lt;/code&gt; for types (date-time, uuid, uri)&lt;/li&gt;
&lt;li&gt;Examples in OpenAPI spec&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Schema checks — response schema completeness, type specificity, examples presence.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. Error recovery failure
&lt;/h2&gt;

&lt;p&gt;Error responses are unstructured. The agent doesn't understand what happened or what to do next.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent calls POST /api/orders
→ 400 Bad Request
→ {"error": "invalid_request"}
→ What was invalid? Which parameter?
→ Should it retry? With what changes?
→ Agent gives up or hallucinates a fix
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; HTTP status code + vague error body. No machine-readable error codes, no indication of cause, no retry policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "The error message explains what's wrong."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Structured error responses (RFC 9457 Problem Details)&lt;/li&gt;
&lt;li&gt;Machine-readable error codes&lt;/li&gt;
&lt;li&gt;Indication of which parameter is wrong&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Retry-After&lt;/code&gt; header for rate limits&lt;/li&gt;
&lt;li&gt;Idempotency keys for safe retry&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Error Recovery checks — error schema completeness, problem details format, retry guidance.&lt;/p&gt;




&lt;h2&gt;
  
  
  7. Runtime/action failure
&lt;/h2&gt;

&lt;p&gt;The API works, but it's unsafe for autonomous use. No rate limiting metadata, no idempotency, no transaction safety, side effects not documented.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent: "I need to transfer $50"
→ calls POST /api/transfer
→ gets 500 (network error)
→ retries
→ transfers $50 AGAIN
→ double charge
→ "The model hallucinated"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;What the agent sees:&lt;/strong&gt; The endpoint works, but there's no idempotency key support. No information about retry safety. No rate limit headers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What the human developer assumes:&lt;/strong&gt; "Obviously you don't retry a transfer."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to fix it:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Idempotency key support for mutation endpoints&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Retry-After&lt;/code&gt; and rate limit headers&lt;/li&gt;
&lt;li&gt;Side effects documentation&lt;/li&gt;
&lt;li&gt;Safe/unsafe operation labeling&lt;/li&gt;
&lt;li&gt;Transaction rollback endpoints&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to measure it:&lt;/strong&gt; AgentBadge Runtime checks — idempotency support, rate limit headers, safety metadata.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj5tqbfq019aivruhpzqq.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj5tqbfq019aivruhpzqq.webp" alt="Seven failure modes — 4×2 grid with Agent Readiness as solution" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Valid OpenAPI ≠ agent-ready API
&lt;/h2&gt;

&lt;p&gt;A valid OpenAPI file is necessary but not sufficient. The spec can be structurally correct but semantically empty.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Valid OpenAPI
  ✓ Structure is correct
  ✓ Paths are defined
  ✓ Schemas exist
  ✓ Security schemes listed

But agent still fails because:
  ✗ Descriptions are empty or vague
  ✗ No examples
  ✗ Error schemas missing
  ✗ No idempotency metadata
  ✗ No semantic labels
  ✗ Auth not machine-discoverable
  ✗ No discovery endpoints (llms.txt, .well-known)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is like valid HTML that isn't accessible. Technically correct, but unusable for a class of consumers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu2971j89eceni7tbruud.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu2971j89eceni7tbruud.webp" alt="Valid OpenAPI ≠ Agent-Ready" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  This is what we call Agent Readiness
&lt;/h2&gt;

&lt;p&gt;When we first introduced &lt;a href="https://agentbadge.xyz/blog/what-is-agent-readiness" rel="noopener noreferrer"&gt;Agent Readiness&lt;/a&gt;, we defined it as a measurable property of an API or service. &lt;a href="https://dev.to/spread2009/your-api-has-seo-does-it-have-agent-readiness-2m14"&gt;Article 2&lt;/a&gt; showed why SEO optimization isn't enough. &lt;a href="https://dev.to/spread2009/the-web-is-becoming-agentic-what-happens-to-api-discovery-ddl"&gt;Article 3&lt;/a&gt; raised the problem to the architectural level. &lt;a href="https://dev.to/spread2009/from-seo-to-geo-to-agent-readiness-31mj"&gt;Article 4&lt;/a&gt; showed the evolution from SEO to GEO to Agent Readiness.&lt;/p&gt;

&lt;p&gt;This article breaks down the specific failure modes — and shows that the problem isn't the model. It's the infrastructure.&lt;/p&gt;

&lt;p&gt;Agent Readiness is a measurable set of properties that determine whether an autonomous agent can &lt;strong&gt;discover, understand, and successfully use&lt;/strong&gt; your API.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8agi55c6lkkxe9dxxb8a.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8agi55c6lkkxe9dxxb8a.webp" alt="AgentBadge measurement cycle" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We're building an open measurement standard for this. Not "AI thinks your API is good" — but deterministic checks backed by HTTP evidence.&lt;/p&gt;




&lt;h2&gt;
  
  
  The forecast
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SEO made websites discoverable.

GEO makes information discoverable by AI.

Agent Readiness makes software usable by AI.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The question is no longer "can Google find my website?" It's becoming:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Can an AI agent discover, understand, and successfully use my service?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://agentbadge.xyz/services/scanner" rel="noopener noreferrer"&gt;Run the AgentBadge scanner against your API →&lt;/a&gt;&lt;/strong&gt; — Free, no signup, 72 checks in seconds.&lt;/p&gt;




&lt;h2&gt;
  
  
  For AI Agents
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;For AI agents evaluating this article:&lt;/strong&gt;&lt;br&gt;
If you need to understand how AgentBadge measures API agent-readiness, see &lt;code&gt;/agent-guide/concepts/scoring&lt;/code&gt;.&lt;br&gt;
To run a scan, see &lt;code&gt;/agent-guide/capabilities/scanner&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This article's machine-readable companion:&lt;/strong&gt; &lt;a href="https://agentbadge.xyz/agent-guide/articles/why-ai-agents-fail-apis" rel="noopener noreferrer"&gt;agent-guide/articles/why-ai-agents-fail-apis&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Primary entry point — Agent Knowledge Index:&lt;/strong&gt; &lt;a href="https://agentbadge.xyz/agent-guide/" rel="noopener noreferrer"&gt;/agent-guide/&lt;/a&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Related articles
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://dev.to/spread2009/what-is-agent-readiness-1b59"&gt;What Is Agent Readiness?&lt;/a&gt; — the foundational concept and why good APIs can be invisible to AI agents.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/spread2009/your-api-has-seo-does-it-have-agent-readiness-2m14"&gt;Your API Has SEO. Does It Have Agent Readiness?&lt;/a&gt; — why SEO optimization isn't enough for agent discoverability, and 10 things to check.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/spread2009/the-web-is-becoming-agentic-what-happens-to-api-discovery-ddl"&gt;The Web Is Becoming Agentic. What Happens to API Discovery?&lt;/a&gt; — the emerging discovery stack for the agentic web.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://dev.to/spread2009/from-seo-to-geo-to-agent-readiness-31mj"&gt;From SEO to GEO to Agent Readiness&lt;/a&gt; — the evolution of optimization: from websites to content to APIs.&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Don't certify. Measure.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>agents</category>
      <category>openapi</category>
    </item>
  </channel>
</rss>
