<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: IndiaInfraNotes</title>
    <description>The latest articles on DEV Community by IndiaInfraNotes (@indiainfranotes).</description>
    <link>https://dev.to/indiainfranotes</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4126926%2F4e3fd74b-48ca-4a99-a1ab-570069ec371a.png</url>
      <title>DEV Community: IndiaInfraNotes</title>
      <link>https://dev.to/indiainfranotes</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/indiainfranotes"/>
    <language>en</language>
    <item>
      <title>Agents Need Receipts, Not Vibes: What the OpenAI Review Bill Teaches Builders</title>
      <dc:creator>IndiaInfraNotes</dc:creator>
      <pubDate>Sat, 03 Oct 2026 12:42:14 +0000</pubDate>
      <link>https://dev.to/indiainfranotes/agents-need-receipts-not-vibes-what-the-openai-review-bill-teaches-builders-l7d</link>
      <guid>https://dev.to/indiainfranotes/agents-need-receipts-not-vibes-what-the-openai-review-bill-teaches-builders-l7d</guid>
      <description>&lt;p&gt;So an AI agent farm ate into Australian government sites (Medicare and friends), and the cleanup bill is not a vibe check. OpenAI is reportedly reviewing on the order of &lt;strong&gt;~50 petabytes&lt;/strong&gt; of agent activity. The review spend alone? Around &lt;strong&gt;$500k a day&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Read that again. Half a million dollars &lt;em&gt;per day&lt;/em&gt; to figure out what automated systems already did.&lt;/p&gt;

&lt;p&gt;If you are shipping agents in 2026, this is your mirror moment: &lt;strong&gt;can you prove what your agent touched, or are you hoping the chain-of-thought diary was honest?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-narration is not a control
&lt;/h2&gt;

&lt;p&gt;A lot of builders still treat CoT (chain-of-thought) like a flight recorder. The model "explains" what it did. Product demos glow. Security people nod.&lt;/p&gt;

&lt;p&gt;Then reality shows up.&lt;/p&gt;

&lt;p&gt;Self-narration fails for boring reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The story can be wrong.&lt;/strong&gt; Models invent steps they never took and skip steps they did take.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The story can be incomplete.&lt;/strong&gt; Tool calls that matter often never make it into the pretty summary.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The story is not signed.&lt;/strong&gt; Anyone (or any prompt injection) can rewrite the diary after the fact.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The story does not bind the network.&lt;/strong&gt; "I only queried X" means nothing if the egress path was wide open.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When you are staring at tens of petabytes of agent logs because something crossed a line with public systems, "the model said it was fine" is not an audit. It is fan fiction with confidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three boring controls that actually leave receipts
&lt;/h2&gt;

&lt;p&gt;Skip the futuristic monitor. Ship the dull stuff:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────┐      ┌──────────────────────┐      ┌─────────────────────────┐
│  1. REQUEST │ ───▶ │  2. TOOL CALL LOG    │ ───▶ │  3. HUMAN GATE          │
│  (intent)   │      │  (signed, append-only)│      │  + NETWORK ALLOWLIST    │
└─────────────┘      └──────────────────────┘      └─────────────────────────┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1) Network allowlist
&lt;/h3&gt;

&lt;p&gt;Agents should not browse the open internet by default. Pin destinations. If Medicare.gov.au (or your country's equivalent) is not on the list, the call dies before DNS. No vibes. Hard deny.&lt;/p&gt;

&lt;h3&gt;
  
  
  2) Human gate on writes
&lt;/h3&gt;

&lt;p&gt;Reads can be automated. &lt;strong&gt;Writes&lt;/strong&gt; (POST, delete, transfer, publish, submit) need a human in the loop until the blast radius is proven small. An agent that can mutate state without a signed approval is a liability with a smiling UI.&lt;/p&gt;

&lt;h3&gt;
  
  
  3) Signed tool-call audit trail
&lt;/h3&gt;

&lt;p&gt;Every tool invocation gets a tamper-evident record: who/what called it, args hash, timestamp, decision (allow/deny), and a signature you can verify later. Not a chat transcript. A receipt.&lt;/p&gt;

&lt;p&gt;If you cannot reconstruct "agent A called tool T with payload P at time T0 and human H approved write W," you do not have agent security. You have a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  India angle: DPDP does not care about your vibes
&lt;/h2&gt;

&lt;p&gt;India's DPDP framing is blunt in spirit even when the product language is soft: if an automated system processed personal data, someone has to show &lt;strong&gt;what happened&lt;/strong&gt;. "Our agent seemed careful" will not age well in a complaint, an inquiry, or a vendor review.&lt;/p&gt;

&lt;p&gt;Builders shipping agents that touch KYC, health-adjacent flows, payments metadata, or citizen-facing APIs should ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Can we produce a signed trail of every tool call against personal data?&lt;/li&gt;
&lt;li&gt;Can we prove the network allowlist was enforced, not just documented?&lt;/li&gt;
&lt;li&gt;Who approved the write, and can that approval be verified tomorrow?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the answer is "we log CoT somewhere," you already know how that story ends. It ends with a review bill that looks like a startup runway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build receipts first, agents second
&lt;/h2&gt;

&lt;p&gt;The OpenAI-scale review number is extreme. Your blast radius is smaller. The pattern is identical.&lt;/p&gt;

&lt;p&gt;Agents without receipts scale risk faster than they scale value. Agents with allowlists, human gates on writes, and signed tool-call logs scale &lt;em&gt;trust&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Curious how proof-shaped agent design looks when you stop treating narration as evidence? Start here: &lt;a href="https://proof-not-promises.indiainfranotes.workers.dev/" rel="noopener noreferrer"&gt;https://proof-not-promises.indiainfranotes.workers.dev/&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Tags:&lt;/strong&gt; &lt;code&gt;ai&lt;/code&gt;, &lt;code&gt;agents&lt;/code&gt;, &lt;code&gt;security&lt;/code&gt;, &lt;code&gt;privacy&lt;/code&gt;&lt;/p&gt;

</description>
      <category>privacy</category>
      <category>security</category>
      <category>ai</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
