<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: pranav-afk</title>
    <description>The latest articles on DEV Community by pranav-afk (@pranavafk).</description>
    <link>https://dev.to/pranavafk</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4037827%2F002b9868-459c-417a-91b4-8b95f1b1e051.jpg</url>
      <title>DEV Community: pranav-afk</title>
      <link>https://dev.to/pranavafk</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pranavafk"/>
    <language>en</language>
    <item>
      <title>Traced a Multi Agent Failure Back 6 Steps. Here's the Instrumentation That Made It Possible.</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Mon, 17 Aug 2026 15:56:41 +0000</pubDate>
      <link>https://dev.to/pranavafk/traced-a-multi-agent-failure-back-6-steps-heres-the-instrumentation-that-made-it-possible-2obb</link>
      <guid>https://dev.to/pranavafk/traced-a-multi-agent-failure-back-6-steps-heres-the-instrumentation-that-made-it-possible-2obb</guid>
      <description>&lt;p&gt;Last week an agent in a production pipeline I was debugging did something that looked, on the surface, completely irrational. Correct tool call, correct parameters, valid output, but the decision to make that call at all was wrong. It took six steps of tracing backward to find out why.&lt;br&gt;
Here's the actual root cause, and why most observability setups wouldn't have caught it.&lt;br&gt;
The setup: a multi agent pipeline where Agent A researches a topic, writes findings to shared memory, and Agent D (four hops later, unrelated task) reads from that same memory scope because the key happened to overlap. Agent A's findings were accurate when written. By the time Agent D read them, the underlying data had changed. Agent D reasoned perfectly, off information that was stale by the time it mattered.&lt;br&gt;
Why this is hard to catch:&lt;br&gt;
The failing agent's logs look completely normal, valid input, valid reasoning, valid output&lt;br&gt;
Standard tracing shows what happened at each step, not when a piece of context was written versus when it was consumed&lt;br&gt;
Nobody flagged it because no individual step was wrong, the failure only exists in the relationship between two steps that happened at different times&lt;br&gt;
What actually solved it:&lt;br&gt;
Timestamp every memory write and read separately, and diff the gap when a failure investigation starts, a large gap between write and read is often the first real clue&lt;br&gt;
Scope memory access explicitly rather than relying on implicit key matching, Agent D shouldn't have been able to see Agent A's write in the first place if the scopes were correctly separated&lt;br&gt;
Make replay possible from any single node, so you can re-run just the suspect step with the memory state as it existed at read time, not the current state&lt;br&gt;
This kind of cross temporal bug is becoming more common as pipelines get longer and memory gets shared across more agents. Most tracing tools are built for single agent debugging and don't surface this class of issue well.&lt;br&gt;
We ended up building this kind of instrumentation into Cartha, our own agent governance layer, after running into this exact problem across a few different customer pipelines. Curious if others have hit similar cross temporal failures, or if this is more of an edge case in typical setups.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>machinelearning</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Debugging Multi-Agent Systems: Why "It Works" Isn't Enough Anymore</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Mon, 17 Aug 2026 02:51:42 +0000</pubDate>
      <link>https://dev.to/pranavafk/debugging-multi-agent-systems-why-it-works-isnt-enough-anymore-3di5</link>
      <guid>https://dev.to/pranavafk/debugging-multi-agent-systems-why-it-works-isnt-enough-anymore-3di5</guid>
      <description>&lt;p&gt;As agent systems move from single-chain prototypes to multi-agent production pipelines, a new failure mode has become common: the agent technically did what it was told, but the outcome is still wrong.&lt;br&gt;
Debugging a single LLM call is straightforward. You look at the prompt, the response, done. Debugging a multi-agent system is a different problem entirely — a request that gets handed off through 4-5 agents can fail because:&lt;br&gt;
The wrong tool got selected (but the tool call itself succeeded)&lt;br&gt;
Retrieved context was accurate but stale&lt;br&gt;
One agent wrote to a memory scope that a different agent read from later, for an unrelated reason&lt;br&gt;
The final output looks fine, but the reasoning path that produced it wasn't&lt;br&gt;
The core issue: most tracing tools log what happened, not why.&lt;br&gt;
A trace showing "Agent B called Tool X with these params" tells you the call happened. It doesn't tell you why Agent B decided to call Tool X, what context it believed was true at that moment, or whether that context was still valid.&lt;br&gt;
What actually helps:&lt;br&gt;
Tag memory events with intent, not just content — log why a write happened, not just what was written&lt;br&gt;
Make every step replayable in isolation — you should be able to fork a run from any node and re-execute just that segment&lt;br&gt;
Separate "tool call succeeded" from "tool call was the right call" — these get conflated in most logging setups, and it's the second one that actually matters for debugging&lt;br&gt;
We ran into this building Cartha, a governance and observability layer for AI agent fleets, and it changed how we think about tracing — treating memory scope and decision rationale as first-class, queryable data rather than something buried in a log line.&lt;br&gt;
Curious how others are handling this — are you building custom instrumentation, or leaning on LangSmith/Langfuse and accepting the gap?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Add Traces, Hard Budgets, and Scoped Memory to AI Agents in Python</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Sun, 02 Aug 2026 04:36:34 +0000</pubDate>
      <link>https://dev.to/pranavafk/how-to-add-traces-hard-budgets-and-scoped-memory-to-ai-agents-in-python-42po</link>
      <guid>https://dev.to/pranavafk/how-to-add-traces-hard-budgets-and-scoped-memory-to-ai-agents-in-python-42po</guid>
      <description>&lt;p&gt;If you ship multi-step &lt;strong&gt;AI agents&lt;/strong&gt; with tools and LLMs, you eventually need more than print statements: &lt;strong&gt;traces&lt;/strong&gt;, &lt;strong&gt;cost ceilings&lt;/strong&gt;, &lt;strong&gt;tool permissions&lt;/strong&gt;, and &lt;strong&gt;memory that does not leak across users&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;This guide shows a practical &lt;strong&gt;Python&lt;/strong&gt; integration pattern for &lt;strong&gt;AI agent observability and governance&lt;/strong&gt; using the open &lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt; platform and the &lt;a href="https://pypi.org/project/cartha-sdk/" rel="noopener noreferrer"&gt;&lt;code&gt;cartha-sdk&lt;/code&gt;&lt;/a&gt; package.&lt;/p&gt;

&lt;p&gt;Official product and dashboard: &lt;strong&gt;&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;https://cartha.in&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Docs: &lt;strong&gt;&lt;a href="https://cartha.in/documentation" rel="noopener noreferrer"&gt;https://cartha.in/documentation&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Pricing / trial: &lt;strong&gt;&lt;a href="https://cartha.in/pricing" rel="noopener noreferrer"&gt;https://cartha.in/pricing&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;
Login / API keys: &lt;strong&gt;&lt;a href="https://cartha.in/login" rel="noopener noreferrer"&gt;https://cartha.in/login&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;


&lt;h2&gt;
  
  
  What is an AI agent control plane?
&lt;/h2&gt;

&lt;p&gt;An &lt;strong&gt;AI agent control plane&lt;/strong&gt; sits beside your agents (CrewAI, LangGraph, OpenAI, custom code). You keep building agents as usual. The control plane:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Records &lt;strong&gt;execution traces&lt;/strong&gt; (tools, LLM calls, errors)&lt;/li&gt;
&lt;li&gt;Enforces &lt;strong&gt;hard budget breakers&lt;/strong&gt; (stop runaway spend)&lt;/li&gt;
&lt;li&gt;Applies &lt;strong&gt;tool allow-lists&lt;/strong&gt; (block unauthorized tools before they run)&lt;/li&gt;
&lt;li&gt;Stores &lt;strong&gt;scoped memory&lt;/strong&gt; (&lt;code&gt;user&lt;/code&gt; / &lt;code&gt;agent&lt;/code&gt; / &lt;code&gt;team&lt;/code&gt; / &lt;code&gt;org&lt;/code&gt;) with server-side isolation&lt;/li&gt;
&lt;li&gt;Exposes an ops &lt;strong&gt;dashboard&lt;/strong&gt; for agents, traces, memory, and costs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt; is built as that control plane for production agent systems—not a replacement for your agent framework.&lt;/p&gt;


&lt;h2&gt;
  
  
  Why Python AI agents need observability and budgets
&lt;/h2&gt;

&lt;p&gt;Search traffic and real teams ask the same questions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Problem&lt;/th&gt;
&lt;th&gt;What breaks without a control plane&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Debugging&lt;/td&gt;
&lt;td&gt;“Which tool failed?” needs step-level traces&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost&lt;/td&gt;
&lt;td&gt;Agent loops can burn API spend overnight&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Safety&lt;/td&gt;
&lt;td&gt;Support bots must not call wire-transfer tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-tenant apps&lt;/td&gt;
&lt;td&gt;Customer A memory must never appear for Customer B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audit&lt;/td&gt;
&lt;td&gt;Compliance needs “what did the agent know?”&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Logging alone is not enough. You want &lt;strong&gt;server-enforced&lt;/strong&gt; budgets, policies, and scopes—visible at &lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;cartha.in&lt;/a&gt;.&lt;/p&gt;


&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Python 3.10+&lt;/li&gt;
&lt;li&gt;Free workspace + API key from &lt;strong&gt;&lt;a href="https://cartha.in/login" rel="noopener noreferrer"&gt;Cartha login&lt;/a&gt;&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Optional: OpenAI key if you use &lt;code&gt;wrap_openai()&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;


&lt;h2&gt;
  
  
  Step 1 — Install Cartha SDK
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;cartha-sdk
&lt;span class="c"&gt;# optional&lt;/span&gt;
pip &lt;span class="nb"&gt;install &lt;/span&gt;openai

&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CARTHA_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"cartha_..."&lt;/span&gt;           &lt;span class="c"&gt;# from https://cartha.in → Settings / Keys&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;CARTHA_API_BASE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://cartha.in"&lt;/span&gt;   &lt;span class="c"&gt;# same host as the product&lt;/span&gt;

Full &lt;span class="nb"&gt;install &lt;/span&gt;notes and examples live &lt;span class="k"&gt;in &lt;/span&gt;the Cartha documentation &lt;span class="o"&gt;(&lt;/span&gt;https://cartha.in/documentation&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt;

───
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h3&gt;
  
  
  Step 2 — Minimal Python integration (copy-paste)
&lt;/h3&gt;

&lt;p&gt;This is the core AI agent instrumentation pattern: init → tool → trace → optional memory + OpenAI wrap.&lt;br&gt;
&lt;/p&gt;

&lt;p&gt;```import os&lt;br&gt;
import cartha&lt;/p&gt;

&lt;h1&gt;
  
  
  Connect to your Cartha workspace (&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;https://cartha.in&lt;/a&gt;)
&lt;/h1&gt;

&lt;p&gt;cartha.init(&lt;br&gt;
    api_key=os.environ["CARTHA_API_KEY"],&lt;br&gt;
    api_base=os.environ.get("CARTHA_API_BASE", "&lt;a href="https://cartha.in%22" rel="noopener noreferrer"&gt;https://cartha.in"&lt;/a&gt;),&lt;br&gt;
)&lt;/p&gt;

&lt;h1&gt;
  
  
  Auto LLM + cost steps for OpenAI chat completions
&lt;/h1&gt;

&lt;p&gt;client = cartha.wrap_openai()&lt;/p&gt;

&lt;p&gt;@cartha.tool()&lt;br&gt;
def crm_lookup(user_id: str) -&amp;gt; dict:&lt;br&gt;
    """Recorded as a tool step; blocked if not on allowed_tools."""&lt;br&gt;
    return {"user_id": user_id, "plan": "pro"}&lt;/p&gt;

&lt;p&gt;@cartha.trace(&lt;br&gt;
    id="support_agent",&lt;br&gt;
    team="support",&lt;br&gt;
    budget_usd=0.50,                 # hard agent-run budget&lt;br&gt;
    allowed_tools=["crm_lookup"],    # tool allow-list (authority)&lt;br&gt;
)&lt;br&gt;
async def handle_ticket(user_id: str, ticket: str) -&amp;gt; str:&lt;br&gt;
    # Scoped memory — user scope is isolated per customer&lt;br&gt;
    await cartha.remember(&lt;br&gt;
        user_id=user_id,&lt;br&gt;
        content=f"Ticket: {ticket}",&lt;br&gt;
        scope="user",&lt;br&gt;
        confidence=0.9,&lt;br&gt;
    )&lt;br&gt;
    hits = await cartha.recall(&lt;br&gt;
        user_id=user_id,&lt;br&gt;
        context=ticket,&lt;br&gt;
        scope=["user", "team"],&lt;br&gt;
        top_k=5,&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;data = crm_lookup(user_id)

r = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{
        "role": "user",
        "content": f"{ticket}\nCRM: {data}\nMemory: {hits}",
    }],
)
return r.choices[0].message.content or ""
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Run the function once, then open the Cartha dashboard (&lt;a href="https://cartha.in):" rel="noopener noreferrer"&gt;https://cartha.in):&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;• Agents — registration / heartbeats&lt;br&gt;
• Traces — step-by-step run replay&lt;br&gt;
• Memory — scoped store/recall&lt;br&gt;
• Costs — spend feeding budget breakers&lt;/p&gt;

&lt;p&gt;───&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

### Step 3 — Hard agent budgets (stop runaway spend)



```from cartha import BudgetExceeded

@cartha.trace(id="support_agent", team="support", budget_usd=0.01)
async def risky_run(user_id: str) -&amp;gt; str:
    # ... LLM / cost events ...
    return "ok"

try:
    await risky_run("user_123")
except BudgetExceeded as e:
    print("Cartha stopped the run:", e)

This is an agent-run budget (per @cartha.trace), not “hope the model stops.” Details: cartha.in/documentation (https://cartha.in/documentation).

───
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 4 — Tool allow-lists (authority before execution)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@cartha.trace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;support_agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;team&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;support&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;allowed_tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;crm_lookup&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;  &lt;span class="c1"&gt;# wire tools NOT listed
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;support_only&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;crm_lookup&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="c1"&gt;# calling a non-listed @cartha.tool → ToolNotAuthorized (before body runs)
&lt;/span&gt;
&lt;span class="n"&gt;Support&lt;/span&gt; &lt;span class="n"&gt;can&lt;/span&gt; &lt;span class="n"&gt;help&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;it&lt;/span&gt; &lt;span class="n"&gt;cannot&lt;/span&gt; &lt;span class="err"&gt;“&lt;/span&gt;&lt;span class="n"&gt;accidentally&lt;/span&gt;&lt;span class="err"&gt;”&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="n"&gt;privileged&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;those&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="n"&gt;are&lt;/span&gt; &lt;span class="n"&gt;outside&lt;/span&gt; &lt;span class="n"&gt;the&lt;/span&gt; &lt;span class="n"&gt;allow&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt; &lt;span class="n"&gt;That&lt;/span&gt;&lt;span class="err"&gt;’&lt;/span&gt;&lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="n"&gt;AI&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="n"&gt;governance&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="n"&gt;suggestion&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;

&lt;span class="err"&gt;───&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 5 — Scoped memory (multi-tenant isolation)
&lt;/h3&gt;

&lt;p&gt;┌───────┬───────────────────────────────┐&lt;br&gt;
│ Scope │ Who can use it                │&lt;br&gt;
├───────┼───────────────────────────────┤&lt;br&gt;
│ user  │ Memories for that user_id     │&lt;br&gt;
├───────┼───────────────────────────────┤&lt;br&gt;
│ agent │ Bound to agent identity       │&lt;br&gt;
├───────┼───────────────────────────────┤&lt;br&gt;
│ team  │ Shared within a team          │&lt;br&gt;
├───────┼───────────────────────────────┤&lt;br&gt;
│ org   │ Org-wide (when policy allows) │&lt;br&gt;
└───────┴───────────────────────────────┘&lt;/p&gt;

&lt;p&gt;Isolation is enforced on the Cartha API, not only in client code. Product overview: &lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;https://cartha.in&lt;/a&gt; (&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;https://cartha.in&lt;/a&gt;).&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Empty Memory Isn’t Missing Data — It’s a Permission Hole</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Sun, 26 Jul 2026 13:48:08 +0000</pubDate>
      <link>https://dev.to/pranavafk/empty-memory-isnt-missing-data-its-a-permission-hole-1ipa</link>
      <guid>https://dev.to/pranavafk/empty-memory-isnt-missing-data-its-a-permission-hole-1ipa</guid>
      <description>&lt;p&gt;&lt;strong&gt;Definition:&lt;/strong&gt; In multi-agent AI systems, an empty memory recall often means &lt;em&gt;this agent is not allowed to see that fact&lt;/em&gt; — not &lt;em&gt;the fact does not exist&lt;/em&gt;. If your stack turns both cases into &lt;code&gt;[]&lt;/code&gt;, the model will invent over a permission hole.&lt;/p&gt;

&lt;p&gt;That problem is one reason we built &lt;strong&gt;&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt;&lt;/strong&gt; — a governance-first agent operations platform (scoped memory, traces, hard budgets, cost attribution). Full setup guide: &lt;strong&gt;&lt;a href="https://cartha.in/how-to-use" rel="noopener noreferrer"&gt;How to Use Cartha&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The multi-agent failure mode
&lt;/h2&gt;

&lt;p&gt;Single-agent demos treat memory like a vector DB:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;embed → top-k → stuff into the prompt → hope.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That breaks as soon as you have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;more than one agent&lt;/li&gt;
&lt;li&gt;more than one user&lt;/li&gt;
&lt;li&gt;more than one team&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Example&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Support agent&lt;/li&gt;
&lt;li&gt;Finance agent&lt;/li&gt;
&lt;li&gt;Shared &lt;code&gt;user_id&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Support asks: “What’s this customer’s plan and last payment?”&lt;/p&gt;

&lt;p&gt;If finance memory is &lt;strong&gt;out of scope&lt;/strong&gt; for support, a naive stack returns &lt;strong&gt;nothing&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The support agent then:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;assumes nothing is known&lt;/li&gt;
&lt;li&gt;invents a plan&lt;/li&gt;
&lt;li&gt;or calls tools it shouldn’t&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You only notice when the customer is wronged — or when someone asks:&lt;br&gt;
&lt;strong&gt;“What did the agent know when it said that?”&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Silent empty vs explicit withhold
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;Agent sees&lt;/th&gt;
&lt;th&gt;Risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;silent&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Empty hits&lt;/td&gt;
&lt;td&gt;Hallucination over a permission hole&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;denied_hint&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Explicit “matched but withheld” (no content)&lt;/td&gt;
&lt;td&gt;Agent can refuse to invent&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For fleets, &lt;strong&gt;explicit withhold is the better default&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Not perfect privacy (existence can leak).&lt;br&gt;
But &lt;strong&gt;silent empty&lt;/strong&gt; trains agents to fake authority.&lt;/p&gt;

&lt;p&gt;A good denial sounds like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A finance-scoped memory matched, but this agent cannot read that scope. Do &lt;strong&gt;not&lt;/strong&gt; invent a substitute — treat this as withheld context.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s a &lt;strong&gt;permission event&lt;/strong&gt;, not missing data.&lt;/p&gt;

&lt;p&gt;On &lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt;, org default denial mode is &lt;strong&gt;&lt;code&gt;denied_hint&lt;/code&gt;&lt;/strong&gt; (&lt;code&gt;silent&lt;/code&gt; is opt-in in Settings). Details: &lt;a href="https://cartha.in/how-to-use" rel="noopener noreferrer"&gt;cartha.in/how-to-use&lt;/a&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Scopes that actually mean something
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Who can recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;user&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Agents serving that user&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;agent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Only this agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;team&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Agents sharing a team id&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;org&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Whole organization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The enum is not the product.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enforcement is.&lt;/strong&gt; If the agent can query a raw index with broader credentials, “scopes” are cosplay.&lt;/p&gt;

&lt;p&gt;Cartha enforces scopes &lt;strong&gt;server-side&lt;/strong&gt; on store/recall — not only in a system prompt.&lt;/p&gt;




&lt;h2&gt;
  
  
  What to log (for humans and for ops)
&lt;/h2&gt;

&lt;p&gt;When recall is denied, log:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;agent id&lt;/li&gt;
&lt;li&gt;requested scopes&lt;/li&gt;
&lt;li&gt;denial reason&lt;/li&gt;
&lt;li&gt;that content was withheld (not the secret content)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Otherwise you only get “the model made something up” with no path to &lt;em&gt;why&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;This pairs with multi-agent debugging: wrongness often starts at a &lt;strong&gt;boundary&lt;/strong&gt; (handoff &lt;strong&gt;or&lt;/strong&gt; memory boundary), not at the final sentence.&lt;/p&gt;




&lt;h2&gt;
  
  
  Minimal pattern (any framework)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
# Store with an explicit scope
await remember(user_id=uid, content="Prefers email", scope="user")

# Recall — still server-clamped
hits = await recall(
    user_id=uid,
    context="contact preference",
    scope=["user", "team"],
)

# Prefer structured denials, not only empty lists
# so the agent branches: "no access" vs "no data"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>python</category>
    </item>
    <item>
      <title>Empty memory isn’t missing data. It’s a permission hole.</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Sat, 25 Jul 2026 15:28:39 +0000</pubDate>
      <link>https://dev.to/pranavafk/empty-memory-isnt-missing-data-its-a-permission-hole-4j8h</link>
      <guid>https://dev.to/pranavafk/empty-memory-isnt-missing-data-its-a-permission-hole-4j8h</guid>
      <description>&lt;p&gt;&lt;strong&gt;Definition:&lt;/strong&gt; In multi-agent systems, an empty memory recall often means &lt;em&gt;this agent is not allowed to see that fact&lt;/em&gt; — not &lt;em&gt;the fact does not exist&lt;/em&gt;. Collapsing both into &lt;code&gt;[]&lt;/code&gt; is how agents invent over permission holes.&lt;/p&gt;

&lt;p&gt;This is one of the core problems we design for at &lt;strong&gt;&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt;&lt;/strong&gt; — an ops/governance layer for AI agent fleets (scoped memory, traces, costs, hard budgets). Docs: &lt;strong&gt;&lt;a href="https://cartha.in/how-to-use" rel="noopener noreferrer"&gt;How to Use Cartha&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The failure mode
&lt;/h2&gt;

&lt;p&gt;Most agent stacks treat memory like a vector DB with vibes:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;retrieve top-k → stuff into the prompt → hope the model behaves.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That works for a single demo agent.&lt;/p&gt;

&lt;p&gt;It fails when you have &lt;strong&gt;more than one agent&lt;/strong&gt;, &lt;strong&gt;more than one user&lt;/strong&gt;, or &lt;strong&gt;more than one team&lt;/strong&gt; sharing infrastructure.&lt;/p&gt;

&lt;p&gt;Classic setup:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Support agent&lt;/li&gt;
&lt;li&gt;Finance / billing agent&lt;/li&gt;
&lt;li&gt;Shared user id or “company memory”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Support asks: “What’s this user’s plan and last payment?”&lt;/p&gt;

&lt;p&gt;If finance-scoped memory is &lt;strong&gt;out of scope&lt;/strong&gt; for support, a naive stack returns &lt;strong&gt;nothing&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The support agent then:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;assumes nothing is known&lt;/li&gt;
&lt;li&gt;invents a plan&lt;/li&gt;
&lt;li&gt;or calls tools it shouldn’t&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You notice when a customer is wronged — or when someone asks: &lt;em&gt;what did the agent know when it said that?&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Silent empty vs explicit withhold
&lt;/h2&gt;

&lt;p&gt;Two denial modes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mode&lt;/th&gt;
&lt;th&gt;What the agent sees&lt;/th&gt;
&lt;th&gt;Risk&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;silent&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Empty hits&lt;/td&gt;
&lt;td&gt;Agent invents over a permission hole&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;denied_hint&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Explicit “matched but withheld” (no content)&lt;/td&gt;
&lt;td&gt;Agent can refuse to invent; may leak existence&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For agent fleets, &lt;strong&gt;explicit withhold is the better default&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Not perfect privacy (existence can leak).&lt;br&gt;
But &lt;strong&gt;silent empty&lt;/strong&gt; trains models to hallucinate authority they don’t have.&lt;/p&gt;

&lt;p&gt;A good denial sounds like:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A finance-scoped memory matched, but this agent cannot read that scope. Do &lt;strong&gt;not&lt;/strong&gt; invent a substitute — treat this as withheld context.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s a &lt;strong&gt;permission event&lt;/strong&gt;, not missing data.&lt;/p&gt;




&lt;h2&gt;
  
  
  Scopes that mean something
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Who can recall&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;user&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Agents serving that user&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;agent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Only this agent (private)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;team&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Agents sharing a team id&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;org&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Whole organization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The critical part isn’t the enum names.&lt;/p&gt;

&lt;p&gt;It’s &lt;strong&gt;enforcement on the server&lt;/strong&gt; — not “please only request allowed scopes” in a system prompt.&lt;/p&gt;

&lt;p&gt;If the agent can query a raw index with broader credentials, your “scopes” are cosplay.&lt;/p&gt;

&lt;p&gt;At &lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt;, memory scopes are enforced server-side; org default denial mode is &lt;strong&gt;&lt;code&gt;denied_hint&lt;/code&gt;&lt;/strong&gt; (&lt;code&gt;silent&lt;/code&gt; is opt-in). See the full walkthrough: &lt;a href="https://cartha.in/how-to-use" rel="noopener noreferrer"&gt;cartha.in/how-to-use&lt;/a&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Minimal shape (any stack)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
# Store with an explicit scope
await remember(user_id=uid, content="Prefers email", scope="user")

# Recall — requested scopes still get server-clamped
hits = await recall(
    user_id=uid,
    context="contact preference",
    scope=["user", "team"],
)

# Prefer structured denials, not only empty lists
# so the agent can branch: "no access" vs "no data"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>security</category>
      <category>programming</category>
    </item>
    <item>
      <title>The bug was at step 2. You noticed it at step 5.</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Fri, 24 Jul 2026 15:21:56 +0000</pubDate>
      <link>https://dev.to/pranavafk/the-bug-was-at-step-2-you-noticed-it-at-step-5-1gb9</link>
      <guid>https://dev.to/pranavafk/the-bug-was-at-step-2-you-noticed-it-at-step-5-1gb9</guid>
      <description>&lt;p&gt;Single-agent debugging is mostly a solved shape: see the prompt, see the tool call, see the output.&lt;/p&gt;

&lt;p&gt;Multi-agent graphs are different. The pattern that keeps biting teams:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Agent A hands off to agent B&lt;/li&gt;
&lt;li&gt;B calls a tool&lt;/li&gt;
&lt;li&gt;Result goes back to A&lt;/li&gt;
&lt;li&gt;A decides something based on it&lt;/li&gt;
&lt;li&gt;Three steps later the final answer is clearly wrong&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The failure is usually &lt;strong&gt;not&lt;/strong&gt; at step 5. The &lt;em&gt;wrongness&lt;/em&gt; happened at the &lt;strong&gt;handoff&lt;/strong&gt; (step 2). By the time you stare at the final output, you’re reconstructing the chain from logs by hand.&lt;/p&gt;

&lt;p&gt;That is the multi-agent debugging tax.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why handoffs break silently
&lt;/h2&gt;

&lt;p&gt;In a single chain, one component owns the state story.&lt;/p&gt;

&lt;p&gt;In a multi-agent run, state is often &lt;strong&gt;implicit&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A logs: “I passed the right thing.”&lt;/li&gt;
&lt;li&gt;B logs: “I received something reasonable.”&lt;/li&gt;
&lt;li&gt;Both look “correct” in isolation.&lt;/li&gt;
&lt;li&gt;The gap between &lt;strong&gt;what A intended&lt;/strong&gt; and &lt;strong&gt;what B actually operated on&lt;/strong&gt; is where silent failures live.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you only log each agent’s own outputs, that gap never shows up.&lt;/p&gt;




&lt;h2&gt;
  
  
  Three things that help (before perfect tooling)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Treat the handoff as a first-class event
&lt;/h3&gt;

&lt;p&gt;Log more than “agent A finished” / “agent B started”:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;decision to hand off (why)&lt;/li&gt;
&lt;li&gt;payload on the wire&lt;/li&gt;
&lt;li&gt;what the receiver was &lt;strong&gt;told&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;what the receiver &lt;strong&gt;could access&lt;/strong&gt; (tools, memory scopes, permissions)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last line matters. Payload can match while &lt;strong&gt;authority&lt;/strong&gt; still diverges (B can’t read the memory A assumed it could).&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Prefer step inspection over full re-run
&lt;/h3&gt;

&lt;p&gt;Re-running the whole graph to debug is slow and dangerous when tools have side effects (email sent, ticket updated, money moved).&lt;/p&gt;

&lt;p&gt;What you want first is: jump to the handoff step and inspect state &lt;strong&gt;there&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Detect divergence early (cheap)
&lt;/h3&gt;

&lt;p&gt;A useful pattern from practice: &lt;strong&gt;snapshot at the boundary&lt;/strong&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sender snapshot: what A put on the wire&lt;/li&gt;
&lt;li&gt;Receiver snapshot: what B bound as working state &lt;em&gt;before&lt;/em&gt; its first tool/LLM call&lt;/li&gt;
&lt;li&gt;Diff them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If they diverge, you catch it at step 2, not step 5. Cheaper than full replay: compare snapshots, don’t re-execute the graph.&lt;/p&gt;

&lt;p&gt;Fork-from-node (test a different tool/policy from that state without replaying upstream) is the recovery half. &lt;strong&gt;Detection&lt;/strong&gt; is the half most teams never build.&lt;/p&gt;




&lt;h2&gt;
  
  
  What this means for “agent observability”
&lt;/h2&gt;

&lt;p&gt;Generic “show me LLM spans” helps single chains.&lt;/p&gt;

&lt;p&gt;Multi-agent systems need &lt;strong&gt;boundary-aware&lt;/strong&gt; ops:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;nested / linked runs across agents&lt;/li&gt;
&lt;li&gt;what memory was visible at a decision&lt;/li&gt;
&lt;li&gt;what tools were authorized&lt;/li&gt;
&lt;li&gt;cost and retries rolled up per logical task&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s the direction we’re building toward with &lt;strong&gt;&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt;&lt;/strong&gt; (SDK-first: traces, scoped memory, budgets, multi-agent nest). We’re still early on full fork-from-node and automatic handoff divergence probes — and we’re honest about that.&lt;/p&gt;

&lt;p&gt;If you only take one idea away: &lt;strong&gt;log the handoff, not just the agents.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Questions for you
&lt;/h2&gt;

&lt;p&gt;If you run multi-agent graphs (LangGraph or otherwise):&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Do you log handoffs as first-class events, or only per-agent I/O?&lt;/li&gt;
&lt;li&gt;When you “replay,” is it full re-run or step-level inspect?&lt;/li&gt;
&lt;li&gt;Have you built any check that sender payload == receiver interpreted state?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Curious what’s worked (or failed) for you in production.&lt;/p&gt;

&lt;p&gt;Docs / try: &lt;a href="https://cartha.in/how-to-use" rel="noopener noreferrer"&gt;cartha.in/how-to-use&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>python</category>
      <category>automation</category>
    </item>
    <item>
      <title>Soft Alerts Don’t Stop Agent Spend. Hard Budgets Do.</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Wed, 22 Jul 2026 13:23:32 +0000</pubDate>
      <link>https://dev.to/pranavafk/soft-alerts-dont-stop-agent-spend-hard-budgets-do-4ifl</link>
      <guid>https://dev.to/pranavafk/soft-alerts-dont-stop-agent-spend-hard-budgets-do-4ifl</guid>
      <description>&lt;p&gt;Soft alerts are comforting.&lt;br&gt;
They are also how you wake up to a surprise OpenAI bill from a retry loop at 3am.&lt;/p&gt;

&lt;p&gt;If you run multi-step agents (tools + LLM calls), you need &lt;strong&gt;two&lt;/strong&gt; systems:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Observability&lt;/strong&gt; — what happened (traces, tokens, errors)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforcement&lt;/strong&gt; — what is allowed to happen &lt;em&gt;next&lt;/em&gt; (budget, tools, memory scope)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Most stacks are strong at (1) and weak at (2).&lt;/p&gt;




&lt;h2&gt;
  
  
  Soft alerts vs hard budgets
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;th&gt;Failure mode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Alert at 80% of monthly spend&lt;/td&gt;
&lt;td&gt;Email / Slack&lt;/td&gt;
&lt;td&gt;Agent keeps running&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Dashboard “cost this week”&lt;/td&gt;
&lt;td&gt;Human looks later&lt;/td&gt;
&lt;td&gt;Overnight burn&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hard budget on the run&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Reject next spend when ceiling hits&lt;/td&gt;
&lt;td&gt;Run stops; you debug&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Hard budgets only work if &lt;strong&gt;LLM/tool cost is recorded on the hot path&lt;/strong&gt; (or immediately after each call). A run with only “start/finish” spans cannot enforce real dollars.&lt;/p&gt;




&lt;h2&gt;
  
  
  Minimal mental model
&lt;/h2&gt;

&lt;p&gt;Agent run starts with budget_usd = 0.50&lt;br&gt;
  → tool call&lt;br&gt;
  → LLM call (+ tokens → $)&lt;br&gt;
  → if spent &amp;gt;= budget → fail closed (raise / stop)&lt;br&gt;
  → else continue&lt;/p&gt;

&lt;p&gt;Enforcement should not be “we'll audit later.”&lt;br&gt;
If the gate is only after the wire transfer (or the refund email), it's already too late.&lt;/p&gt;




&lt;h2&gt;
  
  
  Python sketch (pattern)
&lt;/h2&gt;

&lt;p&gt;You can implement this in app code. The important parts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Attach a &lt;strong&gt;budget&lt;/strong&gt; to the &lt;strong&gt;run&lt;/strong&gt; (not only org monthly total)&lt;/li&gt;
&lt;li&gt;Record &lt;strong&gt;every&lt;/strong&gt; LLM call’s estimated cost&lt;/li&gt;
&lt;li&gt;On exceed: &lt;strong&gt;raise&lt;/strong&gt; (or return a controlled error) — don’t log-and-continue&lt;/li&gt;
&lt;/ul&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
class BudgetExceeded(Exception):
    pass

class RunBudget:
    def __init__(self, limit_usd: float):
        self.limit = limit_usd
        self.spent = 0.0

    def charge(self, usd: float) -&amp;gt; None:
        if self.spent + usd &amp;gt; self.limit:
            raise BudgetExceeded(
                f"spent={self.spent:.4f} + {usd:.4f} &amp;gt; limit={self.limit}"
            )
        self.spent += usd

Wire that into your OpenAI wrapper after each completion (use usage tokens × your price table).
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>python</category>
      <category>ai</category>
      <category>llm</category>
      <category>devtools</category>
    </item>
    <item>
      <title>Hello DEV, I build ops for AI agent fleets</title>
      <dc:creator>pranav-afk</dc:creator>
      <pubDate>Mon, 20 Jul 2026 09:37:26 +0000</pubDate>
      <link>https://dev.to/pranavafk/hello-dev-i-build-ops-for-ai-agent-fleets-3l9o</link>
      <guid>https://dev.to/pranavafk/hello-dev-i-build-ops-for-ai-agent-fleets-3l9o</guid>
      <description>&lt;h2&gt;
  
  
  Hi, DEV 👋
&lt;/h2&gt;

&lt;p&gt;I'm part of the team building &lt;strong&gt;&lt;a href="https://cartha.in" rel="noopener noreferrer"&gt;Cartha&lt;/a&gt;&lt;/strong&gt; — an ops layer for AI agents that actually run in production.&lt;/p&gt;

&lt;p&gt;This is a short intro: the problem we kept hitting, what we shipped, and what I'll write about here.&lt;/p&gt;




&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;We were debugging an agent that did something inexplicable in production. Logs and traces answered &lt;em&gt;what ran&lt;/em&gt;. They didn't answer the question that mattered:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What did the agent know when it made that call — and was it allowed to know that?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you run multi-agent systems, you eventually need more than a pretty timeline:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Need&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Scoped memory&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;User vs agent vs team vs org — without leaks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hard budgets&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Soft alerts don't stop a retry loop at 3am&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Tool limits&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Child agents shouldn't inherit the keys to the kingdom&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;HITL&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Some actions need a human before they execute&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Observability tells you what happened. &lt;strong&gt;Governance&lt;/strong&gt; decides what is allowed to happen next.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Cartha is
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Not&lt;/strong&gt; a chatbot builder. You keep your own agents (OpenAI, tools, business logic).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Yes&lt;/strong&gt; a control plane for fleets:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Run timelines (tools + LLM steps)&lt;/li&gt;
&lt;li&gt;Server-enforced memory scopes (&lt;code&gt;user&lt;/code&gt; / &lt;code&gt;agent&lt;/code&gt; / &lt;code&gt;team&lt;/code&gt; / &lt;code&gt;org&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Fail-closed spend budgets on a run&lt;/li&gt;
&lt;li&gt;Nested agents + attenuated delegation&lt;/li&gt;
&lt;li&gt;Policies + escalations&lt;/li&gt;
&lt;li&gt;Dashboard for agents, traces, cost, memory&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Minimal Python path
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
python
import cartha

cartha.init()  # CARTHA_API_KEY + CARTHA_API_BASE=https://cartha.in
client = cartha.wrap_openai()  # auto LLM steps + cost

@cartha.tool()
def crm_lookup(user_id: str) -&amp;gt; dict:
    return {"plan": "pro"}

@cartha.trace(id="support", team="support", budget_usd=0.5)
def handle(user_id: str, ticket: str) -&amp;gt; str:
    data = crm_lookup(user_id)
    r = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": f"{ticket}\n{data}"}],
    )
    return r.choices[0].message.content or ""

Docs: How to Use (https://cartha.in/how-to-use)
Product: cartha.in (https://cartha.in)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>devtools</category>
      <category>python</category>
    </item>
  </channel>
</rss>
