DEV Community

Orvi Das
Orvi Das

Posted on

Your AI Agent's Kill Switch Returns a Boolean. That's the Bug (Python Circuit Breaker)

Here's the loop that cost me an account:

def check_for_lockout(page) -> bool:
    if page.locator("text=Your account is locked").count():
        alert("🚨 LOCKOUT DETECTED")
        return True
    return False

while True:
    if check_for_lockout(page):
        continue          # skip this tick, try again later
    do_next_action(page)
Enter fullscreen mode Exit fullscreen mode

The detector worked. On 2026-07-21 the log line appeared, Telegram buzzed, and the selector matched exactly what it was supposed to match. Forty-one seconds later the while came back around. My bot kept poking a locked platform for another three and a half days.

Nothing here is an AI problem, and that's what makes it worth writing about. The model never made a bad call. The bug was a function that answered a question when it should have ended the program.

Return values are only suggestions

A detector that returns True hands the decision to whoever called it. That caller might be code you wrote at 2am, a retry wrapper, a try/except someone added later, or a scheduler that just restarts the process. Sooner or later one of them does the polite thing: skip one action, log it, carry on.

If a signal means "stop", it shouldn't be able to get through a caller who disagrees. It should end the process.

The same fix covers the bigger failure class. Most agent disasters aren't one catastrophic decision. They're ordinary actions repeated too many times for too long, with nothing counting them.

Knight Capital is the textbook case. On 1 August 2012 its router turned 212 retail orders into more than 4 million executions across 154 stocks: 397 million shares in 45 minutes, a loss of over $460 million, and later a $12 million SEC settlement (SEC 34-70694). Every order was valid on its own. What was missing was something that compared the orders sent this minute with the number that should have gone out.

The agent-era version: in July 2025, Replit's coding agent wiped a production database with records on 2,400+ executives and companies during a code freeze, then said rollback was impossible. That wasn't true (The Register). It's the same pattern. The loop has the authority and nothing upstream limits it.

What each common guard actually misses

Mechanism What it bounds What gets past it
Retry limit One call A loop of calls that all succeed
Provider rate limit Requests/min per key Slow, steady runaways. It throttles; it never stops
Budget alert Spend, after the fact Everything between the threshold and a human reading the email
Detector returning bool Nothing on its own Whatever the caller decides to do
Pre-flight breaker Actions/window, total actions, wall time, spend before the call Only lanes you forgot to route through it

The bottom row is the one you build.

A breaker that actually exits

# breaker.py
import json, os, sys, threading, time
from collections import deque
from pathlib import Path

class CircuitBreaker:
    def __init__(self, *, max_per_minute, max_total, max_runtime_s,
                 sentinel="BREAKER_TRIPPED.json"):
        self.max_per_minute = max_per_minute
        self.max_total = max_total
        self.max_runtime_s = max_runtime_s
        self.sentinel = Path(sentinel)
        self.started = time.monotonic()
        self.window = deque()
        self.total = 0
        self.lock = threading.Lock()
        if self.sentinel.exists():
            sys.exit(f"refusing to start: {self.sentinel} exists. "
                     "Read it, fix the cause, delete it by hand.")

    def trip(self, reason):
        self.sentinel.write_text(json.dumps({"reason": reason, "at": time.time()}))
        print(f"BREAKER TRIPPED: {reason}", file=sys.stderr, flush=True)
        os._exit(2)

    def acquire(self, lane):
        now = time.monotonic()
        with self.lock:
            if now - self.started > self.max_runtime_s:
                self.trip(f"runtime exceeded {self.max_runtime_s}s")
            while self.window and now - self.window[0] > 60:
                self.window.popleft()
            if len(self.window) >= self.max_per_minute:
                self.trip(f"{lane}: {len(self.window)} actions in 60s")
            if self.total >= self.max_total:
                self.trip(f"{lane}: hit total cap {self.max_total}")
            self.window.append(now)
            self.total += 1
Enter fullscreen mode Exit fullscreen mode

A few choices here look odd, so here's the reasoning.

os._exit, not sys.exit or raise. sys.exit raises SystemExit. A bare except: will catch it, and so will except BaseException. Called from a worker thread, it only ends that thread. os._exit ends the process right away. It skips finally blocks and buffered I/O, which is why the sentinel gets written and stderr flushed first.

Tripping on rate, not throttling. A normal rate limiter would sleep and try again. In an agent, going over your own expected rate usually means something is broken: a retry storm, a duplicated lane, a prompt stuck in a loop. Sleeping hides that. Stopping makes you look at it.

The sentinel file. If launchd or systemd runs with KeepAlive/Restart=always, a plain exit just gets restarted 10 seconds later, back into the same loop. The sentinel makes the trip survive restarts until a person deletes it.

Wiring it in, so the detector can no longer give a soft answer:

breaker = CircuitBreaker(max_per_minute=6, max_total=30, max_runtime_s=6 * 3600)

def check_for_lockout(page):
    if page.locator("text=Your account is locked").count():
        alert("🚨 LOCKOUT DETECTED")
        breaker.trip("platform lockout")   # does not return

while True:
    check_for_lockout(page)
    breaker.acquire("reply")
    do_next_action(page)
Enter fullscreen mode Exit fullscreen mode

Every lane (replies, posts, DMs, LLM calls) goes through one shared acquire. A per-lane cap can't see that three lanes at 80% each add up to 240%.

Spend: why 20 agents "under budget" blow the cap

LLM spend has its own trap. The obvious check looks like this:

if spent + estimate <= cap:
    resp = client.messages.create(...)
    spent += actual
Enter fullscreen mode Exit fullscreen mode

Run that across 20 concurrent workers and all 20 read spent before any of them writes it back. Each one is under budget, and together they're 20× over. You can't fix this by checking more often. You fix it by reserving the money before the call:

class Budget:
    def __init__(self, cap_usd):
        self.cap = cap_usd
        self.committed = 0.0
        self.reserved = 0.0
        self.lock = threading.Lock()

    def reserve(self, est):
        with self.lock:
            if self.committed + self.reserved + est > self.cap:
                raise RuntimeError(f"budget: would exceed ${self.cap:.2f}")
            self.reserved += est

    def settle(self, est, actual):
        with self.lock:
            self.reserved -= est
            self.committed += actual

def guarded_create(client, budget, **kw):
    # worst case: full prompt + max_tokens at output price
    est = estimate_cost(kw["model"], kw["messages"], kw["max_tokens"])
    budget.reserve(est)
    try:
        resp = client.messages.create(**kw)
    except Exception:
        budget.settle(est, 0.0)
        raise
    budget.settle(est, actual_cost(kw["model"], resp.usage))
    return resp
Enter fullscreen mode Exit fullscreen mode

Reserve the worst case, then settle to the real cost. Once the cap is reached, the provider is never called. Catch the RuntimeError at the top of your loop and pass it to breaker.trip() if a blown budget should stop the whole run, which it usually should.

If you'd rather not maintain this yourself, baar-core is the open-source version I use: an atomic pre-flight cap that returns a 402 before any request goes to the provider (pip install baar-core). For teams, noburn.dev builds on it with per-user budgets enforced before each call, so a user over their limit gets blocked instead of showing up on next month's invoice.

The test that would have saved me

Don't test that the detector returns True. Test that the process dies:

import subprocess, sys

def test_lockout_kills_process(tmp_path):
    script = tmp_path / "run.py"
    script.write_text(
        "from breaker import CircuitBreaker\n"
        "b = CircuitBreaker(max_per_minute=99, max_total=99, max_runtime_s=99,"
        f" sentinel=r'{tmp_path}/S.json')\n"
        "b.trip('test')\n"
        "print('STILL ALIVE')\n"
    )
    r = subprocess.run([sys.executable, str(script)], capture_output=True, text=True)
    assert r.returncode == 2
    assert "STILL ALIVE" not in r.stdout
    assert (tmp_path / "S.json").exists()
Enter fullscreen mode Exit fullscreen mode

The test checks the exit code, the missing output line and the sentinel file, and it runs in a few milliseconds. My original detector had tests too, and they all passed, because they only asked whether it noticed the lockout.

What's the dumbest way a safety check in your system has "worked" without actually stopping anything?


Originally published at https://robatdasorvi.com/stories/why-every-agentic-loop-needs-a-circuit-breaker-and-how-i-built-one

Top comments (3)

Collapse
 
pm25coder profile image
pm25coder •

The detector bug is right, and I think the fix moves the ambiguity one step rather than removing it. Two readings, taken by running it.

Exiting relocates the signal into the exit code. The exit code is a namespace you already share with "something went wrong". A supervisor that reads it — which is exactly the "scheduler that just restarts the process" in your own list of callers who disagree — sees the same state:

child intent returncode stderr
sys.exit(1) after the breaker trips stop, do not come back 1 0 B
unhandled exception while checking crash, restart me 1 128 B
crash with stderr discarded crash, restart me 1 0 B

The last row is the one that decides it: the return code cannot separate them, and stderr is not a contract — a process killed by a signal, or one whose output goes to a log, gives the same empty stderr as a deliberate trip. So "it should end the process" is right about the caller, and it hands the new reader a channel it was already using to mean something else.

I hit the production version of this. A tool had reserved exit 1 to mean this run had no failed jobs. It died while printing the failure report — the bytes the console code page could not represent — so the print raised, the process exited 1, and the caller read "no failures" about a run that had failures. 56 of 6,236 lines were unprintable in the caller's decoder. Ending the process is what let the crash impersonate the verdict.

Your sentinel is the right shape, for that reason — it is a channel the supervisor was not already reading. But absence has two causes, and the obvious reader collapses them:

on disk naive read correct read
no sentinel go go
sentinel, parses stop stop
sentinel, half-written go stop — you do not know

A tripped-and-killed process leaves exactly the third row, when the writer is open(path, "w") followed by json.dump(...). "I could not parse it" defaulting to "not tripped" is the same bug one level down: the permissive answer is the one you get by accident. Fail closed on present-but-unreadable, and let only missing count as go.

Collapse
 
axiru profile image
Axiru •

Making the stop end the process is the fix most bots never get. Great Knight Capital example.

We apply the same idea to money: once a refund or payment is denied, a retry wrapper can't send it again without a person.

Does your breaker count actions per hour, or only watch for the lockout text?

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to