DEV Community

Anurag
Anurag

Posted on

I Audited My Own AI-Agent Library Against a Real Bug Report — Here's What I Found

I maintain StateGuard, a small Python library that gives AI agents transactional state — snapshot, check invariants, roll back and undo side effects if something's wrong. Instead of just adding features, I spent a few days trying to break my own design against a real bug report filed against an unrelated open-source project, and found four bugs in my own code in the process. Writing them up because the bug classes are more interesting than the library itself.

1. A module-level "current transaction" is a race condition waiting to happen

The active saga was tracked as a plain global. Fine under one request at a time. Under load, with overlapping sagas in different threads or asyncio tasks, one request's rollback could get attached to a different request's transaction — silently. The fix was a ContextVar instead of a global, scoped per thread/task automatically. I wrote a test that starts two overlapping sagas on purpose, with one committing while the other is mid-rollback, to make sure this can't regress.

2. A "flexible" undo function signature is a foot-gun

I wanted compensations to accept whatever arguments made sense — undo(result), undo(state, result), undo(order_id, result). My first pass matched by position, which meant a compensation with a differently-ordered signature would silently receive the wrong argument. Fixed it to match by parameter name against the original step's arguments first, falling back to position only when names don't line up.

3. A failed "undo" was being logged and then ignored

If the compensating action itself raised (the exchange API is down, the file's already gone), the original code logged it at critical level and moved on. The caller had no way to know the rollback was incomplete. Now it raises a CompensationError chained to whatever caused the rollback, with a hook so you can route it to a retry queue instead of a log line nobody reads until it's too late.

4. Async in a sync context doesn't fail loudly enough

An async compensation run inside a plain with Saga(...) (not async with) can't be awaited. It used to just... not run. Now it's reported as a failed compensation instead of silently skipped.

None of these are exotic — they're the standard shape of "this only breaks under conditions I didn't test for." Repo, tests, and the full write-up of what changed: Github. Genuinely interested if anyone's hit variations of these in their own saga/rollback code.

Top comments (0)