DEV Community

Ryan Cole
Ryan Cole

Posted on

The 4 Things That Kill AI Agents in Production (and the Runtime That Survives Them)

Most AI agent demos die the moment reality hits: a restart, a rate limit, a crashed worker. The model is fine — the plumbing isn't.

I kept rebuilding the same four pieces for every agent project, so I wrote them once, as a small stdlib-only Python runtime. Here's what actually breaks, and how to handle it.

1. No task state -> a restart loses everything

If your agent holds its queue in memory, a single deploy wipes the run. The fix is a durable store. I use SQLite in WAL mode with two tables (tasks + events), and every state transition appends an audit row.

States are explicit:

ready -> running -> done
                 -> failed -> ready   (retry)
                 -> failed -> dead    (exhausted)
Enter fullscreen mode Exit fullscreen mode

2. No retries -> one 429 and the run dies

Transient failures (HTTP 429 / 5xx) should be retried with exponential backoff; client errors (4xx) should fail fast. Exhausted tasks go to a dead-letter state with the last error preserved, so you can inspect instead of guess.

delay = base_delay * (2 ** attempt)
Enter fullscreen mode Exit fullscreen mode

3. No lease/heartbeat -> a crashed worker wedges the queue

This is the one people forget. When a worker claims a task it takes a lease. It renews the lease with a heartbeat. If the worker dies, a reclaim job resets expired leases back to ready.

# run this on a schedule (cron every 5-15 min)
python cli.py --db state.db reclaim
Enter fullscreen mode Exit fullscreen mode

4. Prompts with no schema -> garbage in, garbage out

Free-form LLM output is unparseable at scale. Validate every structured response against a JSON schema before it enters your pipeline. I keep 32 system prompts in 6 categories (orchestration, code review, security audit, extraction, analysis, verification) plus 3 schemas.

The whole thing runs offline

Here's real output from the kit — no API key needed for the demo:

$ python cli.py --db demo.db init
initialized: demo.db
$ python cli.py --db demo.db enqueue --title "audit-report" --body "CHECK=secret-scan"
t_7c3c2ed4
$ python cli.py --db demo.db worker --name worker-1 --demo --once
[DEMO MODE] synthetic offline client - not a live model
{"processed": "t_7c3c2ed4", "status": "done"}
$ python -m unittest discover -s tests
Ran 11 tests in 1.086s
OK
Enter fullscreen mode Exit fullscreen mode

Going live is one config file plus your API key. The live clients refuse to fabricate output when credentials are missing — no silent mocks.

What's inside

  • SQLite WAL task store with lease claim + heartbeat + reclaim
  • Exponential-backoff retry with dead-letter
  • Namespaced memory bridge with value dedupe
  • OpenAI-compatible and Anthropic LLM clients
  • 32 system prompts + 3 JSON schemas + machine index
  • 2 example pipelines + 11 tests, stdlib only

I packaged the blueprint, the runtime, and the prompt library here if it saves you the same weeks:

Production AI Agent Architecture & Autonomous Prompt Chains — https://ancuboy.gumroad.com/l/production-ai-agent-kit/LAUNCH50

Use code LAUNCH50 for 50% off. Happy to answer questions about the retry/lease design in the comments.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.