Most AI agent demos die the moment reality hits: a restart, a rate limit, a crashed worker. The model is fine — the plumbing isn't.
I kept rebuilding the same four pieces for every agent project, so I wrote them once, as a small stdlib-only Python runtime. Here's what actually breaks, and how to handle it.
1. No task state -> a restart loses everything
If your agent holds its queue in memory, a single deploy wipes the run. The fix is a durable store. I use SQLite in WAL mode with two tables (tasks + events), and every state transition appends an audit row.
States are explicit:
ready -> running -> done
-> failed -> ready (retry)
-> failed -> dead (exhausted)
2. No retries -> one 429 and the run dies
Transient failures (HTTP 429 / 5xx) should be retried with exponential backoff; client errors (4xx) should fail fast. Exhausted tasks go to a dead-letter state with the last error preserved, so you can inspect instead of guess.
delay = base_delay * (2 ** attempt)
3. No lease/heartbeat -> a crashed worker wedges the queue
This is the one people forget. When a worker claims a task it takes a lease. It renews the lease with a heartbeat. If the worker dies, a reclaim job resets expired leases back to ready.
# run this on a schedule (cron every 5-15 min)
python cli.py --db state.db reclaim
4. Prompts with no schema -> garbage in, garbage out
Free-form LLM output is unparseable at scale. Validate every structured response against a JSON schema before it enters your pipeline. I keep 32 system prompts in 6 categories (orchestration, code review, security audit, extraction, analysis, verification) plus 3 schemas.
The whole thing runs offline
Here's real output from the kit — no API key needed for the demo:
$ python cli.py --db demo.db init
initialized: demo.db
$ python cli.py --db demo.db enqueue --title "audit-report" --body "CHECK=secret-scan"
t_7c3c2ed4
$ python cli.py --db demo.db worker --name worker-1 --demo --once
[DEMO MODE] synthetic offline client - not a live model
{"processed": "t_7c3c2ed4", "status": "done"}
$ python -m unittest discover -s tests
Ran 11 tests in 1.086s
OK
Going live is one config file plus your API key. The live clients refuse to fabricate output when credentials are missing — no silent mocks.
What's inside
- SQLite WAL task store with lease claim + heartbeat + reclaim
- Exponential-backoff retry with dead-letter
- Namespaced memory bridge with value dedupe
- OpenAI-compatible and Anthropic LLM clients
- 32 system prompts + 3 JSON schemas + machine index
- 2 example pipelines + 11 tests, stdlib only
I packaged the blueprint, the runtime, and the prompt library here if it saves you the same weeks:
Production AI Agent Architecture & Autonomous Prompt Chains — https://ancuboy.gumroad.com/l/production-ai-agent-kit/LAUNCH50
Use code LAUNCH50 for 50% off. Happy to answer questions about the retry/lease design in the comments.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.