On-Call Hero — Durable AI SRE Agent
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built On-Call Hero, an AI SRE agent that helps investigate and respond to incidents.
The main thing I wanted to solve was reliability. What happens when an AI agent is in the middle of handling an incident and the worker crashes or the server restarts?
On-Call Hero uses Temporal to make the workflow durable. It can investigate an issue, propose a fix, wait for human approval, execute it, verify the result, and roll back when needed.
Even if the worker is killed in the middle of an incident, the workflow can resume from where it stopped.
Demo
Everything runs locally.
make up
Then open http://127.0.0.1:8000 and try the S1 Bad deploy scenario.
You can also test recovery with:
make kill-worker
make up
Code
GitHub: https://github.com/Armansiddiqui9/oncallhero
How I Built It
The project uses Python, Temporal, FastAPI, SSE, and LiteLLM. It can also run with local models through Ollama.
I added human approval before actions are executed, along with plan validation, prompt-injection protection, idempotency, budgets, audit history, and rollback handling.
Why Open Innovation Matters
Most of the project is built around open-source tools, which made it possible to experiment with durable AI agents locally and understand every part of the system.
I'm hoping to keep improving it and would love feedback from the community.
Top comments (0)