This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange
What I Built
Servers crash at 3 AM. A human wakes up, SSHs in, types systemctl restart, and goes back to sleep. I wondered if the ticket could do that work instead.
Ops Brain is a server that files its own incident reports in Sanity and heals itself, but only after a human approves.
The loop:
- A small Python agent watches my Ubuntu machine.
- nginx dies. The agent grabs the recent logs and creates an
incidentdocument in Sanity with a title, service, diagnosis, proposed action, risk level and log excerpt. - I open Sanity Studio and find a ticket from a server that reported its own death.
- I flip the status to
approvedand hit Publish. - The agent sees my approval, restarts nginx, checks that it's actually healthy, and stamps the ticket
verifiedwith a resolved time.
The server files the ticket, I sign it, and the server fixes itself.
Demo
{https://drive.google.com/file/d/1JbS2bEXw86HEF122E8LWS6lPCNezKAyf/view?usp=drivesdk}
Live Studio: https://opsbrain-benny.sanity.studio/
Code
{https://github.com/Bencoy09/ops-brain }
How I Used Sanity
The whole project rests on one decision: the workflow is data, not code.
An incident is a document that moves through states: awaiting_approval → approved → verified (or failed). The agent and a human push the same document forward. There's no separate approvals app and no webhook maze. The document is the process.
The schema is a single incident type:
-
service,title,logExcerpt: what broke, plus the evidence -
diagnosis,action,risk: what the agent proposes -
status: the workflow state, shown as radio buttons in Studio so approving takes one click -
detectedAt,resolvedAt: a built-in timeline of how long each outage lasted
Because incidents are structured content, I can ask GROQ questions like "show me everything that failed this week" without grepping a single log file.
My Build Process
The one rule I wouldn't bend: the agent never runs arbitrary commands. It executes only actions from a hard-coded allowlist (restart_service on explicitly watched services). Anything else is blocked and marked failed. An agent that obeys whatever text lands in a database is a remote-code-execution bug waiting to happen, so the allowlist is part of the design.
What went wrong (the honest part):
-
npm kept timing out. My connection dropped
create sanitymid-install. The project and dataset had already been created, so I reran onlynpm installwith--maxsockets=3and it went through. -
A half-written package broke
sanity deploy. An interrupted install leftzodcorrupted. Deleting it and reinstalling that one package fixed it. -
I pasted a shell command into my Python file. Line 1 was
cat > ..., and Python was not amused. I deleted the stray line. - The classic gotcha: the agent only sees published documents. I approved an incident, nothing happened, and the cause was that I hadn't clicked Publish.
-
I leaked my own token while testing. I typed it on the command line, so I revoked it and switched to a hidden prompt (
read -s) for the new one.
Tools: I used Claude to plan the architecture and debug each error as it appeared. The rest is deliberately plain: WSL Ubuntu, nginx, Python's standard library and Sanity Studio.
What I'd build next:
- An Approve button inside Studio using the App SDK, instead of changing a dropdown.
- A learning step where, after a successful fix, the agent drafts a new runbook entry for me to approve.
- More watched services, with risk levels that decide what needs a human and what doesn't.
Sanity Project Details
- Project ID:
hhoxq5y4 - Dataset:
production(public) - Query the incidents yourself: https://hhoxq5y4.api.sanity.io/v2025-02-19/data/query/production?query=*[_type=="incident"]
Top comments (0)