DEV Community

Cover image for My Server Files Its Own Tickets, Then Waits for My Signature to Heal Itself
Benny
Benny

Posted on

My Server Files Its Own Tickets, Then Waits for My Signature to Heal Itself

Sanity Challenge Path Two Submission

This is a submission for the Sanity Challenge, Path Two: Vibe-Code Something Strange

What I Built

Servers crash at 3 AM. A human wakes up, SSHs in, types systemctl restart, and goes back to sleep. I wondered if the ticket could do that work instead.

Ops Brain is a server that files its own incident reports in Sanity and heals itself, but only after a human approves.

The loop:

  1. A small Python agent watches my Ubuntu machine.
  2. nginx dies. The agent grabs the recent logs and creates an incident document in Sanity with a title, service, diagnosis, proposed action, risk level and log excerpt.
  3. I open Sanity Studio and find a ticket from a server that reported its own death.
  4. I flip the status to approved and hit Publish.
  5. The agent sees my approval, restarts nginx, checks that it's actually healthy, and stamps the ticket verified with a resolved time.

The server files the ticket, I sign it, and the server fixes itself.

Demo

{https://drive.google.com/file/d/1JbS2bEXw86HEF122E8LWS6lPCNezKAyf/view?usp=drivesdk}

Live Studio: https://opsbrain-benny.sanity.studio/

Code

{https://github.com/Bencoy09/ops-brain }

How I Used Sanity

The whole project rests on one decision: the workflow is data, not code.

An incident is a document that moves through states: awaiting_approval → approved → verified (or failed). The agent and a human push the same document forward. There's no separate approvals app and no webhook maze. The document is the process.

The schema is a single incident type:

  • service, title, logExcerpt: what broke, plus the evidence
  • diagnosis, action, risk: what the agent proposes
  • status: the workflow state, shown as radio buttons in Studio so approving takes one click
  • detectedAt, resolvedAt: a built-in timeline of how long each outage lasted

Because incidents are structured content, I can ask GROQ questions like "show me everything that failed this week" without grepping a single log file.

My Build Process

The one rule I wouldn't bend: the agent never runs arbitrary commands. It executes only actions from a hard-coded allowlist (restart_service on explicitly watched services). Anything else is blocked and marked failed. An agent that obeys whatever text lands in a database is a remote-code-execution bug waiting to happen, so the allowlist is part of the design.

What went wrong (the honest part):

  • npm kept timing out. My connection dropped create sanity mid-install. The project and dataset had already been created, so I reran only npm install with --maxsockets=3 and it went through.
  • A half-written package broke sanity deploy. An interrupted install left zod corrupted. Deleting it and reinstalling that one package fixed it.
  • I pasted a shell command into my Python file. Line 1 was cat > ..., and Python was not amused. I deleted the stray line.
  • The classic gotcha: the agent only sees published documents. I approved an incident, nothing happened, and the cause was that I hadn't clicked Publish.
  • I leaked my own token while testing. I typed it on the command line, so I revoked it and switched to a hidden prompt (read -s) for the new one.

Tools: I used Claude to plan the architecture and debug each error as it appeared. The rest is deliberately plain: WSL Ubuntu, nginx, Python's standard library and Sanity Studio.

What I'd build next:

  • An Approve button inside Studio using the App SDK, instead of changing a dropdown.
  • A learning step where, after a successful fix, the agent drafts a new runbook entry for me to approve.
  • More watched services, with risk levels that decide what needs a human and what doesn't.

Sanity Project Details

Top comments (0)