I will not let a paging model choose the first command, because the alert class should already have made that choice. The only question that matters in minute one is whether five read-only receipts exist before anyone requests a write. If those receipts are missing, I keep the shell frozen and I escalate on a clock rather than on a feeling. Have you watched a confident paragraph turn a noisy page into a restart that nobody had actually approved?
The page may change attention, not authority
A page is a request for attention, and I refuse to treat it as a grant of write access on the paged host. I want that sentence in the incident channel before the bridge call collects opinions, because opinions arrive faster than evidence. The card in this article is a proposal you can adapt, not a diary of an outage I am pretending to have run last night. If your team already runs a stricter change gate, keep that gate, and steal only the receipt idea from this card.
Three constraints sit above the commands, and I repeat them whenever someone offers a shortcut around the card.
- Intake may name an alert class, but intake may not mutate hosts, flags, or queues.
- The first five commands are read-only, and the class picks them before a person or a model starts improvising.
- Unfreeze covers one named command, and only after receipts plus an escalation ack exist in the channel.
Faster drafts did not retire the freeze
I keep seeing discussion, including this week, about models sounding smoother while the software around them stayed stubborn. I will not quote a leaderboard I have not opened, and I will not treat a headline as a primary source for your pager. The practical lesson I am willing to keep is narrower than the headline, and it is about authority rather than fluency. A cleaner paragraph does not earn a write on a sick host, even when the alert text was copied perfectly into the prompt.
If the draft is fluent and the receipts are still empty, the freeze still wins, and the escalation clock still runs. Why would smoother wording change the owner of a production shell? I want the answer written down before the bridge call turns confidence into a typed mutation.
Pick a class before you open a shell
I use three classes, because a single generic checklist hides the check that would actually explain this page. Class A means a customer-visible failure, or any credible risk that stored data was lost or corrupted. Class B means the error budget is burning, while nobody has confirmed an actual customer-facing break yet. Class C means a detector symptom that might belong to the detector itself, rather than to the service under the page.
Why would I start in a softer class when a wrong downgrade can hide a real customer break? Downgrading the page is itself a decision, and I would rather run one extra read than miss the checkout path. If I cannot defend the class in a single line, I stay on Class A until a human explicitly moves it. That bias is intentional, which is why the table below never shows an open write posture for any class.
| Class | What I need to see | First goal of the reads | Write posture |
|---|---|---|---|
| A | Errors, timeouts, or data doubt | Bound impact and name a secondary | Frozen |
| B | Budget burn, no confirmed break | Confirm scope and the recent trend | Frozen |
| C | Detector-only, no user report | Separate detector fault from host fault | Frozen |
The class changes the questions I ask, and it deliberately does not change the ban on mutation during intake. Would you unlock a restart for Class C merely because a draft called the alert simple? I would not, so the linter below treats every class as frozen until the receipts actually land.
Lint the first five commands on a laptop
I keep the allowlist in a script, so a tired paste cannot smuggle a write verb into the first minute of the page. The listing is unexecuted proposal code, and I would run it against a fixture on a laptop rather than on a production shell. It classifies one alert line, prints the matching card, and returns non-zero when a suggestion looks like a mutation. You should replace the paths before you trust it, because I have not executed this file on your hosts.
#!/usr/bin/env python3
"""Proposal: require read receipts before any unfreeze. Not a prod agent."""
import re, sys
WRITE = re.compile(
r"\b(kill|delete|drop|apply|restart|scale|patch|exec|truncate|systemctl)\b",
re.I,
)
CARDS = {
"A": [
"date -u",
"hostname",
"uptime",
"curl -fsS --max-time 3 http://127.0.0.1/health || true",
"tail -n 40 /var/log/app/app.log",
],
"B": [
"date -u",
"uptime",
"df -h",
"ps -eo pid,pcpu,pmem,comm --sort=-pcpu | head",
"tail -n 30 /var/log/app/app.log",
],
"C": [
"date -u",
"hostname",
"id",
"journalctl -u app -n 40 --no-pager || true",
"tail -n 20 /var/log/app/app.log",
],
}
def classify(text: str) -> str:
t = text.lower()
if any(k in t for k in ("data loss", "checkout", "5xx", "timeout")):
return "A"
if any(k in t for k in ("slo", "burn", "latency")):
return "B"
return "C"
def main() -> int:
alert = sys.stdin.read().strip()
kind = classify(alert)
print(f"class={kind} posture=frozen receipts=0/5")
for index, cmd in enumerate(CARDS[kind], start=1):
if WRITE.search(cmd):
print(f"reject card[{index}] {cmd}")
return 2
print(f"read[{index}] {cmd}")
suggested = " ".join(sys.argv[1:])
if suggested and WRITE.search(suggested):
print("reject suggestion: write seen before receipts")
return 3
print("hold unfreeze until five receipts and an escalation ack")
return 0
if __name__ == "__main__":
raise SystemExit(main())
A fixture committed beside the script should stay short enough to reread on a phone during the page.
alert: checkout 5xx ratio above page threshold for 3 minutes
The local check is supposed to feel boring, which is what I want when the bridge call is already loud.
python3 receipt_card.py systemctl restart app < fixture.txt
echo "exit=$?"
An exit status of 3 means the suggestion tried to skip the receipt gate, so I discard that suggestion immediately. An exit status of 0 means I earned a reading list, and I still have not earned permission to write. After each read, I append a receipt with a tiny helper instead of trusting memory in the middle of the night.
#!/bin/sh
# Proposal: record one read receipt. Usage: ./note_receipt.sh A "date -u" 0
printf 'RECEIPT class=%s cmd="%s" at=%s exit=%s\n' \
"$1" "$2" "$(date -u +%Y-%m-%dT%H:%M:%SZ)" "$3"
RECEIPT class=A cmd="date -u" at=2026-10-12T03:12:01Z exit=0
RECEIPT class=A cmd="hostname" at=2026-10-12T03:12:04Z exit=0
Escalation runs on a clock, not a mood
I start the escalation clock when the class line is posted, because a vague status is not a policy anyone can audit the next morning. In this proposal, Class A pages the secondary at ten minutes, Class B at twenty, and Class C at thirty unless customer chatter appears. Those durations are planning defaults for a small rotation, not measurements taken from your pager data. Edit the numbers on the card before you adopt them, and do not cite this article as if it were a study.
The order of the names matters more than the speed of any generated summary you might paste.
- Primary posts the class, the freeze, and the five-command card in the incident channel.
- Primary names the secondary in that same message, even when the escalation timer has not fired yet.
- If the timer fires without a secondary ack, primary pages them and still refuses writes.
- If impact stays unclear at the next interval, primary pulls the service owner into the channel.
Could a model draft that acknowledgement faster than I can type the same four fields myself? Yes, and the speed helps only when class, freeze, secondary, and deadline all remain intact. If the draft omits the secondary, I do not send it, even when the surrounding prose sounds calm and complete. A missing name is an escalation failure, not a style problem I should polish after the page is already late.
Unfreeze is one human-typed command
Freeze means no write-class command and no temporary scale, including a one-liner that a chatbot offers as the obvious fix. Unfreeze is not a change of mood, and a red graph does not imply permission the channel never granted. I allow one command, typed by the human who holds the page, only after every gate below is already true.
- Five read receipts are in the channel, and each receipt names the same class that intake recorded.
- The secondary, or the service owner, has acknowledged that same class in a written channel message.
- The command is absent from the read-only card, and a human typed it rather than pasting a model line untouched.
- The release names an end time, and the posture returns to frozen when that time passes without a fresh line.
I keep the release on one grep-friendly line, so a later review can see exactly what the page allowed.
UNFREEZE class=A receipts=5/5 ack=lee cmd="service app reload" until=2026-10-12T03:40:00Z by=sam
If any field is missing, I stay frozen, even when the chart looks worse than it did during intake. A worse chart is a reason to escalate the page, not a reason to skip the receipts you already required. Would you sign a blank unfreeze because a draft said the reload looked safe from the alert text alone? I want that answer recorded as no, in the same channel that received the original page.
Draft the card away from production credentials
Disclosure: This article was prepared as part of MonkeyCode's product outreach. I would draft and lint this card on a separate machine, and the operator says MonkeyCode offers free model access plus a free server option for that drafting work. I am not stating a token quota, a hardware shape, or an end date, because no primary source for those figures was attached here. If that option is closed when you read this, run the same linter on any laptop that cannot see production credentials.
What I would send is the alert sentence and the allowlist, with secrets left out of the prompt on purpose. What I would accept back is a class label or a rejection, never a live session on the host that was paged. If the draft inserts a write verb, the script's non-zero exit is the decision, and the drafting host does not get another vote. That split is the point of keeping generation away from the shell this runbook is trying to protect.
I am not using this section to rank tools, and I am not claiming the free option will still exist next quarter. Availability language here is operator-supplied, not a measurement I reran against a dated product page on the day of this draft.
Who should skip this card
This approach fits poorly when you lack an incident channel, a secondary, or a read-only check that stays safe on a sick host. It also fits poorly on safety-critical control gear, where my sample commands could be the wrong shape entirely. I am not claiming a shorter recovery, a better model score, or a replacement for the vendor runbook you already trust.
I keep the limits beside the suggestion, so a tired reader cannot skip them on the way to the code block.
- The classifier is only a keyword stub, and novel alert names will miss until you extend those lists yourself.
- Read-only is not automatically harmless, so the tail and curl examples stay capped, local, and easy to delete.
- A free drafting host may be down, rate-limited, or unfit for secrets, so never paste credentials into the prompt.
- The clocks above are proposals, not results from a published on-call study that you should quote as evidence.
If a command in the card does not exist on your hosts, delete it before the next page instead of discovering the gap live. I would rather you keep a shorter card than run an example that only matched an imaginary layout. The receipt rule still works after you swap in commands your platform team already trusts for diagnosis.
If the paper card has never been linted, run it once against a fixture before the next rotation, on a separate machine you control. A free drafting server is optional for that pass, and the freeze rule does not depend on which editor helped you type the card.
Top comments (0)