DEV Community

Emery Huang
Emery Huang

Posted on

Block Write-Class Commands Until the Page Clock Names Owner and Expiry

I will not let an assistant draft become a shell command until a page clock names the owner, the severity, and a hard expiry. The opening stretch of a page should stay in a read class, because an early write is only a guess with privilege. Have you ever changed a unit file while the alert title was still a blur on the second monitor? I keep that impulse behind a command class gate, and I open the gate only when a human can point at the clock.

What the page clock is allowed to decide

A page clock is a small file that records who is on the hook, how severe the page looks, and when the read window ends. It does not diagnose the outage, and it does not bless a clever patch that a model just printed into chat. I use it as a gate in front of write-class commands, so a draft can sit on scratch storage away from the paging host. Would you trust a suggestion that arrived before anyone was willing to name the severity out loud?

I would not, which is why the clock has to exist before that suggestion is eligible for a human review. The file lives on a scratch path, not beside the unit that woke you up in the first place. I treat the production shell as frozen until a human replaces a closed write lock with an explicit unfreeze. If the clock file is missing, I keep reading and I escalate, rather than improvising a fix from a chat pane.

The clock file I want beside the runbook

The sample below is a proposal for the runbook repository, and I have not executed it against a live fleet. Copy it only into a scratch directory, then replace every placeholder with values taken from the page you received. Leave production hostnames out of this file on purpose, because a clock should never become a quiet map of targets. I would rather see an empty field than a hostname copied from a suggestion that nobody reviewed.

# page-clock.yaml: proposed artifact, not a live incident record
page_id: "replace-with-alert-id"
owner: "replace-with-human-name"
severity: "unknown"
opened_at: "replace-with-utc-timestamp"
read_window_minutes: 15
write_lock: "closed"
escalation_at: "replace-with-utc-timestamp"
draft_state: "held"
rollback_intent: "unset"
Enter fullscreen mode Exit fullscreen mode

First commands that stay in the read class

I start with commands that only observe, and I refuse anything that mutates packages, units, routes, or stored data. These examples assume a Linux host you already administer, and they stay proposals until your access policy allows the reads. I also keep a tiny wrapper that classifies a command string before I consider running that string for real. Have you noticed how a draft mixes a harmless status check with a restart on the next line?

  1. Confirm the page identity in the clock file before you open a process list, because a mismatched alert wastes the window.
  2. Reject the snapshot when owner or severity is still a placeholder, and spend that minute naming the fields instead.
  3. Collect service state into the scratch directory, and do not pipe that output into a restart or a package tool.
  4. Hold any command the classifier marks unknown, because an unknown verb is closer to a write than to a safe read.
# classify.sh: proposed read/write split for a scratch shell you control.
# Unknown commands fail closed. This example has not been run on a fleet.
set -eu
classify() {
  local cmd="$1"
  case "$cmd" in
    "systemctl status "*|"systemctl is-active "*|"journalctl "*|"ss "*|"ps "*|"cat "*|"ls "*)
      printf 'READ %s\n' "$cmd"
      ;;
    "systemctl restart "*|"systemctl stop "*|"apt "*|"dnf "*|"rm "*|"kubectl apply "*|"kubectl delete "*)
      printf 'WRITE blocked until the page clock opens: %s\n' "$cmd" >&2
      return 2
      ;;
    *)
      printf 'UNKNOWN hold for a human: %s\n' "$cmd" >&2
      return 3
      ;;
  esac
}
classify "journalctl -u demo.service --since 20 min ago"
classify "systemctl restart demo.service" || true
Enter fullscreen mode Exit fullscreen mode

I do not smuggle a restart into this window, even when a draft sounds sure about a cause it invented. A restart can hide the evidence the next human needed, and a package change can widen the failure without rollback intent. If the service is already down, the snapshot still comes first, and the write class stays blocked. Have you restarted first and then lost the only log line that explained why the page fired?

How the gate decides an unfreeze

When the escalation time passes and the owner has not updated the clock, I stop polishing drafts and page the secondary. The secondary should inherit a closed lock and a held draft, not a half-applied change that nobody can describe. I would rather spend the read window twice than ship an unnamed edit into the same hour as the alert. If severity is still unknown at that moment, I widen the bridge instead of letting a model guess a rank.

The unfreeze sentence

The unfreeze rule is short enough to read aloud without scrolling through a long bridge novel. A human may open the write lock only after owner, severity, and rollback intent are real values you trust. An assistant may suggest wording for the rollback line, but it may not flip the lock field by itself. If your severe-page policy demands a second person, the gate stays shut until that second name is written down.

# gate.py: proposed local check, not a production agent.
# Exit 2 while the lock is closed or required fields are placeholders.
import sys

PLACEHOLDERS = {
    '',
    'unset',
    'unknown',
    'replace-with-alert-id',
    'replace-with-human-name',
}

def load(path):
    fields = {}
    for raw in open(path, encoding='utf-8'):
        line = raw.strip()
        if not line or line.startswith('#') or ':' not in line:
            continue
        key, val = line.split(':', 1)
        fields[key.strip()] = val.split('#', 1)[0].strip().strip('"')
    return fields

def main():
    fields = load('page-clock.yaml')
    problems = []
    for key in ('page_id', 'owner', 'rollback_intent'):
        value = fields.get(key, '')
        if value in PLACEHOLDERS or value.startswith('replace-'):
            problems.append(key)
    if fields.get('severity') not in ('sev1', 'sev2', 'sev3'):
        problems.append('severity')
    if fields.get('write_lock') != 'open':
        problems.append('write_lock')
    if problems:
        print('write lock stays closed:', ', '.join(problems))
        return 2
    if fields.get('draft_state') != 'reviewed':
        print('draft still held; human review required')
        return 3
    print('human gate passed; use your normal change process')
    return 0

if __name__ == '__main__':
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

Where a free draft bench fits

Disclosure: This article was prepared as part of MonkeyCode's product outreach. I bring it up here because free model access and a free server can hold this scratch clock away from the production shell. I have no measured quota or hardware sheet, so I will not invent either figure or a permanence claim. The part I actually want is a clean split between a draft bench and the host that is paging.

Would I paste production credentials onto that free server just to save a copy step during the page? I would not, and the runbook should forbid that shortcut in the same paragraph as the product note. I would use the free model only after the read commands have already produced a snapshot I can stand behind on the bridge. It may tidy four short lines, and it may not invent the owner, the severity, or the rollback intent.

If you already have a MonkeyCode login, render this clock on the free server and carry the reviewed file back by hand. That is the only invitation I will make, because a runbook should stay useful when the product name is deleted. I still want the classifier and the gate even if you never open that draft bench at all. Does your bridge note still make sense after you remove every product word from the page?

A decision table for the first twenty minutes

I read this table before I accept help from anyone, including myself at minute ten when the alert still feels urgent. The rows are states of the clock, not moods, and the last column is the only next action I allow. If a row says the lock stays closed, I do not negotiate with a confident paragraph from a model. Have you ever talked yourself out of a table you wrote while you were still calm?

Clock state Draft state Next action I allow
File missing Any text Create the clock on scratch, run read-class commands only, and escalate if no owner can be named
Lock closed and severity unknown Held Keep observing, widen the bridge at escalation time, and do not apply a patch
Lock closed and severity named Held Let a human review the note, and keep production writes blocked
Lock open and draft reviewed Reviewed Follow the existing change process, with rollback intent still visible
Lock open and draft still held Any text Fail closed, because model text is not approval to write

A local test I can run without a fleet

I would test the gate on a laptop before anyone pastes it into a shared runbook, because an untested lock is only a comment. Create a temporary directory, copy the sample clock and gate script into it, and confirm a closed lock exits with status 2. Then write real-looking values, open the lock, mark the draft reviewed, and confirm the pass line prints on your machine. Set severity back to unknown afterward, and confirm the script refuses again, since that state is not a real unfreeze.

Exit codes worth trusting

I am not claiming this rehearsal caught a production incident, and I am not publishing a timing number from a borrowed machine. The only result I want is a predictable exit code in a directory you can delete when the check is done. If the closed lock exits zero, the gate is wrong, and it does not belong in the on-call document yet. Would you hand a pager a script that passes while the lock field still says closed?

# Proposed local test. Temporary directory only.
# Expect status 2, then status 0, then status 2 again.
tmpdir=$(mktemp -d)
cp page-clock.yaml gate.py "$tmpdir/"
cd "$tmpdir"
python3 gate.py || echo "closed lock failed closed, exit=$?"
python3 - <<'PY'
from pathlib import Path
path = Path("page-clock.yaml")
text = path.read_text(encoding="utf-8")
text = text.replace("replace-with-alert-id", "alert-1001")
text = text.replace("replace-with-human-name", "bridge-owner")
text = text.replace('severity: "unknown"', 'severity: "sev2"')
text = text.replace('write_lock: "closed"', 'write_lock: "open"')
text = text.replace('draft_state: "held"', 'draft_state: "reviewed"')
text = text.replace('rollback_intent: "unset"', 'rollback_intent: "revert the last config commit"')
path.write_text(text, encoding="utf-8")
PY
python3 gate.py
python3 - <<'PY'
from pathlib import Path
path = Path("page-clock.yaml")
text = path.read_text(encoding="utf-8")
text = text.replace('severity: "sev2"', 'severity: "unknown"')
path.write_text(text, encoding="utf-8")
PY
python3 gate.py || echo "unknown severity failed closed, exit=$?"
Enter fullscreen mode Exit fullscreen mode

Who should skip this pattern

This pattern is a poor fit when your charter already authorizes a break-glass write, such as a documented failover that does not need a fresh clock. It is also a poor fit when you are alone on a severe page and the existing runbook already names the immediate action to take. Do not place secrets, customer payloads, or production kubeconfigs on a free server because the draft step feels convenient today. Do not use the gate to delay a human page, and do not treat model confidence as a fill-in for the owner.

Free model access can change, rate-limit, or vanish, so the runbook still has to make sense with the assistant powered off completely. If your organization forbids third-party tools on an incident bridge, skip the product path and keep the clock file plus the classifier. I would rather keep a boring read window than defend a fast draft that nobody on the bridge can explain. Are you adopting a gate you cannot test on a laptop this week, or are you only collecting another note?

What I leave in the bridge note

I end the read window with four lines, and I do not let the note grow into a novel while the lock is still closed. Those lines are the page id, the owner, the severity, and the next human action after the escalation time you already named. An assistant may tidy the grammar once those lines exist, but it may not invent them from a vague alert title alone. If you cannot fill those four lines, are you really ready to run a write-class command on the host?

Top comments (0)