DEV Community

Sam Rivera
Sam Rivera

Posted on

Build a Fail-Closed Ship Card Before an AI CLI Leaves Dry-Run

I stare at a green dry-run and still hesitate to ship. One passing fixture is not a write-mode signal for me. Would you let that helper touch a shared branch tonight?

A solo CLI can look finished and still be unsafe to run. The missing piece is saved evidence, not another clever prompt. I keep a ship card beside the tool for that pause.

What the card decides

The card is a small JSON file committed next to the CLI. It names gates, proof files, and a stop rule. If any proof is missing, the process exits.

I am not asking the model to bless itself. I am asking a local script to refuse the run. Can a missing proof file count as a feature here?

This is a template you can copy tonight. It is not a benchmark from my laptop. Fill every field from a run you actually saved.

Four gates, then write mode

1. Lock the paths

Name every directory the CLI is allowed to read. Name every file the CLI is allowed to write. Anything outside those lists is out of scope.

I keep the write list painfully short on purpose. A helper that may touch the whole repo is not ready. Why hand it the whole tree this early?

2. Save the dry-run

Run the helper once with writes turned off. Save stdout, the patch, and the exit code. Put those three files under an evidence directory.

An empty evidence directory means you stop immediately. A chat log is not a patch you can review. Did you actually save the diff this time?

3. Bound the blast radius

Set a maximum changed-file count on the card. Set a maximum added-line count beside it. Refuse deletes unless the card explicitly allows them.

A surprise rename should fail this gate immediately. A secret-shaped string in the patch should fail it too. Would you merge that change while half asleep?

4. Declare the exit

Write the exact command that restores the tree. Write the condition that abandons the whole experiment. Write what you do if the runner is gone.

No exit plan means the card is still incomplete. I would rather shelve the branch than improvise. Shipping without rollback is how tiny tools linger.

Decision table

Gate Evidence you commit Fail closed when
Path lock allowed reads and writes a list is empty, or a patch path is outside
Dry-run stdout, diff.patch, exit code any file is missing or the code is not zero
Blast radius max files, max lines, delete flag caps are blank, or a delete appears
Exit plan rollback command, abandon rule either string is empty or untested

Read the row before you argue with the script. The script is the boring coworker in this workflow. Are you willing to lose an argument to it?

Steps before write mode

  1. Create a branch and a fresh evidence directory.
  2. Run the CLI in dry-run and redirect stdout to a file.
  3. Save the git diff, even when the diff is empty.
  4. Write the exit code into exit_code.txt with no extra words.
  5. Fill ship-card.json from those files, not from memory.
  6. Run the checker and accept a red result as a real gate.
  7. Consider write mode only on that same allowlist.

Skip a step and the card becomes theater. A skipped step turns the later review into mush. Which step do you usually skip under time pressure?

Copy the checker

This script is a starter, not a sandbox product. It checks presence, a zero exit, deletes, and write paths. It does not call a network, and it does not score a model.

#!/usr/bin/env python3
"""Fail-closed ship card checker. Template, not a benchmark."""
import json
import sys
from pathlib import Path

REQUIRED = [
    "allowed_reads",
    "allowed_writes",
    "evidence_dir",
    "max_changed_files",
    "max_added_lines",
    "allow_deletes",
    "rollback_cmd",
    "abandon_if",
]

def fail(msg: str) -> None:
    print(f"SHIP_CARD_FAIL: {msg}")
    raise SystemExit(2)

def main() -> None:
    card_path = Path(sys.argv[1] if len(sys.argv) > 1 else "ship-card.json")
    if not card_path.is_file():
        fail(f"missing card: {card_path}")
    card = json.loads(card_path.read_text(encoding="utf-8"))
    for key in REQUIRED:
        if key not in card or card[key] in ("", [], None):
            fail(f"empty gate: {key}")
    for key in ("max_changed_files", "max_added_lines"):
        if not isinstance(card[key], int) or card[key] < 0:
            fail(f"{key} must be a non-negative integer")
    evidence = Path(card["evidence_dir"])
    needed = ["stdout.txt", "diff.patch", "exit_code.txt"]
    for name in needed:
        file = evidence / name
        if not file.is_file() or file.stat().st_size == 0:
            fail(f"missing evidence: {file}")
    code = (evidence / "exit_code.txt").read_text(encoding="utf-8").strip()
    if code != "0":
        fail(f"dry-run exit was {code}")
    diff = (evidence / "diff.patch").read_text(encoding="utf-8")
    if "diff --git" not in diff:
        fail("diff.patch has no git diff header")
    if card["allow_deletes"] is not False and card["allow_deletes"] is not True:
        fail("allow_deletes must be a boolean")
    if not card["allow_deletes"] and "\ndeleted file mode " in f"\n{diff}":
        fail("delete found but allow_deletes is false")
    writes = card["allowed_writes"]
    for line in diff.splitlines():
        if not line.startswith("+++ b/"):
            continue
        path = line[6:]
        allowed = path in writes or any(
            path.startswith(prefix.rstrip("/") + "/") for prefix in writes
        )
        if not allowed:
            fail(f"write outside allowlist: {path}")
    print("SHIP_CARD_OK")

if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Sample card

{
  "allowed_reads": ["src/", "tests/fixtures/"],
  "allowed_writes": ["src/format.py"],
  "evidence_dir": "evidence/dry-run-01",
  "max_changed_files": 1,
  "max_added_lines": 40,
  "allow_deletes": false,
  "rollback_cmd": "git checkout -- src/format.py",
  "abandon_if": "runner missing, empty diff, or secret-like token in patch"
}
Enter fullscreen mode Exit fullscreen mode

The numeric caps must be integers, or the checker stops. It still does not count lines inside the patch. Add that counter before you trust write mode.

Commands

Make the evidence folder, then run the checker. A missing patch should end with exit code 2. Do not catch that error and keep going.

mkdir -p evidence/dry-run-01
python3 ship_card.py ship-card.json
echo $?
Enter fullscreen mode Exit fullscreen mode

Capture a real dry-run with the commands below. Adjust the CLI name to your own tool. Keep writes off until the card is green.

python3 rewrite_cli.py --dry-run src/format.py > evidence/dry-run-01/stdout.txt
printf '%s\n' "$?" > evidence/dry-run-01/exit_code.txt
git diff -- src/format.py > evidence/dry-run-01/diff.patch
Enter fullscreen mode Exit fullscreen mode

Save the CLI status before any later command runs. Otherwise you will record git, not the helper. rewrite_cli.py is only a placeholder name.

Failure fixture

Use a bad patch when you want a guaranteed red run. This one deletes a file the card should protect. Your card should reject it if deletes are forbidden.

diff --git a/src/format.py b/src/format.py
deleted file mode 100644
index 1111111..0000000
--- a/src/format.py
+++ /dev/null
@@ -1,3 +0,0 @@
-def normalize(text):
-    return text.strip()
Enter fullscreen mode Exit fullscreen mode

Point evidence_dir at that folder and rerun the checker. Confirm the printed failure and the non-zero exit. Then point the card back at the good folder.

Also search the patch before you widen writes. The starter script does not scan for secrets. A green card can still hide a leaked key.

rg -n "AKIA|sk-|BEGIN PRIVATE" evidence || echo "no obvious token"
Enter fullscreen mode Exit fullscreen mode

Add that search to your own local wrapper. I would not call the card done without it. Did the bad fixture fail loudly on your machine?

A free runner is just another field

Disclosure: This article was prepared as part of MonkeyCode's product outreach. I treat MonkeyCode free model access as one optional dry-run runner. A free server option can host that small canary off your laptop.

I will not invent a token cap, a model name, or a hardware tier. Those two availability claims are only useful after you read them today. A stale screenshot is a failed evidence gate.

Quotas, names, and duration can change without this article updating. If you use that path, add only fields you personally verified. Leave them blank when the page will not load.

Blank means abandon the write, not guess a limit. The checker never calls MonkeyCode on your behalf. You still run your own command on the machine.

{
  "model_access": "free-model-access",
  "compute": "free-server-option",
  "limits_checked_on": "YYYY-MM-DD",
  "limits_source": "account page reviewed today"
}
Enter fullscreen mode Exit fullscreen mode

The card only records that the limits were looked up. Remove the product name and the four gates still stand. Path locks do not need a vendor to matter.

If MonkeyCode is the runner you picked, open the account page first. Paste only the free-model and free-server facts you can see today. Then let the checker argue, not your memory.

Time box and rollback

Give this experiment one evening, not an open-ended week. I would stop after two red fixtures in a row. More retries usually mean the task itself is muddy.

Practice the rollback before you actually need it. Dirty status after rollback is an abandon signal. Do not continue on a tree you cannot explain.

git status --short
git checkout -- src/format.py
git status --short
Enter fullscreen mode Exit fullscreen mode

If the second status is not clean, shelve the branch. Write the reason into the abandon_if field. Tomorrow you should not have to guess why it stopped.

Who should not use this

Do not use this card on customer data, payments, or auth code. Do not use it as a security audit or a compliance pack. Do not point it at secrets or production credentials.

Skip it when your team already has a required change tool. A second checklist will rot within a week. Also skip it if you cannot store a local diff.

Free model access can be wrong, slow, or simply unavailable. A free server option can vanish during the canary. This card does not promise uptime, capacity, or a lasting free tier.

Limits you should know

The path check is plain string matching, not a real sandbox. A tricky relative path can still slip through it. Treat the allowlist as a seatbelt, not a vault.

Line caps and file caps are stored, not counted, in this version. That is a real gap, not a footnote to ignore. Add counters before write mode, or stay in dry-run.

I am not publishing timings, prices, or model comparisons here. I do not have a fresh primary measurement to cite. If you need numbers, measure them yourself and date the note.

One pass today

Copy the checker, the sample card, and the bad fixture. Run the red case first, then one dry-run of your own CLI. Stop when a required field is still empty.

What evidence file is still missing before you allow writes? That gap is the next build, not a new feature.

Top comments (0)