DEV Community

Avery Lin
Avery Lin

Posted on

Fail the Docs Build When CLI Examples Drift From a Fixture Ledger

Generated command documentation fails in review when example lines are invented rather than recorded from a controlled run. A fixture ledger should freeze the argv list, the exit status, and a hash of stdout before any model writes surrounding prose. The drafting model may explain flags, suggest section order, and propose cautionary notes, but it must not own the recorded commands. Reviewers then accept or reject the page by comparing fenced examples with the ledger, rather than by trusting the model's confident wording.

Why help examples drift after a clean draft

Help text and short tutorial snippets often outlive the exact binary build that originally produced them. A renamed flag or a changed default can leave a polished paragraph pointing at a command that now exits nonzero. Hashing the captured output makes silent format drift visible without requiring the prose to quote every byte. The same ledger also tells a later editor which examples are safe to paraphrase and which lines are frozen.

Manual screenshots and pasted terminal blocks do not, by themselves, stop this kind of example drift. A paste can be edited during drafting, and a screenshot cannot be diffed when a flag moves between releases. A small JSON ledger is easier to review in pull requests because each field has a single owner. The prose file then becomes a view over that ledger, not a second source of truth.

What the model may draft, and what a human must own

Ownership of each field should be decided before the first generated paragraph is opened for editing. Human owners keep every field that a user might copy into a shell or cite in an incident note. The drafting model may write connective prose only after those owned fields already exist as committed data. If a sentence would change user behavior when it is wrong, it stays outside the drafting pass entirely.

Field Human owner Model role
Argv list Yes, from the fixture run Must not invent or reorder arguments
Exit code Yes, from the fixture run Must not guess success or failure
Stdout and stderr hashes Yes, from the fixture run Must not describe unseen bytes as facts
Flag names and defaults Yes, from source or a machine help dump May mention a flag only if the inventory lists it
Section order and flag rationale Reviews the draft May draft the wording
Canonical URL and anchor id Yes May suggest visible link text only
Support windows and warranty lines Yes Must not draft those sentences

That table is a review contract for the docs pull request, not a claim about any particular generator. A missing human owner for exit codes is enough reason to block the documentation pull request. A missing owner for introductory tone is not a blocking defect, because tone does not change the command result. The distinction keeps the drafting pass narrow, reviewable, and clearly separate from the recorded command facts.

A numbered workflow you can rerun

1. Dump a flag inventory from the binary

Start from the command itself, or from the parser definition that the release binary actually ships. Write one JSON object per flag, including the canonical spelling, the default, and whether the flag is stable. Do not ask a model to reconstruct that inventory from memory or from an old README. Commit the dump beside the docs so a later release can diff added, removed, and renamed flags.

2. Specify fixtures before any prose exists

Each fixture names an id, an argv list, optional stdin, and the doc files allowed to cite it. Keep the first fixture set small, covering version output, a dry-run lint, and one deliberate failure. Reject any fixture that needs production credentials, live tenant identifiers, or unbounded network calls to complete. The spec is human-owned input, and the model should not be allowed to append commands to it.

3. Record the ledger on a clean runner

Run the recorder in a disposable environment so local shell aliases and developer config files cannot leak into the examples. A disposable server or a locked CI job is enough for this recording pass when the binary and fixtures fit that environment. Confirm those limits in the provider's own console before you depend on them for a release. Store exit codes and hashes, not a private stdout body, unless the example is already public.

4. Draft prose only from the ledger and the inventory

Give the drafting model the ledger, the flag inventory, and a short style note as its only inputs. Ask it to explain when to use each recorded command, and to leave argv lines untouched. Require every fenced example to carry the fixture id, the argv list, the exit code, and the stdout hash. If the model adds a command that is not in the ledger, discard that block instead of editing it into truth.

5. Fail the docs check when cited fields drift

The checker loads the ledger, scans documentation for fixture fences, and compares argv, exit code, and stdout hash. A field mismatch fails the pull request even when the surrounding paragraph reads smoothly and completely. Unknown fixture ids fail as well, because a new example must be recorded before it is explained. Prose-only edits pass this mechanical gate and can continue through the ordinary editorial review that follows.

6. Review the human-owned sentences by hand

After the checker passes, a human still reads rationale, warnings, and any sentence that states a guarantee. The checker cannot tell whether a caution note overclaims safety or omits a required privilege for the command. That reading stays on the docs owner, not on the model and not on the hash comparison. Merge the page only when both the ledger gate and that human reading are fully complete.

Reference checker and ledger shape

The following Python is an unexecuted reference implementation, not a measured result from a production repository. It shows the recording step and the comparison step in one file so a team can adapt the paths. Command timeouts, argv allowlists, and secret scanning belong in the wrapper that calls this reference script. Do not point it at commands whose stdout may contain tokens, customer data, or unredacted paths.

#!/usr/bin/env python3
"""Unexecuted reference: record CLI fixtures and check doc fences."""
import hashlib
import json
import re
import subprocess
import sys
from pathlib import Path

FENCE = re.compile(r"```

fixture\n(.*?)

```", re.S)

def record(spec: dict, timeout: int = 30) -> dict:
    proc = subprocess.run(
        spec["argv"],
        input=spec.get("stdin"),
        text=True,
        capture_output=True,
        timeout=timeout,
        check=False,
    )
    stdout = proc.stdout.encode()
    stderr = proc.stderr.encode()
    return {
        "id": spec["id"],
        "argv": spec["argv"],
        "exit_code": proc.returncode,
        "stdout_sha256": hashlib.sha256(stdout).hexdigest(),
        "stderr_sha256": hashlib.sha256(stderr).hexdigest(),
        "stdout_bytes": len(stdout),
    }

def check(ledger_path: Path, docs: list[Path]) -> list[str]:
    ledger = json.loads(ledger_path.read_text())
    by_id = {row["id"]: row for row in ledger["fixtures"]}
    errors = []
    if not ledger.get("binary_version"):
        errors.append("ledger: binary_version is missing")
    for path in docs:
        for block in FENCE.findall(path.read_text()):
            claimed = json.loads(block)
            recorded = by_id.get(claimed["id"])
            if recorded is None:
                errors.append(f"{path}: unknown fixture {claimed['id']}")
                continue
            for key in ("argv", "exit_code", "stdout_sha256"):
                if claimed.get(key) != recorded.get(key):
                    errors.append(f"{path}: {claimed['id']} {key} drifted")
    return errors

def main() -> int:
    if len(sys.argv) < 3:
        print("usage: fixture_gate.py record spec.json ledger.json", file=sys.stderr)
        print("       fixture_gate.py check ledger.json docs.md", file=sys.stderr)
        return 2
    if sys.argv[1] == "record":
        spec = json.loads(Path(sys.argv[2]).read_text())
        rows = [record(item) for item in spec["fixtures"]]
        payload = {
            "binary_version": spec.get("binary_version"),
            "fixtures": rows,
        }
        Path(sys.argv[3]).write_text(json.dumps(payload, indent=2) + "\n")
        return 0
    errors = check(Path(sys.argv[2]), [Path(p) for p in sys.argv[3:]])
    print("\n".join(errors))
    return 1 if errors else 0

if __name__ == "__main__":
    raise SystemExit(main())
Enter fullscreen mode Exit fullscreen mode

A spec file can stay small while the team proves that the gate fails closed on a bad example. The ids below are placeholders for the workflow, not observed runs, and the exit codes are not measurements. Replace the argv entries with the binary you actually ship before anyone records a ledger from this sample. Record that ledger only after the replacement, then copy recorded fields into documentation fences without retyping them.

{
  "binary_version": "REPLACE_FROM_BUILD",
  "fixtures": [
    {"id": "version", "argv": ["tool", "version"], "stdin": null},
    {"id": "lint-dry-run", "argv": ["tool", "lint", "--dry-run", "sample/"], "stdin": null},
    {"id": "lint-missing-path", "argv": ["tool", "lint"], "stdin": null}
  ]
}
Enter fullscreen mode Exit fullscreen mode
python3 fixture_gate.py record fixtures/spec.json fixtures/ledger.json
python3 fixture_gate.py check fixtures/ledger.json docs/cli/lint.md
Enter fullscreen mode Exit fullscreen mode

A documentation fence then cites the recorded row instead of a loose terminal paste from a chat transcript. The hash is a merge gate for editors, not a value that readers are expected to type into a shell. Editors can still show a short redacted excerpt beside the fence when that excerpt comes from the same recorded run. If the excerpt and the hash disagree, the excerpt is discarded and the ledger row remains the cited result.

{
  "id": "lint-missing-path",
  "argv": ["tool", "lint"],
  "exit_code": null,
  "stdout_sha256": "REPLACE_FROM_LEDGER"
}
Enter fullscreen mode Exit fullscreen mode

Where a free model and a free server fit

MonkeyCode's free model access and free server option can sit inside this workflow without becoming the source of truth. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The free server is a reasonable place to record fixtures when a clean process environment is required and current plan limits are confirmed. The free model is a reasonable drafter for rationale paragraphs once the ledger and the flag inventory are the only behavioral inputs.

Those two availability claims are not a capacity plan, a hardware specification, or a promise that a named model remains on the free tier. This article does not report quotas, durations, benchmark scores, or supported binary sizes, because those facts were not measured here. Treat the provider console and the current product terms as the primary sources before a release depends on either option. If the server cannot run the binary, record the ledger in the existing CI runner and keep the same checker in the repository.

The drafting prompt should stay narrow so the model cannot widen the set of recorded commands on its own. Pass the ledger JSON, the flag inventory, and an instruction to refuse any command that is not already recorded. Ask for explanations, failure meaning, and related reading, while forbidding new flags, new exit codes, and new URLs. Save the model output as a proposal file, then run the checker before a human edits tone or order.

Limitations of the ledger gate

Hashes do not prove that an example is wise, only that the cited run still matches the recorded ledger. Nondeterministic output, timestamps, and unordered maps will fail the gate unless a normalizer strips them before hashing. A command that passes today can still be a poor teaching example if it requires privileges the page never states. The human review in the last workflow step exists because those judgments are not hash comparisons at all.

The checker also trusts whatever environment recorded the ledger, which may not match the release under review. A fixture recorded on a developer laptop can embed aliases, extra path entries, or a locally patched binary. That risk is why the recording step prefers a clean server or a locked CI image with a noted binary version. If the version field is absent, reviewers should treat the hashes as incomplete evidence rather than as a release proof.

Secret handling needs a separate control from the documentation checker that the workflow above already described. Hashing stdout does not make it safe to execute a command that prints credentials, because the process still ran. Keep fixture argv free of tokens, and block any fixture that reads production configuration or private environment files. If a team cannot meet that bar, this workflow is the wrong place to generate public help pages from live runs.

Who should not use this approach

Skip the ledger gate when public output is inherently unstable and cannot be normalized without hiding the lesson. Interactive installers, full-screen pagers, and commands that wait on a live service fit this method poorly. Also skip it when legal or support commitments dominate the page, because those lines need specialist owners. A small internal tool with two stable subcommands is a better first candidate than a platform CLI with remote side effects.

Teams that cannot store a binary version beside the ledger should wait before adopting the gate. Without that version, a green check only says the docs match some earlier machine, not the release under review. The method is also a poor fit for marketing pages that intentionally compress commands into slogans or screenshots. Use the ledger for reference pages where a copied command must exit as the documentation claims it will.

Close the review with one failing example

Start with one subcommand, three fixtures, and a pull request that fails on a deliberately wrong exit code. That single failure shows reviewers what the gate catches before anyone drafts a full help chapter around it. After that failure is understood, allow the model to draft rationale from the passing ledger and review human-owned sentences by hand. If a free recording environment is available, use it for the clean run, and keep the checker in the repository for the next release.

Top comments (0)