Generated troubleshooting pages stay trustworthy only when a human-owned failure contract pins every status, retry class, and compensation rule before prose exists. A model may draft explanations, examples, and sequencing language that cite those pins, but it must not invent operational facts. A small checker then rejects any draft sentence that asserts a code, a wait, or a side effect absent from the pin file. Teams that skip the pin will publish fluent pages that disagree with the tests that actually define the API.
The split that keeps drafts from owning behavior
Public API docs fail in a narrow way: the narrative sounds specific while the numbers and side effects are unverified. Readers follow a retry hint, a status table, or a compensation step, then hit production behavior that the page never extracted from tests. That gap is a documentation ownership problem, not a wording problem, because fluent prose cannot repair a missing contract. The workflow below treats the contract as the only mergeable source for operational claims in the published page.
A useful page still needs readable recovery language, sample requests, and a short sequence of what the caller should try next. Those parts are draftable once each sentence points at a pin identifier that a human already accepted. The model is a drafter of exposition, and the human remains the owner of behavior, scope, and publication. This split is stricter than a style review, because style review never checks whether a status code exists.
Claim classes in one decision table
| Claim class | Example in a page | Drafter | Owner | Merge rule |
|---|---|---|---|---|
| Transport status | HTTP 409 on duplicate submit | Extract from tests | Human | Prose must quote the pin row |
| Retry class | safe to retry with the same key | Human label only | Human | Draft may not upgrade the class |
| Wait or backoff | two seconds, three attempts | Human if stated | Human | Omit unless the pin records it |
| Compensation | void the hold, then refresh | Human sequence | Human | Draft may explain, not reorder |
| Caller explanation | why the cart looks unchanged | Model | Human edits tone | Allowed if it cites a pin id |
| Sample payload | redacted request body | Model from fixture | Human confirms redaction | Fixture path must be in the pin |
| Escalation severity | page the on-call tier | Nobody but the human | Human | Forbidden in generated drafts |
| Support promise | response within an hour | Nobody but the human | Human | Forbidden in generated drafts |
The table is the editorial policy for this workflow, and it should live next to the pin file in the same review. Rows marked Forbidden stay out of generated drafts, and a later edit may not soften them into hints. If a fact is missing from the pin, the page must stay silent rather than guess a helpful number. Reviewers should treat a missing cell as an omission, not as permission for the drafter to fill a plausible value.
What the model may draft
The model may draft opening context, the plain meaning of a pin row, and a short example copied from a named fixture. It may also draft transition sentences that connect two pin identifiers without adding a new status, a new wait, or a new side effect. It may propose headings, a glossary sentence, and a short what-you-will-see paragraph when every concrete token appears in the pin. It may not draft escalation policy, billing outcomes, legal remedies, security workarounds, or any promise about support response time.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. Free model access is enough for this drafting pass, because the task is bounded exposition over a pin the human already wrote. Use the free server option only as a review surface after you confirm what that option actually provides. Do not assume retention, isolation, quotas, or hardware from this article, and confirm those limits with the operator before you depend on them.
What a human must own
- Extract failure rows only from the tests or handlers that define the behavior, and reject rows that survive only in old wiki pages.
- Assign a stable pin identifier, a retry class, and an explicit compensation list, including an empty list when no compensation exists.
- Record fixture paths for every sample the page may show, and strip secrets before that fixture becomes eligible for drafting.
- Review the drafted essay only after the checker passes, then sign the pin hash so later prose edits cannot drift silently.
Human ownership here means the person can defend each row against the test suite, not merely approve the tone of the paragraph. If nobody can point to the asserting test, the row stays out of the pin and out of the page. That rule slows the first page and prevents a second page from laundering the same guess. The signature belongs on the pin, not on the model draft, because the draft is replaceable and the contract is not.
Build the pin file
The pin format below is a proposal you can commit beside the docs tree, and it is not a measured production schema. It stays small on purpose, so a reviewer can diff behavior without reading the surrounding essay. Each row names the asserting test, which keeps the contract tied to executable evidence rather than to memory. Empty compensation arrays are valid and should stay explicit, because silence is not the same as an empty list.
{
"pin_version": 1,
"surface": "checkout.submit",
"rows": [
{
"id": "FC-409-DUP",
"status": 409,
"retry_class": "same_key_safe",
"compensation": [],
"fixture": "fixtures/dup_submit.json",
"asserted_by": "tests/checkout_test.py::test_duplicate_submit",
"caller_may_see": "duplicate submit"
},
{
"id": "FC-422-FIELD",
"status": 422,
"retry_class": "fix_then_new_attempt",
"compensation": ["drop the invalid field"],
"fixture": "fixtures/bad_field.json",
"asserted_by": "tests/checkout_test.py::test_unknown_field",
"caller_may_see": "unknown field"
}
]
}
Check drafts before they can merge
The checker is an unexecuted example, not a measured production gate, and you should run it in your own repository before you trust it. It loads the pin, collects allowed tokens, and fails when the markdown invents a status code or a forbidden phrase. It also requires every cited pin identifier to exist, and it requires every status in the page to be listed. Extend the forbidden list for your own support promises rather than weakening the defaults to quiet a noisy draft.
#!/usr/bin/env python3
"""Unexecuted example: reject troubleshooting drafts that outrun the pin."""
import json
import re
import sys
from pathlib import Path
FORBIDDEN = (
"page the on-call",
"within an hour",
"guaranteed",
"we will refund",
"severity 1",
)
def main() -> int:
pin = json.loads(Path(sys.argv[1]).read_text())
draft = Path(sys.argv[2]).read_text()
rows = pin["rows"]
allowed_status = {str(row["status"]) for row in rows}
allowed_ids = {row["id"] for row in rows}
errors: list[str] = []
for code in set(re.findall(r"\b[1-5]\d{2}\b", draft)):
if code not in allowed_status:
errors.append(f"status {code} is not in the pin")
for pin_id in set(re.findall(r"\bFC-\d{3}-[A-Z0-9-]+\b", draft)):
if pin_id not in allowed_ids:
errors.append(f"unknown pin id {pin_id}")
lowered = draft.lower()
for phrase in FORBIDDEN:
if phrase in lowered:
errors.append(f"forbidden phrase: {phrase}")
mentioned = set(re.findall(r"\bFC-\d{3}-[A-Z0-9-]+\b", draft))
if not mentioned:
errors.append("draft cites no pin id")
if errors:
print("\n".join(errors))
return 1
print(f"ok: {len(rows)} pin rows, {len(mentioned)} ids cited")
return 0
if __name__ == "__main__":
raise SystemExit(main())
python3 check_failure_pin.py failure_contract.json docs/checkout-failures.md
The commands above are the local gate, and they do not require a hosted model at all. Run them in continuous integration on every docs pull request that touches a troubleshooting page in the repository. If the checker fails, fix the pin or delete the sentence, and do not ask the model to invent a matching row. A green run only means the draft stayed inside the pin, not that the pin matches production behavior today.
Numbered drafting procedure
- Freeze the pin and record its hash before any model call, so the draft cannot quietly edit the contract underneath the review.
- Send the model only the pin rows, the fixture excerpts, and a written instruction to cite pin identifiers and omit unknown numbers.
- Write the draft to a branch, run the checker locally, and run it again in continuous integration before a human reads for tone.
- If a free server option is available, use it only for the drafting review, and merge only after the pin hash is unchanged.
Step two is where free model access fits, because the input is a closed set of rows rather than an open invention request. Keep the prompt narrow: explain the listed rows, quote status values only from those rows, and leave escalation blank. If the model adds a backoff interval, treat that as a failed draft even when the sentence sounds cautious. The human adds waits only by editing the pin first, then regenerating the explanation from the updated rows.
Review notes that belong in the pull request
A reviewer should see four artifacts in one change: the pin, the asserting tests, the draft, and the checker output. The review comment should name the pin hash and the ids the page cites, rather than saying that the page looks complete. Completeness is not a property of prose length, and a short page citing two real rows is safer than a long silent page. If the tests changed since the pin was frozen, the pin update is a separate commit with its own owner.
Compare the draft against the decision table before you approve wording, and allow explanation to move only when the cited fixture is unchanged. Retry class, compensation order, and status codes cannot move unless the pin diff shows the same change in the same review. That comparison takes minutes when the table lives in the repository, and much longer when the policy lives only in chat. Keep the table versioned with the pin so a future reviewer can see which policy governed the merged page.
Limitations you should keep visible
This checker catches status-shaped numbers and a short forbidden list, and it will miss paraphrases that smuggle a new side effect without a code. A sentence such as "the hold is released automatically" can pass the regex while still inventing compensation the pin left empty. A human reading pass remains mandatory, and the checker is a floor rather than a proof of behavioral accuracy. You should add project-specific patterns for currency amounts, region names, and retry counts if those tokens appear in your pages.
The pin is only as true as the tests it names, so a stale test will stamp a stale contract with false confidence. Free model output can still be wrong inside the allowed vocabulary, for example by swapping two pin explanations while citing both identifiers. A review on the free server option does not validate behavior, execute the suite, or replace the local checker in continuous integration. Nothing in this workflow measures latency, cost, or quality scores, and you should not invent those numbers to justify the gate.
Who should not use this approach
Skip this workflow when the failure behavior lives only in production logs and nobody can point to an asserting test or handler. Skip it for security advisories, legal terms, billing promises, and live incident notices that need a named owner. Skip it if the team will treat a green checker as permission to publish without reading the compensation column. Also skip it when the API surface changes hourly and the pin cannot be frozen for the length of one review.
Small libraries with one obvious error type may not need the full table, and a single owned example in the README can be enough. Large platforms with several client languages should still keep one pin per surface, then let language pages cite it instead of restating codes. Duplicating the contract into every SDK guide recreates the drift this gate is meant to stop. If two pages must differ, the difference belongs in the pin as a client note, not in untracked prose.
Keep the essay replaceable
The practical conclusion is narrow: freeze behavior in a pin, and block merge when drafted sentences outrun that file. Free model access fits the drafting step, and the free server option is only a review surface, but neither one owns the status map. The checker and the human signature are the parts you should keep even if you later swap the drafting tool. If asserting tests already exist, start with two pin rows, then try one free model pass only after the hash is frozen.
Top comments (0)