Configuration documentation should merge only after every sentence is classified as an extracted fact, a human recommendation, or narration that makes no operational promise. A model may explain keys that a schema extract already lists, and it may restate defaults that the same extract records. A human must own recommended values, secret handling, environment overrides, and any sentence that implies support. This split keeps generated guidance useful without letting drafted prose invent a contract the repository cannot prove.
Why generated config pages drift
Most configuration pages fail quietly because reviewers judge fluency while the evidence grade of each sentence stays invisible. The page names a real key, then slides from the shipped default into a preferred production value without a second source. Operators treat that slide as a supported promise, and a later incident can inherit a setting nobody agreed to own. The useful fix is a frozen schema extract, a separate recommendation file, and a checker that rejects mixed claims before merge.
Three evidence grades on one page
An extracted fact is a key name, type, or default that appears in the current schema extract and nowhere else. A recommendation is an owner-approved value for a named environment, with a review date and an explicit non-goal. Narration may define a term, point to a file path, or explain why two recorded facts sit near each other. Narration may not introduce a number, a port, a timeout, or a security outcome that neither input file contains.
| Claim in the draft | Required evidence | Writer | Merge rule |
|---|---|---|---|
| Key exists, with type | Schema extract row | Model may restate | Fail if the path is absent |
| Shipped default | Extract default field, copied exactly | Model may restate | Fail on any mismatch |
| Recommended value | Human entry with owner, environment, and review date | Human writes the entry; model may quote it | Fail if the quote or owner fields are missing |
| Secret handling | Human file names the key and the store, never the value | Human only | Fail if a literal secret appears |
| Support or compliance sentence | Out of scope for this page | Nobody in this workflow | Fail closed |
| Navigation and definitions | Extract descriptions plus file paths | Model may draft | Allowed only without new numbers |
The table above is the merge policy for this workflow, and it is intentionally smaller than a full style guide. Rows describe claim types rather than page sections, because a single paragraph often mixes more than one grade. A sentence that fails its row should be deleted or moved into the human file before anyone rewrites tone. This policy is a proposal for review gates, and it does not describe the behavior of any particular documentation platform.
What the model may draft
The model may draft a definition sentence for each key present in the extract, using the type and description fields already stored there. It may restate a default only by copying the extracted value, including units and nullability, without rounding or substituting a common port. It may write a short navigation sentence that points to the recommendation file when a key has an owner-approved override. It may not choose that override, estimate a safe timeout, or describe compliance effects that the inputs do not state.
What a human must own
A human owner must record every recommended value, the environment it applies to, and the date when that recommendation expires or needs review. The same owner must state which keys are secrets, which files must never be committed, and which defaults are unsafe outside local development. Support boundaries belong in that file too, including which combinations the team will debug and which combinations are explicitly unsupported. Generated prose may quote those recorded lines, but it may not soften, extend, or generalize them into a broader promise.
Workflow
1. Freeze a schema extract from a real input
Start from a schema file that the repository already treats as input to validation, not from an old README table. The extract should keep key path, type, default, and a content hash so later prose can be tied to one revision. If no schema exists, build the extract from a checked-in fixture that a test already loads, and label that fixture as the source. Do not ask a model to infer missing defaults from memory, because that inference is exactly the claim this gate exists to block.
python3 tools/extract_config_keys.py \
--schema config/app.schema.json \
--out build/schema-extract.json
#!/usr/bin/env python3
"""Build a flat key extract from a small JSON Schema document.
Labeled example: allOf, $ref, and remote schema URLs are not handled here.
"""
import argparse
import hashlib
import json
from pathlib import Path
def render_default(spec):
if 'default' not in spec:
return None
if spec['default'] is None:
return 'null'
return str(spec['default'])
def walk(node, prefix, out):
props = node.get('properties') or {}
for name, spec in props.items():
path = prefix + '.' + name if prefix else name
if isinstance(spec, dict) and 'properties' in spec:
walk(spec, path, out)
continue
out.append({
'path': path,
'type': spec.get('type', 'unknown') if isinstance(spec, dict) else 'unknown',
'default': render_default(spec if isinstance(spec, dict) else {}),
'description': spec.get('description', '') if isinstance(spec, dict) else '',
})
def main():
parser = argparse.ArgumentParser()
parser.add_argument('--schema', required=True)
parser.add_argument('--out', required=True)
args = parser.parse_args()
raw = Path(args.schema).read_bytes()
keys = []
walk(json.loads(raw), '', keys)
payload = {
'source': args.schema,
'content_hash': 'sha256:' + hashlib.sha256(raw).hexdigest(),
'keys': keys,
}
Path(args.out).write_text(json.dumps(payload, indent=2) + '\n')
if __name__ == '__main__':
main()
The walker above is a proposed local tool, and it only reads the schema file you pass on the command line. It records a hash of that file so the documentation job can show which revision the prose describes. Nested objects become dotted paths, which keeps the later checks simple without inventing values that the schema never stored. Schema combinators, remote references, and network URLs stay out of scope for this first gate on purpose.
2. Record recommendations in a human-owned file
The recommendation file is short on purpose, and a human edits it in the same pull request that changes supported settings. Each entry names an owner, an environment, a review date, a value, and a reason that points to an issue or decision record. Absence is meaningful: if a key has no entry, the doc may describe the extracted default and must not recommend a replacement. Secret values never appear here; the file may only name the key and say that the value comes from a secret store.
{
"recommendations": [
{
"path": "http.timeout_ms",
"value": "5000",
"environment": "staging",
"owner": "ops-docs",
"review_date": "2026-12-01",
"reason": "placeholder decision id, replace with a real record",
"non_goal": "not a production support promise"
}
]
}
Treat the JSON above as a fixture shape, not as a setting any running service uses. The review date should be a real calendar date your owner accepts, and the reason should point to a record your team can open. The non-goal line is there so a later quote cannot be read as a broader warranty. If you cannot fill owner and review date, leave the key out rather than asking a model to invent them.
3. Draft narration only from those two inputs
Give the model only the schema extract, the recommendation file, and a sentence budget for each key. Ask it to label each paragraph as extracted, recommended, or narration, and to omit any key missing from the extract. Run that draft on a host you already trust, then discard any sentence that cites a file the prompt did not include. Save the raw draft beside the extract hash so a reviewer can see which inputs the narration was allowed to use.
Require three surface patterns in the draft so the checker can stay small, deterministic, and easy to explain. State a shipped default only with the fixed phrase for extracted defaults, followed by the exact value in backticks. State an owned override only with the fixed phrase for recorded recommendations, followed by the human value in backticks. Any other sentence may name backticked key paths, but it must not add digits, support language, or recommendation verbs.
4. Run the lane checker before review
The checker reads the extract, the recommendation file, and the markdown draft, then fails the build on four conditions. A mentioned key path must exist in the extract, and a stated default must match the extracted value exactly. A sentence that uses recommendation language must quote a value present in the human file for that key. A sentence that mentions a secret value, a support promise, or an unlisted number must fail even if the surrounding prose sounds cautious.
python3 tools/check_config_doc.py \
--extract build/schema-extract.json \
--recs owners/recommendations.json \
--doc docs/configuration.md
#!/usr/bin/env python3
"""Fail config docs that mix extracted defaults with unowned recommendations.
Proposed local gate. It does not call a model and does not open a network connection.
"""
import argparse
import json
import re
import sys
from pathlib import Path
KEY_RE = re.compile(r'`([A-Za-z0-9_.-]+)`')
DEFAULT_RE = re.compile(r'extracted default is `([^`]+)`')
REC_VALUE_RE = re.compile(r'recorded recommendation is `([^`]+)`')
REC_WORD_RE = re.compile(
r'\b(recommend(?:ed|ation)?|should set|prefer(?:red)?|production value)\b',
re.I,
)
SUPPORT_RE = re.compile(
r'\b(we support|supported combination|service level|guaranteed)\b',
re.I,
)
SECRET_LITERAL_RE = re.compile(
r'\b(api[_-]?key|password|secret|token)\b.{0,40}`[^`]{6,}`',
re.I,
)
NUMBER_RE = re.compile(r'\d')
def load(path):
return json.loads(Path(path).read_text())
def check(doc, extract, recs):
errors = []
by_path = {item['path']: item for item in extract['keys']}
rec_by_path = {item['path']: item for item in recs['recommendations']}
known = set(by_path) | set(rec_by_path)
for lineno, line in enumerate(doc.splitlines(), 1):
mentioned = KEY_RE.findall(line)
for key in mentioned:
if '.' in key and key not in known:
errors.append(f'{lineno}: `{key}` is not in the schema extract')
for shown in DEFAULT_RE.findall(line):
keys = [key for key in mentioned if key in by_path]
if not keys:
errors.append(f'{lineno}: extracted default without a known key')
for key in keys:
expected = by_path[key]['default']
if shown != expected:
errors.append(
f'{lineno}: default `{shown}` for `{key}` != extract `{expected}`'
)
recommends = bool(REC_WORD_RE.search(line) or REC_VALUE_RE.search(line))
if recommends:
keys = [key for key in mentioned if key in known]
if not keys:
errors.append(f'{lineno}: recommendation language without a known key')
for key in keys:
rec = rec_by_path.get(key)
if rec is None:
errors.append(f'{lineno}: `{key}` has no human recommendation entry')
continue
quoted = REC_VALUE_RE.findall(line)
recorded = str(rec['value'])
if recorded not in quoted:
errors.append(
f'{lineno}: `{key}` must quote recorded value `{recorded}`'
)
for field in ('owner', 'environment', 'review_date'):
if not rec.get(field):
errors.append(f'{lineno}: human entry `{key}` missing {field}')
if SUPPORT_RE.search(line):
errors.append(f'{lineno}: support language is outside this config page')
if SECRET_LITERAL_RE.search(line):
errors.append(f'{lineno}: possible secret literal; name the store, not the value')
allowed_numbers = set(DEFAULT_RE.findall(line)) | set(REC_VALUE_RE.findall(line))
residue = line
for value in allowed_numbers:
residue = residue.replace('`' + value + '`', '')
if NUMBER_RE.search(residue):
errors.append(
f'{lineno}: number is neither an extracted default nor a recorded recommendation'
)
return errors
def main():
parser = argparse.ArgumentParser()
parser.add_argument('--extract', required=True)
parser.add_argument('--recs', required=True)
parser.add_argument('--doc', required=True)
args = parser.parse_args()
errors = check(Path(args.doc).read_text(), load(args.extract), load(args.recs))
if errors:
print('\n'.join(errors))
return 1
print('config doc claims match extract and human recommendations')
return 0
if __name__ == '__main__':
sys.exit(main())
5. Review the human file before the prose
Review time should land on the recommendation file and the extract hash, not on whether the generated sentences sound friendly. If the hash changed, regenerate the draft from the new extract instead of patching numbers by hand inside the markdown. If the owner, environment, or review date is missing, the page is not ready, even when every key name is correct. This order keeps the model in a drafting role and keeps the support boundary in a file a person can diff.
Expected results on the fixtures
The fixtures below are illustrative inputs for the method, and they are not measurements from a production service. You can place them next to the script and compare the exit status with the two cases described here. The failing page is written to report a default mismatch, an unowned recommendation, and a key missing from the extract. The passing page is written to print the success line and exit with status zero, because every number is an exact copy.
Minimal schema fixture, config/app.schema.json:
{
"type": "object",
"properties": {
"http": {
"type": "object",
"properties": {
"port": {
"type": "integer",
"default": 8080,
"description": "Listen port for the local process."
},
"timeout_ms": {
"type": "integer",
"default": 2000,
"description": "Request timeout in milliseconds."
}
}
},
"log": {
"type": "object",
"properties": {
"level": {
"type": "string",
"default": "info",
"description": "Log verbosity name."
}
}
}
}
}
Failing draft, written so the checker should reject it:
`http.port` is the listen port. The extracted default is `80`.
Prefer a production value `443` for `http.port`.
`cache.ttl_seconds` should set the value to `30`.
The script is written to emit a default mismatch for http.port, a missing human entry for that same key, and an unknown-key error for cache.ttl_seconds. It is also written to flag the extra digits 443 and 30, because those numbers are neither extracted defaults nor recorded recommendations. Run the command on your machine before you trust the gate, and treat a different message as a bug in the local copy. These lines are specified behavior of the example, not a benchmark from a hosted documentation run.
Passing draft, which stays inside the two files:
`http.port` is the listen port. The extracted default is `8080`.
`http.timeout_ms` has a staging override. The recorded recommendation is `5000`.
`log.level` controls verbosity. The extracted default is `info`.
How to read a failed check
Read the line number first, then decide whether the fix belongs in the schema, the human file, or the draft. A missing key means the prose invented a path, so delete the sentence or add the key through the normal schema change. A default mismatch means the draft drifted, so copy the extract value again instead of editing the schema to match the prose. An unowned recommendation means a person must either record the value with owner fields or remove the advice entirely.
Limitations of a phrase gate
This checker is a text gate, so it misses recommendations that avoid words like recommended, should, and prefer. A synonym pass will still miss implied advice, which means a human read of narration remains necessary on every change. The extract can also drift from the running binary if your build does not regenerate it in the same job that validates docs. Do not use this workflow for security advisories, legal terms, or incident updates, which need a different ownership record.
Who should not start with generated guidance
Skip this approach when the repository has no schema, no fixture, and no person willing to sign the recommendation file. Skip pages that must state a service level, compliance status, or a customer-specific exception outside config narration. Teams that want the model to pick production values will find the gate frustrating, and that frustration is the intended result. A docs-only repository with unchecked examples should first capture an extract, then adopt the checker, rather than generating a full guide on day one.
Put the gate on the merge path
Add the checker to the same job that already builds docs, and fail the job when the exit status is non-zero. Keep the recommendation file beside a named code owner so generated markdown cannot update supported values unnoticed. Regenerate the draft whenever the extract hash changes, and do not patch numeric claims by hand in the markdown. The page is ready only when the extract hash, the recommendation review date, and the checker result match the commit you intend to merge.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
If both files already exist, MonkeyCode's free model access is a fitting place to draft narration from those files alone. The free server option is a reasonable place to run this checker before review when Python 3 is already available. Neither option chooses recommended values, and neither option replaces the human owner named in the recommendation file.
Top comments (0)