The most confident sentence in an article draft of mine this year was also wrong. The piece claimed that vendor SDK imports in the open-source project it dissected were confined to "exactly four files." It was specific, it was verifiable, and it had survived my own first pass — which is exactly the profile of error that ends up published, screenshotted, and corrected by a stranger in the comments. In this case, a separate claim-checking pass caught it before publication. The real number was five files plus a type-only one. In the very next article's publish cycle, a different check caught the same class of error again — the inventory had quietly grown to six — and a third check caught me asserting that a CI workflow's runs were "all manually dispatched" when the API showed three scheduled runs and a push-triggered one in the history.
Three errors, three catches, zero reached readers. I'm writing about them because I've come to believe the catches are the only credible evidence a fact-checking process can offer. Anyone can claim rigor. A process that documents its own near-misses is showing you the rigor instead.
This is the pipeline behind my DEV series Engineering WorldScript Studio — nine long-form technical articles dissecting one open-source repository, written with AI assistance, with every load-bearing claim verified against the code itself. It's not a product and there's nothing to install. It's a set of habits that happen to be machine-enforced, and it's entirely copyable.
1. The problem: speed compounds in both directions
AI assistance changed my drafting economics: what took a week takes an afternoon. It did the same for my error rate's potential, for three structural reasons.
First, AI-assisted drafting produces inventory claims at a scale manual writing rarely attempts — "exactly N files," "every request passes through X" — because enumeration feels like analysis and models are happy to oblige. Every one of my three near-misses was an inventory claim. Second, technical articles have a drift window: the draft is verified against commit A, but by publication day the repository is at commit B, and everything subtle has had days to rot. Third, some claims are time-dependent by nature — CI run statistics, moderation states, dashboard numbers — and go stale even while the draft sits still.
None of these are model failures. They're process failures that speed makes expensive. So the process is where the fix went.
2. The pipeline: three gates, one ledger
Gate A is a package, not a draft. Before anything goes near the publish button, each article exists as a small dossier: brief, outline, evidence file, and the piece that carries the discipline — a claim ledger. Every factual claim gets a row: the claim, its source file or API, the commit SHA it was verified against, and a status. If a claim can't get a row, it doesn't get a sentence. The fact-check report then walks the ledger row by row against a fresh checkout of the repository — not memory, not the draft's own assertions.
Gate B is staging with verification. The draft goes to the platform unpublished, and the stored body is read back and compared byte-for-byte with the local source. (Why that paranoia is justified: I once pushed a "tiny" post-publish fix using a cached article body and briefly unpublished two live articles in one move — the cached front matter still said published: false. The fix propagated a rule: never write from stale state; always fetch the body fresh from the server. It's in the ledger as D-29, and it will outlive every tool I currently use.)
Gate C is the publish-day drift guard, and it's my favorite gate. Before the publish flip: re-resolve the repository's HEAD, byte-diff every evidence file against the verified baseline, and — the hard-won part — re-run inventory greps without path filters. My second near-miss lives here: the draft's vendor-SDK inventory was built from a grep that, unknown to its author, had a blind spot (a hooks directory nobody thought to include). A repo-wide re-grep on publish day caught a sixth file the staging grep had missed. The rule now reads like a superstition and works like a charm: no inventory claim survives contact with a path-filtered grep.
Time-dependent claims get the same treatment: the CI workflow statistics in one recent article were re-pulled from the API on publication day and compared against the staged numbers — unchanged, so the staged fact-check carried; had they moved, the draft would have been re-dated, not silently stale.
3. Two instances, on purpose
The structural decision that caught near-miss number one: the writer and the claim-checker are different sessions. My drafting session produces the article and its ledger; an independent session — different context, no shared assumptions, instructions to be adversarial — receives the load-bearing claims and checks them against the sources. It flagged "exactly four files" by re-running the enumeration itself and finding a fifth.
This matters because self-review has a known failure mode: you proofread your beliefs, not your text. A second instance has no beliefs to protect. Notably, near-misses two and three were caught by the first instance — the publish-day drift guard and a routine API re-check — because the process had internalized what the second instance taught it. That's the learning curve you want to see: external checks becoming internal habits, with the external check still there anyway.
4. The quieter disciplines
A few rules that don't show in the gate diagram but do the daily work:
- Released truth only. One planned article is blocked right now because its subject is merged but unreleased — the series' rule is that claims describe what a reader can download, not what main contains. Blocked beats impressive-but-wrong.
- A retired-claims register. Every claim that ever died goes on a list with its obituary: what was asserted, why it's false, what the verified replacement says. Absolute claims ("every AI action passes through one service") are effectively banned unless they carry an inventory appendix, a grep protocol, and a SHA.
- As-of dating. Anything that moves — run counts, moderation states — gets an explicit "as of" date in the text. Dated honesty ages; undated claims just become wrong.
- Recommendation is not roadmap. When an article proposes an improvement, it says so in those words. Describing unshipped work as shipped is the one sin the whole apparatus exists to prevent.
- Disclosure, visibly. Every article carries the platform's AI-assistance badge and a footer saying the same in words. The badge isn't a confession; it's part of the provenance, like the SHA.
5. The numbers, such as they are
Across articles #2–#9, the fact-check tables hold 109 numbered checks, each tied to a file, an API response, or a measured value at a named commit. Three self-caught errors pre-publication (the two inventory misses, one data-staleness catch). One operational incident (the accidental unpublish) converted into a standing rule. Zero reader-flagged corrections after publication — so far, and I'd write "so far" even if the number were larger, because the process assumes its next failure is already scheduled.
I'm aware of the irony that this article makes factual claims about a fact-checking process. It went through the same gates: the numbers above come from counting rows in the actual ledgers, and its own claim ledger exists. The process doesn't get a day off for writing about itself.
6. What to steal
None of this needs my tooling. It needs four decisions:
- Ledger every claim. If it can't cite a source at a specific state, it doesn't ship. The ledger is boring to maintain and priceless at 11pm before publication.
- Verify against fresh state. Repositories drift, APIs change, caches lie. Check out, re-pull, re-run — on publication day, not drafting day.
- Forbid path filters in inventory checks. "Every file that…" claims demand repo-wide enumeration, every time. Your blind spots live exactly where you didn't look.
- Separate the writer from the checker. Another person, another session, another model — anything without your assumptions. Give it the claims, not the conclusions.
The comfortable story about AI-assisted writing is that the model gets things right. The durable one is that the process catches things getting wrong — and writes the catch down. My three near-misses are, in the end, the best articles I never published.
Written with AI assistance — as is everything in the series this describes, which is rather the point. All numbers refer to the author's own publication ledgers as of 2026-09-28; the series discussed is "Engineering WorldScript Studio."
Top comments (0)