This entry is about the checks and balances. Nine build steps into Sentinel, five processes have grown up around the specification, and each one catches a different way the code and the documents describing it come apart. Most of the work is done by AI coding agents, but nothing in these processes depends on that. They would work the same way on a team of engineers.
They all start from one assumption: things will diverge. A specification written before the code will be wrong in places. The code will move after someone has described it. A note that was true on Tuesday will be stale by Thursday. On a project that runs for months you cannot stop that, and trying to is how a team ends up trusting documents that stopped being true a while ago. The realistic goal is to find divergence soon after it happens and to correct it somewhere the next person, or the next agent, will see it.
The rest of this entry lays out the five: what each one watches, what keeps it honest, and where it stops. First, an example of the kind of divergence they exist to catch, found eleven days ago by the newest of them.
A setting nothing read
The part of the specification that went out last Wednesday, Ingest Path and Log Schema, says this about writing to the database:
Inserts are batched - multi-row
INSERTflushed on 500 rows or 100 ms, whichever comes first.
The code had both numbers. An accumulator took batch_max_rows and batch_max_delay_ms from configuration, 500 and 100, and offered a method, due(now), that answered exactly the question the sentence asks: is this batch full, or old enough, to flush?
Nothing in the daemon ever called it.
The ingest loop passed now=0.0 into every row it added, so the batch could never age, and then took the whole batch unconditionally. Meanwhile the two settings travelled faithfully through the system: from the settings class into the dataclass, into .env.example, into the deploy workflow, into the GitHub Environment the deploy reads from. Every layer carried them, and none of them did anything. The register row that eventually recorded it put it better than I can:
The two fields travelled from
settings.pyinto the dataclass and affected nothing, which is worse than absent: a knob that appears to work.
The part I did not enjoy: due did have a caller, just not in the application. The throughput benchmark had built its own loop on top of it.
Entry one held up a row as the register earning its keep: CP-1, ingest throughput, 4,369 observations a second, with a field reading Do not pre-decide: that the margin holds. I was quite pleased with that row. The benchmark behind it was measuring a batching strategy this system has never run.
It was honest by coincidence. The unused knob said 500 rows, and the setting that actually bounds a batch, how many messages one pull from the bus may return, also says 500. When the benchmark was repointed at the real one the figure did not move, so the reading stands. But nobody knew that until somebody looked, and nothing was looking.
What found it
The tests missed it. test_batch.py passed throughout, because it tested the accumulator's flush policy directly; in the words of the row that closed it, it
asserted a flush policy nothing consulted, which is how a suite comes to certify behaviour the system does not have.
The audits missed it, because no exit criterion asks whether a configuration field is read. So did the drift check, which reads the code against the specification, and would have found a class implementing the sentence perfectly well.
It turned up when I sat down to write, from the code, a description of what the ingest path does. There is no way to write "the batch flushes when due returns true" with a citation next to it and not go and look for the caller.
That was on the day the fifth process, the development guide, was created. Writing the guide found five things on its first pass, and this was one of the gentler ones.
I have been describing these processes one at a time for three entries, and from here on I call them instruments. Here are all five together.
Five instruments, five tenses
| Instrument | Tense | Its normative home | What keeps it honest |
|---|---|---|---|
| the specification | what the system will be | itself | nothing internal; everything else comes at it from outside |
| the four registers | what is open now | the specification |
test_register_integrity.py, for the registers' shape only |
| the audits | what happened | the build step | a person, once per step, looking back |
| the seam reviews | what nobody said | the joins between documents | a rule in Introduction that every step ends with one |
| the development guide | what exists today | the code |
test_dev_docs.py, and a marker on every claim |
The specification is where it starts, and it is the only one of the five that defines anything. It is also the only one with nothing inside it checking it. A document cannot audit itself; the best it can do is be written so that the others can.
The registers number what the specification says, so that it can be cited, and hold what the build keeps producing that the specification has no home for. Entry one was about why they exist. They are under test, and entry three drew the edge of that: a register can be machine-checked for everything that is a statement about itself, and nothing that is a statement about the world outside.
The audits are the retrospective. One per build step, in two halves: what changed during implementation, and what the step did not deliver. They are kept honest by a person, and only as far as the exit criteria reach. Entry two was the first of them finding silence; entry three was one of them failing to mention a fork in the road, because there was no exit criterion for it to fail.
The seam reviews are the most recent thing the specification was amended to require, and the oldest complaint. Entry one ended on an open question, Nothing reads the documents against each other, and three weeks ago it was answered by making it a step's own work: Introduction now says that "every step ends with a review whose unit is the seam rather than the document". It is the only instrument whose unit is a pair of documents. The one-owning-document rule optimises for consistency inside a document and provides nothing at the joins, and every cross-document gap found before it existed sat between two documents that were each perfectly consistent on their own.
It has a joke in its history that I am still enjoying. The paragraph describing the drift check went on calling the seam review "standing, and currently unowned" for two days after the amendment that gave it an owner. In the document that defines the drift check. The repository's own note on it at the time:
a pointer is only as current as the last person to follow it.
What the first run actually found waits for the documents it found things in; that one is coming.
The development guide is the fifth, and it is the one that runs the other way round.
The one that points at the code
Everything else in that table is about the specification. The registers point at it. The audits measure the code against it. The seam reviews read it against itself.
Nothing said how the thing that exists behaves. Nine build steps in, the daemon was about 36,000 lines across two repositories, and every piece of work, in every session, was paying to reconstruct it. Since nearly all of that work is done by coding agents, "every piece of work" really does mean every one: an agent opening a session has read the repository and none of my recollection, so it rebuilt its understanding of the ingest loop from scratch each time and, presumably, slightly differently each time.
(It did not help that the project's own CLAUDE.md, the first file an agent reads, described two application directories as populated. Both hold a .gitkeep. They always have. They are the layering from a template for a different kind of app, and this process is a daemon. Writing the guide found that too.)
So docs/dev/ was written as fifteen chapters, around 3,300 lines, and it defines nothing, exactly as the registers define nothing. The difference is only what it points at. A register row points at the specification and is cited by id. A chapter points at the code and cites path.py::symbol. The constraint that makes that legitimate rather than a second definition was already in the specification:
Prose in these notes explains them; it does not define them once the code exists.
And where the code and the specification disagree, the guide describes the code and links the register row recording the disagreement. Where there is no row, the disagreement is a finding and gets routed like any other, instead of being written around.
Proved by, or unverified
This is the part I would steal for any project.
Every mechanism section in the guide ends in one of two ways. Either with Proved by: and the test that would fail if the claim stopped being true, or with the bare word unverified. There is no third option, and in particular there is no option for explaining why checking was not necessary. The reasoning, from the guide's own introduction:
a marker reading "not checked" invites the check; one reading "trivially true" ends it.
Two days in, sixty-five claims carried a Proved by:. Three were marked unverified.
And every chapter names the commit it was derived from,
because four pull requests means four different truths about
main.
A chapter is as current as that line says, and no more.
Then there is a test, test_dev_docs.py, which asserts that every cited path exists and defines the symbol it names, that a list marked as an enum's members holds exactly that enum's members, and that a table marked as a mapping parses to exactly that mapping.
That makes it the fifth place in the repository where a document is coupled to the code by a test, and the first that runs this way round. The other four assert that the specification still describes the code. This one asserts that the documentation still describes the code, which sounds like a small distinction until you notice that nothing checked that before.
It rotted in two days
Of course it did.
Two days after the guide was written, it was audited against the code, the way the audits check a build step. Eleven commits had touched it in the meantime, and not with the same care.
Eleven chapters claimed a currency they did not have. Every edited chapter still carried its original derived-from commit, in the layer whose own README makes that line the definition of how current a chapter is. Three described code that had changed under them. One citation credited a function with publishing messages when all it does is format a subject string; the claim underneath was true, the citation resolved, so the test passed. A citation hung on a symbol that made the sentence look checked.
That last one is the docs test's edge, in exactly the way entry three found the register test's. It can prove the symbol exists. It cannot prove the symbol is the one doing the thing.
My favourite goes straight back to the batching. The ingest chapter had to say what really decides a batch, now that the knob was gone: one pull from the bus is one batch is one transaction. It cited test_ingest_pipeline.py for that. Nothing in test_ingest_pipeline.py asserts it. So the audit demoted it, and the section now ends:
For the cadence itself - that
run_oncetakes once per fetch, unconditionally - unverified.
followed by the test that would close it.
I like that more than any Proved by: in the guide. It is the marker working in reverse: a claim that had a check pinned to it losing the check, and saying so, where the reader will trip over it. The specification described a batching rule that never ran. The guide describes the one that does, and is honest that nothing yet proves it. That is the right way round for those two documents to be wrong.
Different ways to rot
Every one of the five goes stale in its own way, including the checks themselves. The tenses are the useful way to hold all this, because each one decays differently, and so each needs something different watching it.
- The specification rots by being overtaken. It is amended as the build discharges its open items, which is why the copies published here are a frozen snapshot, and why the gap between the snapshot and the current text is what this diary is for.
- The registers rot by saying things about the world they cannot see, like entry three's sentence placing a test somewhere it had never been, which stood for three pull requests.
- The audits do not rot so much as squint. They are history, but they only see as far as the exit criteria reach.
- The seam reviews rot like any pointer: the rule was two days old before the paragraph describing it caught up.
- The guide rots by the code moving underneath it, and the derived-from commit is the only reason anyone can tell.
None of them can do another's job. The register tests cannot look outside the registers. The audits cannot see a fork that no criterion asks about. The drift check reads the code against the prose, but only from one side. The seam review reads document against document and never touches code. The guide reads code into prose, and the one thing it cannot do is say what the code should do; that stays with the specification, which is where this started.
Every one of the five was added because the one before it had an edge, and something had just fallen off it. I had assumed the newest would spend its first week finding small things in the newest code. Instead, on its first day, it found a hole under entry one's proudest number.
That is the pattern I would take to any long-running project, with or without agents. Assume the documents and the code will drift. Give each kind of drift something that looks for it, and when one of those checks finds its own edge, add the next one.
This Wednesday - The Watchdog and Derived Points. The specification's fifth part: one timer wheel for every freshness deadline, and what a derived point is allowed to say when its inputs go stale.
Next Monday - the Dev Diary. Whatever the build has sent back to the specification in the meantime; so far that has never been nothing.
Start of the diary: Four Registers and a Drift Check. Why the registers exist at all, and the one rule about where a sentence is allowed to live.
Top comments (0)