Originally published at https://cristiandeluxe.dev/blog/running-a-company-on-an-agent-fleet/.
I use concurrent coding-agent sessions to work on my companies' software. I built the shared infrastructure those sessions use, including a private knowledge wiki and a pipeline for adding sources to it. I also operate the memory service behind that work.
The useful details are in the failures. In August 2026, I was dealing with a growing local memory store, competing writers and model fallbacks that could undermine review. These are the checks and measurements from that period.
Give each session a place to start
My context layer is a cross-linked Markdown wiki, versioned in Git. It holds project setup and past decisions, with links back to the sources. A new session can look up why something was done before proposing to change it.
I built an ingest pipeline for captured web content. A bulk model classifies the captures; a stronger model synthesizes the selected material into wiki pages. Validation rejects pages with no incoming links and writes outside the allowed paths. The workflow also requires a human diff review before publication.
On August 7, I captured a batch of 103 open research tabs. The log records 94 threads reaching classification: 50 selected for ingestion and 44 rejected. I kept the raw captures alongside the synthesis so I could check a page against what I had actually read.
Keep the author and verifier separate
The synthesis workflow uses a second model to grade pages against their sources. I enforce author/verifier separation in code: the grader must not resolve to the model that wrote the page. That includes fallback selection. Otherwise a routing failure can quietly turn an independent check into the author reviewing its own output.
I route large classification passes to a model tier with enough quota to finish them. Synthesis and grading use the scarcer tier. The selectors reject unrecognized model names, so a typo stops the job instead of silently changing its routing.
That separation gives me a check to inspect. It does not make the generated pages correct by itself; I still need the source links and the diff.
Measure the memory service under load
At the time, the local semantic store contained roughly 700,000 entries mined from repositories and session transcripts. I ran a fork of the memory engine because several operational fixes needed to live in the code.
One project-mine measurement on August 4 took 5.3 seconds against an empty store and 190 seconds against a store with 653,000 entries, for the same 448-item project. That was roughly 36 times slower. The mining path checked each file with a collection-wide query. I replaced those repeated lookups with a prefetch scan for larger projects, keeping a final check under the write lock.
A separate curation run showed why adding writers was the wrong response to a queue. Serial processing produced about 7.5 entries per minute. Parallel reads with batched writes reached about 33 entries per minute, a 4.4-fold increase. The writes still went through one writer. The gain came from batching work through that writer, not from letting multiple processes update the vector index.
The store had already suffered index divergence twice. I kept the single-writer lock and treated it as a constraint when changing throughput.
Make recovery stop where it should
I added a watchdog that checks the daemon and index every 15 minutes. It can restart a wedged daemon, but it abstains when a command-line mining job holds the writer lock. The recovery path depends on what failed; a running process alone is not enough to call the service healthy.
Concurrent sessions also share Git repositories. I require each session to stage only its own paths and push completed commits promptly. A formatting hook catches another source of avoidable conflicts. Before restoring a file, the session has to check whether another session changed it.
Those rules let me keep several sessions working without treating another session's unfinished work as something to clean up. When a check fails, I want the job to stop with enough evidence to tell me which operation failed and what it had already written.
Top comments (0)