After my document-ingestion pipeline died three nights in a row, I stopped trusting "it should work" and rebuilt the ingestion layer around one assumption: it will die.
It has since survived 5 session kills and 19 watchdog restarts across a 127-conversation corpus (~8.2M characters, ~11,400 chunks) — zero conversations lost, zero duplicates.
What actually mattered, in order
- Idempotent writes (upserts) — re-runs never duplicate, so crashes cost nothing
- Checkpoint written BEFORE the work it describes — the sidecar resumes exactly where it stopped
- Each vector-store write in a subprocess — a wedge poisons one conversation, never the whole run
- Per-chunk logging — a watchdog that kills "quiet" processes kills working ones; the log must tell the truth
- A watchdog with a staleness timeout calibrated to the logging rate — one known coupling, documented
What did NOT matter
The embedding model, the vector DB choice, the framework. All swappable. The resilience architecture is the actual product.
The stack
Fully local: Ollama + nomic-embed-text + ChromaDB on an 8GB MacBook Air, no cloud, no API keys, no accounts. Airplane-mode safe.
The suite is three tools, AGPL, published: Scribe (ingest engine + sidecar checkpoint), Warden (watchdog), Hand (subprocess-isolated writes)
The failure story, with receipts
The full design rationale — including the night a three-day wedge was diagnosed as a logging failure, not a systems failure, and my own records had to be corrected for counting restarts instead of completions — is published: The Cache Is Not the Corpus
A record that hides its own errors is publicity, not documentation.
Next in this series: what the watchdog-vs-heartbeat design teaches about building AI systems that survive their operators. Follow if that's your lane.
Top comments (0)