I keep one log line in a note on my phone. Least dramatic entry I own. Also the most useful thing in there.
[02:03:12] sync start target=test-cluster
Here is what sat behind it. A six-hour sync job pushed a dev box into a working directory, and the script carried rsync --delete. The exclude list had gone stale. A refactor moved a pile of private modules to a new path and nobody amended the file. rsync ran the arithmetic it always runs — file sitting on the target, no matching entry in the exclude list, delete — and dozens of entries disappeared in one sweep. Exit code 0. A clean bill of health from the exact process that had just finished deleting things.
Recovery took a bit over two hours. Laptop copies, an old branch, a tarball a colleague had mailed himself before a conference. Most of it came back. Some of it didn't come back clean.
Nothing in that pipeline crashed. Every component did what it was told to do. That's the part I keep chewing on, and it took a while to name, because nothing was actually broken.
I spend my days on multi-model agent systems now, and a jury has the same silhouette as that sync script: a fault-tolerant decision circuit made of parts that all report success. If I can't trace each juror, I can't tell you what the verdict is worth. So I instrument the thing like a crime scene, and the habit traces straight back to that exclude file.
When a decision hits the jury it fans out into N model calls. Each call becomes a span. I want juror_id, model_id, model_family, prompt_hash, temperature, seed, token_in, token_out, latency, the raw verdict text, the self-reported confidence, and the calibration bucket that confidence lands in. The parent span carries decision_id. A separate adjudication span records quorum, the individual votes, the final verdict, and the escalation reason if one fired. That's the chain of custody. Without it you have a verdict and a shrug. Somebody asks which juror was lying, and the honest answer is: no idea.
Voting rules are where self-deception gets formal. Simple majority for categorical verdicts. Quorum q, two of three on a normal day. No quorum, escalate to an adjudicator or the on-call human. Weighted voting uses w_i derived from holdout calibration — ignore what the model claims about itself. Self-reported confidence is a witness statement: useful, unsworn, frequently rehearsed. Track the Brier score per juror instead. I watched a model declare 0.93 confidence and then miss the same class of prompt four calls running. The confidence field was theater. The Brier score was the rap sheet.
Expected cost per decision:
C = Σ_j (c_in * tok_in_j + c_out * tok_out_j) + P_esc * C_esc + P_err * C_err
P_err is a guess. Pull it from a labeled canary set, or from whatever the last incident taught you. Treat it as ground truth and you will have a bad quarter. Two numbers matter day to day: cost per accepted decision, and cost per caught error. Re-estimate both every time the jury changes. Models get swapped, prompts drift, providers quietly change things underneath you with no changelog. A control strip saved me from shipping a bad configuration once — one cheap juror and the full jury over the same traffic sample, side by side. The expensive jury was paying for agreement. Accuracy wasn't in the receipt. I'd rather see that bill on my own dashboard than in a finance review.
Tiering keeps the bill honest. Tier 0: one fast model, accept if its confidence clears the cutoff and the prompt isn't sitting in a known dissensus cluster. Tier 1: a second model from a different family, accept if the two agree. Tier 2: a third diverse model or an adjudicator, which buys coverage where error cost is high. N-of-N stays reserved for high-risk intents. I've watched teams route every decision through five models because it felt safer. It felt safe right up until the latency tail crossed the user-facing timeout and the invoice landed. Cross-validation is a tool. Don't make it a personality.
Diversity is the part everyone nods at and nobody measures. Correlated errors kill you quietly. Two prompts to the same model family can agree because they share blind spots, and on a dashboard agreement looks like consensus. Log model_family and provider on every span. Compute pairwise error correlation on the canary set. High correlation means the jury is decorative. Bring in a different architecture, a different retrieval path, or a tool trace that can contradict the model outright. I once watched three jurors drawn from two families agree on a wrong answer. Same blind spot, dressed up twice. That's a chorus. You wanted a jury.
When a decision goes sideways I reach for: dissensus rate by intent, agreement matrix by model pair, escalation rate with reason codes, cost per decision against error rate, calibration drift per juror. Forensics first, dashboard later. Confessions are cheap. Spans hold up. If every juror agreed and the outcome was still wrong, the problem lives in the shared blind spot — different investigation, different fixes.
The trade-offs don't go away. Latency tails grow with N, especially when the calls run sequentially. Parallel calls cut latency and raise peak rate. Weighted voting adds machinery, and stale weights can be gamed. Early exit saves money and weakens the cross-validation you just paid for. Three jurors catch a lot of single-model errors. Shared blind spots survive them intact. More jurors will not fix a bad prompt. Start with a 2-of-3 jury where the error cost clearly beats the jury cost. Instrument first. Tune quorum later. The verdict is a cost-bounded approximation. Trace it.
Back to the sync job. That script was a jury of one. No independent witness anywhere in the chain. The exclude file was the prompt, the target was the output, and everybody agreed. The agreement deleted dozens of private modules.
We chipped at it in the gaps between normal work for a while before it stopped being able to surprise us. The guard came in passes. None of them landed clean the first time.
First pass: make it print rsync --delete --stats and alert on deletions. Observability with no brake. The script exited 0, the alert landed in a channel, and people read it after the damage. A witness that can't stop the crime is a diary.
Second pass: hash the exclude file, block whenever the hash changed. That made noise. The exclude file changes for legitimate reasons, and engineers started appending --force to get their own work done. A guard everyone bypasses is a speed bump with a story attached.
What finally held was the dry-run gate. rsync --dry-run --itemize-changes --delete runs first. The wrapper parses the output for *deleting. If any file currently on disk is about to be removed, it exits 97, writes blocked into the run marker, and leaves the target untouched. The real rsync --delete only runs after the dry run comes back clean. The dry-run output gets logged with run metadata in the same shape I use for decision_id: intent, target, exclude file hash, who kicked it off. Now the run has a trace, the deletion has evidence, and the blocked marker carries a reason code you can grep. The dry run costs a few seconds on a tree this size. Cheap, compared to the alternative.
We did try blocking every deletion for a while. That broke legitimate cleanup, so we compromised: deletions allowed only when the path sits inside the managed set and the run carries an explicit --allow-deletions flag. Semi-automatic, deliberately annoying, and it has held up against our own sloppiness so far. It has not been tested against an exclude file that is actively hostile. I'd rather keep it that way.
The lesson that survived: most of the nodes believed they were right. The sync believed the exclude file. The exclude file believed the last refactor. The target accepted whatever arrived. No single node was broken, and there was no independent juror anywhere in the chain.
So before the next destructive sync, run the dry run first. On the actual target. With the actual exclude file.
rsync -ani --delete --exclude-from=/path/to/exclude /src/ /dst/ | grep '^\*deleting'
If that prints a path you don't recognize, stop. The dry run is the evidence. The real run can wait. Build the model jury to the same standard — every juror traced, same spans, same reason codes, same refusal to accept agreement as proof. Disagreement is the evidence. The verdict is a cost-bounded approximation. Price it. Trace it. Keep the line from 02:03:12 somewhere you can still read it.
Top comments (0)