I run a personal agent system for my own job search, built on adopted frameworks I configured and hardened rather than wrote from scratch (Hermes for reasoning, OpenClaw for execution, Hiring OS as the single source of truth for job data). For a while I had a nagging feeling that I was building more infrastructure into it than I needed. Feelings like that are useless on their own, so I stopped trusting it and went and counted instead.
Here's what was actually sitting on disk:
- 32 config backup files across two systems, spanning 67 days, under four different naming conventions running in parallel. Three of them, labeled "STABLE" and cut 40 minutes apart on the same day, are byte-identical when diffed.
- A duplicate test harness born from a 2-line diff. A test script needed one change the next day: a different output path, one added environment variable. What it got instead was a full second copy of the entire harness, same eighteen output files, same support scripts. Diffed against the original, the two scripts differ by two lines out of about twenty.
- Unused scaffolding for a multi-agent workboard: a goal/card model, a verb-routing dispatcher, a stub safety net, a receipt step, built for a system that has run exactly one demo goal and zero live ones, with one of its two core verbs never implemented.
- 34 archived creative-tooling skills, bulk-moved out of the live system on a single day, one of them alone documenting 38 shaders and 24 character palettes, none with any recorded reason for being set aside.
- Three separate fine-tuning runs of the same idea in one day: a resume-screening model, trained at three different scales, one of which measurably overfit partway through. None of the resulting model files ever got referenced in any live config.
- Three parallel copies of the same career database across three directories, one of them two months stale.
None of that is dramatic by itself. A backup file is not a bug. A second training run is not a crime. What the list shares is a shape: build first, notice the duplication later, and usually fix the pointer without deleting the older copy. That lag, repeated enough times, is the actual cost.
The audit that mattered more than the inventory
Counting artifacts tells you something happened. It doesn't tell you what it cost. For that I had to compare output against consumption, not activity against intention, because activity always looks productive from the inside.
I pulled up one full working session and checked what came out of it: model upgrades, a pass through roughly twenty unrelated GitHub repos, a tool installed and never touched again, an unrelated hardware video watched somewhere in the middle. Zero job-search output. In that same window, two job-discovery automations, the actual point of the system, had been sitting disabled for three weeks. Unnoticed, while I kept building things next to them instead of running them.
Separately: 692 jobs scored, 105 sitting untouched in a save queue with the oldest six weeks old, my last real application six days prior, outreach drafts that had never once left a "simulated" status. Building more automation on top of that picture would have made the backlog worse, not better. The bottleneck was never scouting volume.
Why "it looks useful" isn't a good enough bar
Once the numbers were visible, the mechanism behind them was ordinary, not some deep personal flaw: building has fast feedback, a clean stopping point, and zero risk of rejection. Job search, networking, interview practice have none of those properties. They're slow, ambiguous, and personally exposed. Given a choice between a task with guaranteed positive feedback and one without it, the easier task wins by default unless something structural stops it. That includes work done with an AI coding assistant. The loop doesn't care what's generating the code; it only cares that finishing something feels good and starting the hard thing doesn't.
So I wrote a gate directly into the system instead of relying on noticing it again later. The orchestrating agent's persona document now carries a hard stop condition: once the core pipeline works end-to-end, don't add new frameworks, bots, models, or frontends until the current setup has run reliably in normal use. A separate frozen-state document states the same rule at the architecture level: don't add infrastructure without a real usage failure that proves it necessary.
That's an inversion of the old default, not a tweak to it. Under the old rule, "a memory layer might help" was reason enough to install one. Under the new rule, "might help" doesn't clear the bar; something has to have already broken in a way only new infrastructure fixes.
A worked example of the gate, before the gate existed
Four days before I wrote that rule down, I'd already been applying it by instinct on one specific dependency: a memory tool called GBrain. It installed cleanly and looked functional. Writes worked; gbrain remember succeeded and gbrain doctor independently confirmed the fact was stored. Recall didn't: asking for a fact I'd just written, by exact string, by paraphrase, by provenance tag, returned "No matching facts," every time.
Rather than assume the bug was in my integration, I isolated it. A 9-step script first exercised a different, already-working memory tool end to end (store, recall, correct, recall the correction, all four steps clean), then hit GBrain's own CLI directly with no dispatch layer in between. Same failure. I upgraded GBrain a full point release and reran with a brand-new fact, specifically to rule out a stale-record false pass. Same result. At that point the evidence pointed at GBrain's own recall engine, not my wiring, so I stopped: removed the connection cleanly, left the local install in place in case a future version fixes it, and moved on. No bug filed, no fork, no further debugging on a dependency I don't maintain.
That same week two other things got cut for unrelated reasons: a messaging channel whose "disabled" flag didn't actually stop outbound sends until three escalating fixes later, and two agent personas an audit found had zero real production dependency. Three kills, three different root causes, four days, none of them individually dramatic. Together, they're the gate's own worked example, in reverse: each of those things existed without ever meeting a "proven necessary" bar, and each got cut once that became visible.
What I'm not claiming
The freeze document is one day old as I write this. Nothing confirms the rule has held past today, and I'm not going to write an ending where the discipline "worked": that's not a claim the evidence supports yet. What exists is a gate, written into the system instead of left as intention, and a number I can check later: whether the next thing I'm tempted to build actually clears "a real usage failure proved this necessary," or whether it's just the easiest room in the house, again.

Top comments (0)