I use AI coding agents every day for real client work. A few months ago I noticed a pattern, and once you see it you can't unsee it:
The agent says "done". I push. Tests fail. Or lint breaks. Or the same wrong assumption comes back next session, because nothing was ever written down anywhere.
Each instance was easy to fix. But the pattern kept costing me. And the pattern wasn't a model problem — it was a verification problem. Nobody was checking the work before it left the machine.
The dumb fix that worked
I stopped trying to make the agent smarter and built a gate instead. The gate is deliberately dumb — no intelligence in it at all. It's a Node CLI called forgekit (I built it, it's MIT, zero runtime dependencies):
1. forge verify — make "done" mean something checkable
It detects the repo's test suites, runs them, and reports the truth: PASS, FAIL, INCOMPLETE, or PARTIAL. The key decision: a run that only covered part of the repo is reported as partial — never as green. Most of my "done but broken" failures died right here.
2. forge precommit — check before it lands
Looks at what's actually staged and scans for secrets before anything pushes. The unglamorous stuff that saves you exactly once, and that once is worth it.
3. forge impact — "if I touch this, what breaks?"
A blast-radius estimate before an edit. Honest caveat, and I put it in the README too: it's a heuristic code graph, not a sound call graph. It misses files as well as over-warns. I treat its output as advisory. It's the part I trust least — real repos are the only way to harden a heuristic.
4. forge ledger / forge remember — memory with receipts
Lessons and decisions live in the repo as claims, each carrying its evidence. Only tests, CI, or a human raise a claim's confidence. So the agent stops re-learning the same thing every session — and when it "remembers" something, you can see why it believes it.
The workflow rule
No push without forge verify + forge precommit green. It's a rule, not a suggestion. The repeated-failure class that motivated the whole project is gone from my workflow.
What I deliberately didn't build
A sandbox — guardrails reduce risk, they're not a security boundary. And anything claiming formal verification: "proof-carrying memory" is the feature's name, not a theorem. When you're selling reliability, naming things honestly is part of the product.
Try it
npm install -g @codewithjuber/forgekit → forge init. One init emits native config for Claude Code, Codex, Cursor, Gemini, Aider, Copilot, Windsurf, Zed, Continue and OpenClaw. (Claude Code is what I use daily, so that's the deepest-tested integration.)
Repo: https://github.com/CodeWithJuber/forgekit
If you try it, tell me where the impact analysis misses. That's the backlog I'm building against.



Top comments (1)
Completely agree with your philosophy here.