In April I wrote about a system I built on Claude Code that takes an epic to PR-ready code across 80+ microservices. If you read that post, I sounded like someone who had finished something.
I hadn't. This is the follow-up I didn't plan to write. It covers what went wrong, what I got wrong, and how a tool I built for myself ended up being used by the whole tech team, every new joiner, and people who don't write code at all.
I had written "DO NOT" eight times in one file
A couple of months in, I stopped trusting my own rules.
The one that stung most was commits. My prompts said, in capitals, that the system must never run git commit. It is the rule I care about most, because I want to read every diff before it becomes history. One day I went looking for why that rule felt shaky, and I found two things. My own settings file was explicitly allowing git commands. And one of my own reference documents contained a ready-made git add && git commit snippet. I was telling the model NEVER in eighteen places while handing it permission, and an example, in two others.
Until then, my fix for anything that went wrong had been the same every time. I opened a prompt file and added another line in capital letters. MANDATORY. NEVER. A warning emoji, to be safe.
So one weekend I sat down and audited everything I had built. The same set of rules (don't commit, don't lint, don't skip a phase) was pasted into 18 different files. One of those files said "DO NOT run" eight times. The main orchestrator prompt was 1,464 lines long, and all of it was loaded at once.
I felt a bit silly. This is exactly what I flag in other people's pull requests: duplicated logic and no single source of truth. I would never have approved it. And I had written it myself, one frustrated line at a time.
The damage was what you would predict from any codebase with copy-pasted rules:
- Drift. I changed a rule in one file and 17 others quietly disagreed with it. The model had to guess which copy was the real one.
- Dilution. When everything is MANDATORY, nothing is. The model could not tell a rule that protects a loan calculation from a rule about where the braces go.
- Cost. Every turn, it re-read the same prohibitions when it should have been reading my code.
Asking is not the same as preventing
The fix had two parts.
The first was boring: one rulebook. A single file where every rule has a short ID, like GIT-1 for "never commit" and ISO-1 for "only write inside the isolated workspace". Every other file now says "obey GIT-1 and ISO-1" and nothing more. My main instruction file went from 370 lines to about 145.
The second part changed how I think about these tools.
I split the rules into two kinds. Some are judgment calls, like matching the code style already in a service. The model applies those with discretion, and that is fine. Others are invariants. Never commit. Never write outside the workspace. Those have to hold every single time.
For the invariants, I stopped asking. Claude Code has a permission system, and I moved those rules into it. The command is now denied before it runs. The model's opinion about it no longer matters.
A prompt is a request. A deny rule is a wall. I had spent months making the same request in a louder and louder voice, when what I needed was a wall.
It is not a perfect wall. A command wrapped inside another command can still slip past the matcher, so I kept the written rule as well and I check the permission file from time to time. But the difference in day-to-day behaviour was immediate.
It learned to debug, which I never planned
The April version did one thing: turn an epic into code.
But new features are not most of my week. A large part of it is support. Something failed in production, a number is wrong, and somebody needs to know why before the end of the day.
So I gave the system read-only agents. One searches the logs. One queries MongoDB. One runs SQL. One reads metrics. None of them can write anything.
The first version of the log agent taught me something. I would describe a problem, and it would search for what it imagined the log message looked like. It found nothing, and it reported that nothing with complete confidence.
The fix is embarrassing in its simplicity: before searching the logs, search the code for the actual log string. That is what I do myself when I debug. I had just never written it down, because it never occurred to me that it was a step.
Every incident we solve now becomes a runbook. There are 29 of them. The next time the same thing breaks, the system starts from the runbook instead of from zero, and so does whoever is on support that week.
The day a teammate opened a pull request
In May, a colleague opened a pull request against the framework. It was a skill that spins up a local cluster of our services for debugging. I had not asked for it.
A few weeks later someone else added the database and log agents I just described. Three teammates have contributed so far.
This mattered more to me than any token number. In April I wrote that this was "just something the team needed". If I am honest, at the time it was something I needed, and I was hoping the team would too. A pull request from someone else was the first real evidence.
I remember reading that diff twice. It was not complicated. I read it twice because someone had gone through my prompt files closely enough to extend them, and had decided the thing was worth their own evening. Until then I had been the only person who knew where anything was.
Then people I never built it for started using it
Today the whole tech team works out of this repo. I did not announce that or push for it. It spread one person at a time, usually after someone watched a colleague get an answer in two minutes that would have taken them an afternoon.
New joiners start here now. Anyone who has joined a company with 80+ services knows the first month: you don't know which service owns what, and you feel like you are interrupting a senior engineer every time you ask. Now a new joiner asks the system first. It answers from the service map and from the code itself, and it never sighs.
For any new joining, the first question used to be some version of "where do I even start reading?". I never had a good answer. Ninety-odd services is not something you read. Now I tell them to open the repo and ask it how a loan moves from application to disbursal, then follow whatever it points at. They come back with better questions, and they come back sooner.
The part that surprised me most is the product team. They do not write code, and I never imagined them opening this repo. But they use it to ask plain questions, like which third-party vendors we depend on and which flows each one sits behind. That used to be a message to an engineer and a wait for whoever had time. Now they ask directly and get an answer grounded in the actual code.
I built a pipeline for delivering code. The thing people reach for most is simpler: a way to ask the system how it works and get a straight answer.
Smaller changes that added up
Reviewers. There are now review agents that read a plan or a diff before I do, including one that checks specifically for fintech risk and can block the work. That block used to be a sentence in a prompt. Now it is a piece of state the workflow checks, so it cannot be talked around. One rule I settled on: the reviewer must be at least as capable a model as the author.
The right model for the job. Fetching logs does not need the most expensive model. Deep code analysis and test-driven implementation do. Each agent now runs on the model that fits its work.
A second codebase. The system now works across two separate sets of repositories, where before it knew only one.
What I would tell myself in April
- If you are writing the same rule for the third time, you have a structure problem. The model is not the issue.
- Decide which of your rules are invariants, and enforce those outside the prompt.
- Build for the work you actually do. For me that was debugging, even though code generation makes the better demo.
- Share it before it feels ready. The best parts of this system were not written by me.
- Watch who actually uses it. I built this for engineers shipping epics. New joiners and the product team get as much out of it, and they use it for something I treated as a side feature.
The numbers
I do not have a tidy before-and-after benchmark, and I would rather not invent one. Here is what I can count:
- 92 services are mapped today, up from the 80+ I wrote about in April.
- 31 features have their own workspace in the framework.
- 29 runbooks have come out of real production incidents.
- 3 teammates have contributed code to it, and the whole tech team uses it.
- The token figure from the first post still holds: roughly 75,000 down to 15,000 per epic, because the heavy reading happens in separate agents.
How it all fits together
If you skipped everything above, this is the whole system in one picture. Three kinds of request go in: a question, an epic, or an incident. All three read from the same shared knowledge and run inside the same guardrails, and what each one produces goes back in for the next person.
Some things have not changed. I still approve every phase. I still read every diff. I still raise every pull request myself. I do not expect that to change, and I do not want it to.
The interesting part was never removing the engineer from the loop. It was reducing how much context the engineer has to hold in their head.
If you are building something similar, I would like to know one thing: which of your rules turned out to be walls, and which were only ever requests?






Top comments (0)