For the past year my consulting assignment didn't allow AI agents. So I coded with them in my free time instead: evenings, weekends, a lot of hours. Not to see whether agents can write code (they can), but to see whether they could write code I would actually want to keep.
During my thirty years as a dev I've come to care about clean architecture and code quality more than is probably healthy. That turned out to be the whole problem.
What the usual setup gets you
I did what everyone does. A good model, an AGENTS.md with my rules, a set of skills for the things I do often. It worked, in the sense that features appeared. But I kept running into the same three walls.
The agent guessed. A feature description written for a person leaves gaps a person fills from context. The agent filled them with something plausible. The code compiled, its tests passed, and it did the wrong thing.
The rules were requests. "Read a file before you edit it." "Stay inside the workspace." "Don't add that dependency." All of it lived in a prompt, and a prompt is a suggestion. Most of the time the model followed it. Most of the time is not a rule.
I was the only check. Every change ended with me reading a diff, plus a confident summary written by the same agent that wrote the code. The agent was grading its own homework, and I was the bottleneck.
More instructions didn't fix any of this. Each new line in AGENTS.md made the agent slightly better and the file slightly longer. I needed the guarantees to come from somewhere other than the model.
First, measurements the agent couldn't argue with
So I started building my own MCP servers to measure code quality, code duplication and more, so the agent got facts about the code instead of its own impression of it. They're all on GitHub.
Then I built my own agent
It became KiwiPow Agent, a VS Code extension. I kept iterating on it, and every iteration moved a rule out of the prompt and into the code around it.
- Plan from intent, not from code. The planner reads the docs and approved specs, never the code, and writes a spec of named rules. A planner that reads the code inherits its bugs as requirements.
- Disagreements become decisions. A separate step reads the code and turns every place it disagrees with the spec into a decision. The agent doesn't quietly pick a side; I rule.
- Done comes with evidence. Every task names the test that proves each rule it delivers. No test, not done.
- Read before write, enforced. A session cannot write a file it has not read, or one that changed since it last looked. Not discouraged: refused.
The results got much better.
Five products in a year
With my MCP servers and later my own agent, in a year, on my own, I shipped five products. Not because I needed five products, but to test the approach across different industries and different
ways of working. They are all online if you search ;)
- Ottoweb, homepages for small companies,
- the Kiwisonic suite: a DAW, a synth, a drum designer, a bass designer and a vocal transformer,
- Bikermike, road-trip planning for motorcyclists,
- Sanndir, an AI coach, perfect for venting co-worker troubles ;),
- Bankbook, accounting.
One person. A team of agents. Every decision about what got built, approved and accepted was still mine.
This isn't my first time chasing a neglected loop
My first company, Coderr AB, came from the same itch. I kept seeing teams treat errors as a chore, something to clear rather than something to learn from. So Coderr caught errors and collected telemetry automatically, and ranked them by what mattered, such as how many customers each one hit.
Agent-written code has the same problem one step earlier. The loop between what was asked, what was built and what was accepted is the part nobody has time for, so it doesn't happen. The agent writes faster than anyone can check it against the requirement.
From an agent to KiwiPow
An agent works while I sit in a session with it. What the agent taught me is that the important work shouldn't wait for that: requirements and code should keep getting refined in the background, and the people who own the requirement should be part of the loop, not just the developer. Today the
neglected loop runs between what the business asked for, what the agent built, and what someone actually accepted.
So KiwiPow isn't the agent. It's two products built on what the agent taught me.
KiwiPow Server turns an ask into a spec planned from your own documents, with a citation on every behaviour and open questions that block approval. The product owner approves a numbered revision. Approved features wait in a ranked queue: a developer takes the next one, no sprint planning needed.
KiwiPow IDE is the editor where the agent builds exactly the approved behaviours and proves each one with a test. The rules are enforced by the IDE itself, the same on every model, and every tool call is recorded before it is answered. What the agent knows about the code comes from analyzers built on the language's own toolchain, stored per commit, so it doesn't get to invent a call graph. Where the code disagrees with the spec, the decision travels back to the person who owns the requirement. The tester accepts behaviour by behaviour, against the tests.
And more of it runs on its own. Background loops keep refining: the planner critiques specs before anyone reads them, the analysis turns code-health findings into explained work, and what the code teaches flows back into the requirement.
One license instead of a tracker, an editor and a code-quality tool. And KiwiPow is built with KiwiPow: one founder, a team of agents, every decision owned by a person. Same rules for me as for you.
Where it is now
It's not released yet. If you code with agents and want them held to what was actually asked, sign up for the release. Early signups lock the founding price for as long as they stay customers:
https://kiwipow.com/signup.html
I'd also love to hear how you deal with this today. What's in your AGENTS.md or skills that the model still ignores?
Top comments (0)