What I wanted was not a system that only works when it is written down. What I wanted was a team that ships the same quality. While using AI, company-wide.
Someone comes up with a good practice. Someone sets a standard. Someone builds discipline. And then it stays with that one person. A common story. But if it stays there, there is no point in developing as a team.
One person builds a good way of working. That way reaches every project. It reaches every member. Only then does it become the team's value. In the age of AI, I believe this "spread" is what value looks like.
This article is about how we turned that spread into a mechanism.
What I wanted: a state where everyone ships the same quality
Why not documents?
Documents are kept only by the people who read them. People who do not read them do not keep them. A style guide sitting on a shelf moves nobody's hands. What I wanted was not more documents. It was the same quality, reproduced in everyone's hands.
Left alone, good practices become personal property. One person's project is careful; the neighboring project is sloppy. Same company, uneven output. One side writes tests; the other does not. One side keeps evidence; the other runs on gut. This is a loss for the team. The effort of the one person who built a good way stops with that one person.
For a long time, the assumption was: write it down and it will be kept. But documents alone are not kept. They are kept only when they are built into the working steps. So I decided not only to write, but to distribute. To every project, we distribute the same spec process. The same quality standards. The same operating discipline. The distribution channel is Claude Code. We built it as an in-house plugin platform, used only by our company. A plugin is a component you attach to a tool afterward to add capability.
"Distribute" may sound like pushing from above. It is the opposite. What gets pushed is rules nobody can keep. What gets distributed is a way of working that someone already used on the ground, and that worked. Not an ideal decided in a meeting room. Only the practices that were actually used, and turned out good, get distributed. So they do not float above the floor. They fit the user's hands. Distributing what worked in reality, not what sounds ideal — that difference decides whether it lasts.
Start from "looks convenient," stop with evidence
Why did we start using AI in the first place? The reason is simple. It looked convenient.
We did not start from a difficult theory. It looked faster; it looked like it would free our hands. That was the level of motivation, and I think that is fine. You do not need a cool reason. But if you run on "looks convenient" alone, there is one pit waiting.
AI says "done." Smoothly, and with full confidence. And that "done" is sometimes not to be trusted. It says something works when it does not. It says something is fixed when it is not. So many tests passed; the build succeeded; the state is such and such. It reads those proxy numbers as completion. Without looking at the evidence, the report runs ahead. Leave this alone, and quality actually drops. You just make mistakes faster.
So I decided not to believe this "done." Instead of believing, I look at evidence. I look at logs. I look inside the database. I look at test output. Only what was actually observed counts as done. Counts and statuses are not substitutes for completion. I wrote about this in the published article "AI's 'done' cannot be trusted — a quality gate that stops false completion with evidence."
Why do numbers fool us? Because numbers look exactly like evidence. Hear "a hundred tests passed," and you relax. But those hundred may not touch the one case that matters. A successful build does not mean it works. Numbers reflect a slice of the state. They guarantee nothing about the whole. So I stop just short of the number. What did this number actually verify? Only after that question does it become evidence. Looking reassured and having verified are different things.
"Done" does not pass without evidence
Saying "stop it with evidence" is easy. Words alone are not kept. People hurry. AI wants to move on too. So we made it a shape where you cannot move on without keeping it.
Concretely, we made it a skill — a small executable procedure. A skill is a small set of steps that works automatically inside AI development. When something tries to declare completion, this procedure cuts in. Are there logs? Are there database records? Is there test output? Declare a high-risk completion with none of them, and it stops right there.
The stopping is the point. If it only prints a notice, people skim past it. A yellow warning is invisible to hurried eyes. Even if seen, if you can still move on, the false completion flows downstream. The next person builds on top of it. Rework costs more the later it is found. So we made it impossible to move on. Until the evidence exists, you cannot enter the next step. Quality changed from a verbal promise into a gate you cannot pass.
Some people find this stifling. But with the gate, you actually move faster. You can look only forward, without worrying that things will collapse later. It is not a restraint. It is the ground that lets you move fast with peace of mind.
What is evidence, concretely? It differs per task. If you changed data, it is the database itself, after the change. If you fixed a screen, it is a record of actually opening the fixed screen. If you ran a process, it is the log that process emitted. What they share: they sit on reality's side, not on the report's side. It is not a person saying "I did it." It is reality showing "this is how it is." The gate passes only that reality. It does not pass words.
One person's way, distributed to everyone
One gate alone is still one person's effort. Here is the main point.
One person builds a good way. They register it on the platform. Then it is distributed to every project via Claude Code. The next person who starts a new project has that way from the start. No need to reinvent it from zero. The discipline a teammate found yesterday works in your hands today. On the receiving side, the only operation is a single approval, once.
This is different from training. Training spends the time of both the teacher and the learner. The more people, the more it costs to convey. Word of mouth changes shape with every retelling. A distribution mechanism is different. Build a good way once, and the platform distributes it. Headcount grows; the cost of conveying does not. The shape does not drift. One person's good judgment becomes everyone's starting state.
Think about a new joiner. Install the platform on day one, and the company's quality standards work in their hands from that day. No months of watching a senior's back to absorb how things are done. Good ways stop being something you memorize; they are something you start with. The loss called "practices becoming personal property" is closed here. A good way someone found is not theirs alone. From the next day, it is everyone's. That is what I think team development means.
Good ways compound as they grow. Add one, and the next person to join gains one more. Add ten, and they gain ten. The later you arrive, the richer the ground you stand on. This does not happen when you work alone. Your improvement helps the project of someone you have never met. Someone else's improvement helps your today. This give-and-take runs automatically on the platform. The longer time passes, the thicker the ground gets. That is exactly why building as a team is worth it. The team becomes larger than the sum of its members.
This article was made by this very mechanism
Saying one thing and doing another — that is the worst. So this article itself was made with the same mechanism.
An AI wrote it. But we do not ship what an AI wrote as-is. We make another AI doubt it. Another AI asks back: are the facts right, are there grounds, are there leaps? The writer and the doubter are deliberately different AIs. Let the same AI write and grade itself, and the grading goes soft.
And the final decision belongs to a human. A human reads, a human judges, a human decides to publish. The AIs are crossed so that what one misses, the other picks up. On top of that, a human holds final responsibility. We borrow writing speed from AI. We borrow the doubting eye from AI too. But judgment is the one thing humans do not let go. This order holds for article-making as well. The very text that preaches the mechanism came out through that mechanism. To me, that is the evidence.
Why set up a doubter at all? This approach is called adversarial review. It reads what was written on the assumption that an error is hiding somewhere, and goes looking for counter-evidence. The reason to appoint a doubter is that the author cannot read their own text that way. The grounds you wrote look correct to you. Human or AI, it is the same. So the doubting eye comes from outside. The other AI does not know our circumstances. It does not defer. It calls the odd parts odd. We take the findings and rewrite. If it does not pass, we rewrite again. This back-and-forth raises quality. Passing on the first try is not the goal. The value is in not passing.
Quality is kept by mechanisms, not by attention
Finally, the claim I care about most. Quality is kept not by human attention, but by mechanisms. When we find a contradiction, we do not leave it alone — we turn it into a mechanism, on the spot, and close it.
I did not want more rules. I wanted quality to be the same in everyone's hands. Ideals do not spread as ideals. Good ways stop with one person while they stay personal. The word "done" cannot be trusted without evidence. Efficiency alone does not return time. Every one of these, left alone, stays a contradiction.
So we closed them one by one. "It does not spread" — closed with a distribution mechanism. Verbal promises — closed with a gate you cannot pass. False completion — closed with an evidence gate. The softness of self-grading — closed with another AI and a human's judgment. Saying it keeps nothing. Things are kept only when they are built into the steps.
There is one more good thing about mechanisms. You stop blaming people. When quality drops, hunting for a culprit prevents nothing next time. What you should hunt for is which step was missing. Find the gap, add a gate there. Then the same failure does not repeat from there. Fix the mechanism instead of blaming the person. I believe that is the way that lasts.
Contradictions do not disappear on willpower. They disappear through mechanisms. Turning one person's improvement into everyone's quality — that spread is the value of a team in the age of AI. That is why I built this platform.



Top comments (0)