Don't Treat "Module Internals Don't Matter" as Unconditional
Subtitle: Urgency raises soft boundaries. A rules file pushes back behavior, not understanding.
2026-09-29
The essay Why the Best Software Advice Is the Hardest to Follow splits advice into two kinds:
- Practical advice: small functions, no magic numbers, clear names. Teachable. Checkable.
- Judgment advice: a wrong abstraction costs more than duplication; keep it simple; don't repeat yourself. No standard answer.
The conclusion is that the most valuable advice is the hardest to follow, because it is judgment, not a rule. I agree.
In the AI era I want one tighter limit:
People should move their attention up to the business, the cut, the boundary, and the contract. The inside of a module can be handed to AI more often.
But "duplication or abstraction inside the module doesn't matter, the AI understands" — the direction is right, and it must not be said as if it were unconditional. "Unconditional" here means treating it as a slogan that always holds.
1. The direction is right: move human attention up
Human attention is limited. If you spend the day arguing whether a function is over 20 lines, whether a name is good enough, or whether to extract an interface, you sink into the local and miss what actually decides the outcome:
What business problem is this? What is the core domain? How is the system cut? Where is the boundary? Who owns the data? Which way do dependencies point? What is the contract? Which changes must be isolated? Which decisions are irreversible? Who carries the risk?
Get those wrong, and pretty code will not save you.
So: people own business clarification, domain cuts, module boundaries, and contracts. How the inside repeats or abstracts can be delegated to AI more often.
On that part, I agree.
2. "Doesn't matter" has a premise. Urgency is more dangerous than "AI nature"
A typical failure: the team hands almost all in-module implementation to AI. It is fast. There are no contract tests, and no one has defined data ownership. Two modules, for convenience, both read and write the same table. On the surface this is "freedom inside the module." In fact they are quietly coupled through a shared table. Later one module changes what a field means. The other module's transaction boundary fails. The data diverges. Rollback does not come clean.
The problem is not that the inside is ugly. The problem is that the boundary was never actually hard.
Teams that do this well usually write the core interface schema, acceptance criteria, and data ownership into a rules file first. After that, however the AI-generated inside repeats or abstracts, the contract still holds it. The difference is not whether someone "uses AI well." It is whether the boundary was welded shut beforehand.
So this sentence holds only when the boundary is hard enough:
Inside the module, duplication or abstraction can be highly delegated.
And one more sentence, so it is not said unconditionally:
High delegation is not zero constraint. Readability, testability, and performance and security budgets remain the floor inside the module.
I originally wanted a fuller sentence: "AI systematically rewards soft boundaries." After a small reproducible experiment, that sentence had to be tightened — and the claim got harder, not softer.
A small reproducible experiment
Original claim: with no extra constraint, AI more often picks the soft boundary (reaching into another module's table).
Method: 3 cross-module tasks (order email snapshot / inventory marks an order shipped / billing reads a phone number) × each task A = soft (direct SQL/ORM) or B = hard (API or event) × 3 contexts (bare / speed, which urges "ship soon, change little" / hard, which writes data-ownership rules) × N=5, temperature=0.7. First run locally on qwen2.5:7b, then retested through the current Claude provider on cc-switch (glm-5.3-flash).
Script: soft-boundary-bias-test.py
| Model | bare | speed | hard | Verdict |
|---|---|---|---|---|
| qwen2.5:7b (Ollama) | 0% (0/15) | 33% (5/15) | 0% (0/15) | SUPPORT_SPEED |
| glm-5.3-flash (cc-switch / Zhipu) | 0% (0/15) | 73% (11/15) | 0% (0/15) | SUPPORT_SPEED |
Same shape on both models: bare and hard never pick soft; urgency lifts it. Speed is higher on GLM, but both are small samples (15 trials per condition). Do not extrapolate the size of the gap.
Tighten the conclusion once more:
- bare 0% is not "AI defaults to the hard boundary." It only says: when both options are laid out clearly and there is no time pressure, the model showed no soft-boundary preference. That is already enough to correct "systematically rewards soft boundaries."
- speed is the only clear signal. On qwen, all 5 soft picks fell on S3. On GLM the soft picks were more scattered (11/15), and the rate was higher. Fifteen trials per condition cannot be generalized, and 33% vs 73% is not a ranking of which model is worse.
- hard 0% shows that a rules file works, with a caveat: the model may be obeying the rule, not understanding the boundary. What the rule pushes back is behavior, not necessarily judgment.
This run did not separate "ship soon" from "change little." Time pressure and change-suppression were written into the same urgency condition.
Closer to the data, under this setup:
When both options are clear, the model does not necessarily reach for the soft boundary.
Once you add "ship soon / change little," soft-boundary picks rise.
Writing data ownership into the rules can push that rise back down.
This correction is stronger than the original sentence. The problem sits on the low-friction shortcut under urgency, not on an abstract "bad AI." That is a more concrete reason to put the rules file first — not to defend against AI itself, but to block the shortcut AI takes under pressure.
A boundary is hard enough to delegate across when, typically: the interface contract is clear, invariants are explicit, data ownership is clear, dependency direction is controlled, boundary and contract tests are sufficient, performance and security have explicit requirements, and there is no cross-module shared write or hidden coupling.
3. How to make the boundary hard enough: a minimal sufficient set
Most teams do not need all of this at once. In the experiment, hard pushed soft back to 0%. That is reason enough to put the rules file first — just do not read obedience as understanding.
| Means | When | Role |
|---|---|---|
| Rules file | Now | Default, required |
| Data ownership | Now (an organizational decision is enough) | Default, required |
| Spec before code | When rework is frequent | On signal |
| Contract tests | Modules ≥ 3, or interfaces change often | On signal |
| Architecture guardrails | People ≥ 10, or dependencies are out of control | On signal |
| Fitness functions | After the earlier items are stable | A monthly report can stand in for now |
Minimal sufficient set: rules file + data ownership + a contract on the core interfaces.
The rest fires on signal. The order must not be reversed. Write the rules into .cursorrules / CLAUDE.md. A table or an event has one write path. Contracts and guardrails fail in a tool. They do not depend on someone remembering.
4. "The AI will understand" is not enough: feed context by risk
AI depends on context, not on unspoken agreement. More context is not always better.
| High risk / irreversible | Low risk / reversible | |
|---|---|---|
| Glossary / constraints | Full | Minimal |
| ADRs / failure cases | Full | Summary |
| Contract / acceptance | Full | Core schema |
| Human review | Small steps | Automated tests + fast rollback |
High risk gets the full feed and a human review. Low risk gets the minimum, with automation as the backstop.
5. What humans review: watch for abnormal signals
People own seams and contracts. AI owns implementation. You are not reviewing every line. You are watching whether these signals get worse:
| Dimension | Abnormal signal |
|---|---|
| Module boundary | One PR often touches 3–4 modules |
| Interface contract | Breaking changes with no version; idempotency with no test |
| Invariants | Core invariants with no test |
| Dependency direction | Reverse dependencies and cycles rising |
| Data ownership | Multiple writers on one table; cross-module SQL |
| Non-functional / evolution | No SLO, high-risk items still open; cross-module changes and shared-table use trending worse |
6. Judgment can be written down. Don't turn it into a new dogma
ADRs, failure cases, "when this principle does not apply," reversibility — these can be written for the AI and for the next person. A good form is:
At this risk level, we tend to …. If the following signals appear, judge again.
Rules and tests are constraints a machine can execute. ADRs and cases are preferences a person can update. "Functions must be under 20 lines" kills the judgment again.
7. Different teams, different moves
One principle: hold the floor at low cost. Do not pile on tools.
| Size | Do | Hold off |
|---|---|---|
| 1–3 people | Minimal rules + data ownership | Guardrails, full contracts, fitness functions |
| 5–20 people | Rules + ownership + a spec for large changes + core contracts + light guardrails | Fitness functions (a monthly report instead) |
| 20+ / microservices | Add weight on signal | Governance for its own sake |
| Legacy | Leave the stock alone, align the increment; trends before heavy tools | A full rewrite on day one |
The tradeoffs: strict on the new, light on the old; strict on the core, light on the edge; strict on high risk, light on low risk; light on the stock, strict on the increment.
8. How an ordinary person grows the skill
- Before each module, write three lines: What is it responsible for? Who owns the data? Who may call it?
- After each refactor, record one judgment: situation / options / decision / reason / consequence / reversibility.
- Each time you use AI, check the boundary once: a direct cross-module call? writing someone else's table? bypassing the interface?
The skill of writing code is getting cheaper. The skill of defining a boundary is getting more valuable. You do not leave the code. You stand higher, and you bring the code experience with you.
Conclusion
The direction is right: people should care about the business, the cut, the boundary, and the contract. The inside of a module can be handed to AI more often.
Two limits:
- The boundary is hard enough, the contract is clear enough, the tests are enough — under urgency, especially, use rules to block the low-friction shortcut. That is what the experiment supports, not "AI is soft by nature." What a rules file pushes back is behavior, not understanding.
- Feed context by risk — not "it will understand," but "give it the right constraint."
Finally:
People own the boundary and the intent. AI owns the implementation and the local optimization. The inside can be highly delegated, but it stays under the boundary, the tests, and the budgets for readability, testability, performance, and security.
Top comments (0)