DEV Community

Cover image for Who Made This Decision?
Augusto Chirico
Augusto Chirico

Posted on Originally published at augustochirico.dev

Who Made This Decision?

We built a product that, at the old pace, would have taken a year before there was anything to show. We had it in two months. It worked, it looked good, and every feature that was asked for showed up.

Then it was time to evaluate a small change. It meant touching 80 files.

Not because the change was big, but because the code had no structure. It was one big mass that had grown with a single criterion: make the feature exist. There was dead code left over from one change to the next. There were stores scattered everywhere: sometimes one was used, other times a new one had been created that coupled to the previous one but behaved differently. And a core feature, one that can be built on several technologies of very different nature, had no abstraction separating what it does from how it does it. Nothing prevented the dependency rules from being broken, and they were.

And there were tests everywhere. Thousands. Many of them checked things like whether a CSS class was present in the HTML, and they were going to break as soon as we wanted to change the UI. Others existed because the TDD cycle, applied without judgment, forced us to fix and extend them with every change. More than a safety net, they were a toll.

Looking at that code, the question I was left with wasn't "who wrote this?". It was a different one: who made these decisions?

Nobody. Or rather, the model, feature by feature, without anyone reviewing them.

โšก The speed is real

It's worth starting here, because it's easy to read all of the above as a complaint about AI. It isn't.

Two months versus a year isn't marketing. AI writes code fast, and most of the time it writes it well. There's some romanticism among those of us who have been doing this for years, and I include myself, when we resist accepting that an agent can write all of our code. It can.

The data backs this up. Google's 2025 DORA report found that, for the first time, AI adoption is associated with higher throughput: teams ship more. But the same report has another finding that tends to stay in the background: AI adoption is still associated with lower stability in production. The time saved writing goes into auditing, verifying, and fixing.

DORA's conclusion is that AI works as an amplifier. It multiplies what you already have. If there's judgment, it multiplies that. If there isn't, it multiplies that too.

The problem isn't who writes the code. It's who decides.

๐Ÿ” What the demo doesn't show

Every feature an agent implements comes with implicit decisions. Where the state lives. Which module depends on which. Whether to reuse what exists or create something new. Whether a technology is used directly or behind an interface.

Nobody asked the agent to make those decisions. It made them because it had to in order to close the task. And an agent optimizes for closing the task.

GitClear analyzed 623 million code changes between 2023 and 2026, and summed it up better than I could:

"The headline is not 'AI writes bad code.' It is that today's default AI workflow is incentivized to deliver atomic code (a happy path, a passing test, a closed ticket) while quietly taxing the invisible and the deferred."

The numbers from that same study describe almost exactly the product I was talking about:

Signal Change
Refactoring (% of changed lines) 21% in 2022 โ†’ 3.8% in 2026
Duplicated code blocks +81% since 2023
Cross-file calls (reuse) โˆ’35% since 2023
Error-masking constructs +47%

That's where the scattered stores come from. For an agent, creating a new store is cheaper than understanding the one that already exists. It does that a hundred times, and each one makes sense on its own. Together, they're the 80 files behind a small change.

And we're not reading them. In Sonar's State of Code 2026 survey, 96% of developers say they don't fully trust AI-generated code. Only 48% always verify it before committing. We distrust it, and we accept it anyway.

Addy Osmani put a name to what builds up: comprehension debt. It's the gap between how much code exists in your system and how much of it any human actually understands. Technical debt warns you: slow builds, modules nobody wants to touch. Comprehension debt doesn't. It lets you move forward with confidence until the day a small change touches 80 files.

๐Ÿงฉ The feature without an abstraction

The clearest case was the core feature. I'll simplify it with an example that isn't the real code, but has the same shape. Say saving a draft can go to the cloud or stay local, depending on the context.

This is what it looks like when nobody decides the abstraction:

// Every consumer knows the concrete technologies
import { uploadToS3 } from '../infra/s3';
import { saveToIndexedDb } from '../infra/idb';

export async function saveDraft(draft: Draft, online: boolean): Promise<void> {
  if (online) {
    await uploadToS3(`drafts/${draft.id}.json`, JSON.stringify(draft));
  } else {
    await saveToIndexedDb('drafts', draft);
  }
}
Enter fullscreen mode Exit fullscreen mode

None of this is wrong on its own. The problem shows up when that function exists in ten places, each with its own variant, and adding a third technology means finding all of them.

This is what it looks like when someone decides that the domain defines the behavior and the infrastructure implements it:

// The domain defines what it needs
export interface DraftStorage {
  save(draft: Draft): Promise<void>;
  load(id: DraftId): Promise<Draft | null>;
}

// The infrastructure decides how
export class CloudDraftStorage implements DraftStorage {
  // ...
}

export class LocalDraftStorage implements DraftStorage {
  // ...
}
Enter fullscreen mode Exit fullscreen mode

An agent can write either version. If nobody tells it the second one is the rule, it almost always writes the first, because it's the shortest path to closing the task.

And telling it in a prompt isn't enough. A rule written in a markdown file is something the agent can ignore. A failing check isn't:

// .dependency-cruiser.cjs
module.exports = {
  forbidden: [
    {
      name: 'domain-does-not-know-infra',
      severity: 'error',
      from: { path: '^src/domain' },
      to: { path: '^src/infra' },
    },
  ],
};
Enter fullscreen mode Exit fullscreen mode

With that rule in CI, the shortcut stops being the shortest path.

๐Ÿงช Thousands of tests, little confidence

Something similar happened with the tests, and it's harder to see because it looks like a virtue. A project with thousands of tests sounds like a safe project.

But most of the tests we found were checking implementation details:

// Checks that the HTML has a CSS class
it('renders the save button', () => {
  const { container } = render(<Editor />);
  expect(container.innerHTML).toContain('btn-primary');
});
Enter fullscreen mode Exit fullscreen mode

That test protects nothing a user cares about. The only thing it guarantees is that, the day the button's design changes, someone will have to fix it. Multiply that by a thousand, and every UI change drags along a massive rewrite of tests that weren't testing behavior.

The same case, testing what matters:

// Checks that the user can save a draft
it('saves the draft when the user clicks save', async () => {
  const saved: Draft[] = [];
  const storage: DraftStorage = {
    save: async (draft) => { saved.push(draft); },
    load: async () => null,
  };

  render(<Editor storage={storage} />);
  await userEvent.click(screen.getByRole('button', { name: /save/i }));

  expect(saved).toHaveLength(1);
});
Enter fullscreen mode Exit fullscreen mode

This test survives any redesign and only breaks if saving stops working. And notice that it exists thanks to the abstraction from the previous section: without DraftStorage, there would be nothing to swap in the test.

It's not TDD's fault. It's TDD without judgment: the agent applies the "test first" rule to every detail, including the color of a button, because nobody told it which behavior is worth protecting. The number of tests becomes a vanity metric, and every change pays the cost.

๐Ÿญ Not all software is the same

None of this matters much in a prototype. A proof of concept exists to validate an idea, and if the idea holds up, you throw it away and build it properly. There, vibe coding is exactly the right tool.

Prototype or PoC Product with 100,000 daily users
Why it exists To validate an idea To sustain change for years
Expected lifespan Weeks Years
Cost of bad design Throw it away and start over Every feature costs more than the last
What to delegate Almost everything Implementation, within defined boundaries
What to review That it works What the agent decided to make it work

The real risk sits in between: when the prototype becomes the product without anyone deciding it. That's what happened to us. The code from the first days, written to show something fast, became the foundation everything else was built on.

๐Ÿงญ The decision budget

The most useful way I've found to think about this is to separate the decisions I don't delegate from the ones I do.

The ones I don't delegate:

  • The data model and where state lives. How many stores there are, who owns each piece of data.
  • Module boundaries and the direction of dependencies. What can import what.
  • The abstractions behind core features. What's domain and what's infrastructure.
  • Public contracts. APIs, events, schemas that others consume.
  • New dependencies. Every package that comes in is a decision with consequences.
  • Security and error handling. Where shortcuts stay invisible until it's too late.
  • What gets tested. Which behavior is worth protecting, and what's a detail that can change.

The ones I do delegate:

  • Implementation within those boundaries.
  • Writing the tests, once I've decided which behavior they cover.
  • Mechanical refactors, with a clear target.
  • Boilerplate.

AI still writes most of the code. What changes is that the decisions that give it structure are mine, and they come first.

๐Ÿ› ๏ธ How I do it today

It's not a finished method. It's what's working for me.

1. Design before the prompt. Before asking for code, I write a spec, then a plan, and only then does implementation happen. The most important review isn't the one on the diff, it's the one on the plan. That's where changing course costs a paragraph, not 80 files.

2. Rules written where the agent reads them. Architecture decisions live in CLAUDE.md and in project skills, not in my head. If a rule isn't written down, it doesn't exist for the agent.

3. Guardrails that fail. Strict types, lint rules, dependency rules, all in CI. A markdown file can be ignored. A broken build can't.

4. Two-layer review. An agent reviewer, in my case Copilot on PRs, finds bugs. But the question no automated reviewer asks is the one that matters: what did this change decide that I didn't ask for? A new store, a new dependency, an abstraction that didn't exist. That part I review myself.

5. Ask the agent to declare its decisions. Before implementing, I ask it to list what it's going to decide on its own:

Before implementing, list every design decision this change needs
that is not already in the plan: new files, new state, new
dependencies, new abstractions. Wait for my approval.
Enter fullscreen mode Exit fullscreen mode

Something I didn't expect almost always shows up. Better to see it in a list than in a 40-file diff.

6. Small PRs. A diff that can be read in full is a diff someone actually reads.

๐Ÿข Slower on purpose

All of this makes me slower. On purpose.

I don't say it out of nostalgia. Two months to have a product is a huge result. But the full math isn't "two months versus a year". It's two months, plus what every later change costs on top of a foundation nobody designed. Maybe four months with judgment would have been faster than two without it.

Simon Willison proposed a name for this way of working: vibe engineering. Using agents to go faster, while staying accountable for every line that reaches production. Tests, planning, documentation, review. The same as always, with a new collaborator who writes very fast and decides without telling you.

I wrote about the more personal side of this in another piece. Here I'm interested in the practical side: AI as an amplifier, not as a replacement for judgment.

๐Ÿ”ญ What's next

The future of this craft is probably learning to manage agents: designing the harnesses they follow, the guardrails that contain them, the context each product needs. We're already there.

But today AI still makes decisions for us, and it does so in a way that doesn't show in the demo. It shows three months later, when someone tries to change something small.

๐Ÿ“‹ TL;DR

Idea In practice
The speed is real Two months versus a year isn't marketing
The cost is somewhere else Less stability, more duplication, less refactoring
The problem isn't who writes It's who decides
More tests isn't more confidence A test that checks a CSS class is a toll, not a net
Not all software is the same In a PoC you delegate almost everything; in a product, you don't
Decision budget State, boundaries, abstractions, contracts, dependencies, security, and what gets tested aren't delegated
Guardrails that fail A rule in markdown gets ignored; a check in CI doesn't
Slower on purpose Four months with judgment can be faster than two without it

AI can write all of our code. What it still can't do is own the decisions it made to write it.

Top comments (0)