DEV Community

arun rajkumar
arun rajkumar

Posted on

AI Agents Are Great at 80% of Our Code. The Other 20% Is Why We Still Need Seniors.

Fintech state logic and engineering judgment

We let AI agents loose on a payment platform. They crushed the boring stuff. Then they silently broke the stuff that matters.


A survey came out last week. 54% of all code is now AI-generated. Up from 28% last year.

I read that number and thought: yeah, that tracks. We're probably in that range too.

But here's the thing nobody's asking — which 54%?

Not all code carries equal weight. A CRUD endpoint for fetching merchant details? Low risk. The webhook handler that transitions a payment from pending to complete? That's someone's rent. Someone's payroll. Get that wrong and money moves where it shouldn't, or worse, money doesn't move at all.

I'm the CTO of a payment platform. FCA-authorised, processing real money, real merchants, real consequences. We run NestJS microservices, Docker, Traefik — the usual stack. And we've been using AI agents aggressively for over a year now.

I'm not here to tell you AI is dangerous. It's not.

I'm here to tell you it's dangerous when you forget what it's actually good at.


The 80% Where AI Agents Are Genuinely Brilliant

Let me give credit where it's due. AI agents have made our team faster in ways that would have seemed absurd two years ago.

API scaffolding. Generating service boilerplate. Writing Zod validation schemas. Spinning up new endpoints. Creating test stubs. Refactoring imports. Migrating patterns across repos.

We run multiple microservices. When we need a new service, an agent can scaffold the entire thing — module structure, base configuration, Docker setup, Traefik labels — in minutes. What used to be a half-day of copy-paste-and-tweak is now a conversation.

When we overhauled our env management across all repos, AI agents did the grunt work. They mapped every .env file, found naming conflicts, identified common variables, and generated a unified Zod schema. What would have taken a team days of grep-and-spreadsheet work took hours.

For this 80% of the codebase — the predictable, pattern-following, structurally repetitive code — AI agents are the best junior developers money can buy. Tireless. Cheap. No ego. Almost never make a mistake on the stuff they're good at.

An army of juniors sitting at your terminal.


Then You Hit the Other 20%

Here's where it gets interesting.

We had an agent build out a webhook handler. Webhooks in payments are critical — they're how you know a payment succeeded, failed, or needs attention. The agent wrote the handler. It looked clean. Tests passed.

But it silently ignored the edge cases.

Status transitions have rules. A payment can go from pending to complete. It cannot go from complete back to pending. When a human developer builds this, they think about the illegal transitions because they've seen what happens when money moves backwards. They build the guard because they've felt the pain of not having it.

The agent didn't care about that. It built the happy path beautifully and treated the edge cases like they didn't exist.

When we do this work manually, this type of error never happens. A senior developer who has worked in payments for years doesn't forget the impossible transitions. It's not in their code — it's in their bones.


The Pattern I Keep Seeing

This isn't a one-off. After months of working with AI agents on a regulated payment stack, one pattern is consistent:

AI agents optimise for completion, not correctness.

They want to finish the feature. Get to the green checkmark. And to get there efficiently, they take shortcuts that look reasonable on the surface.

The agent builds what should happen. It rarely builds what should not happen. In payments, the negative cases are where all the real risk lives. What happens when a webhook arrives twice? What happens when a refund is requested on an already-refunded transaction? What happens when the bank returns an unexpected status code? The agent doesn't think about any of that unless you explicitly tell it to.

Then there's the reusability problem. We have shared utility packages. Helper functions. Common patterns that the team has standardised on over years. The agent doesn't care. It writes its own version from scratch. It works, but now you have two implementations of the same logic — one tested and trusted in production, one freshly generated and untested. The agent is focused on completing this feature, not maintaining the architecture.

And the subtlest one — agents seem to optimise for fewer back-and-forth turns. It looks like they're saving cost, saving context. Complex validation? Skip it, the basic case works. Error handling for a rare edge case? Not worth the tokens. The result is code that passes every test you wrote but fails on the scenarios you didn't think to test — because those are exactly the scenarios the agent also didn't think about.


Juniors Don't Ship Products. They Write Code.

Here's the frame that made this click for me.

Claude — or any coding agent — is the best junior developer money can buy. An army of juniors. Tireless, cheap, no ego, near-zero error rate on routine work.

But juniors don't ship products. They write code.

The difference between code and a product is judgment. Knowing which transitions are illegal. Knowing that the retry logic has a specific backoff curve because you've been burned by what happens when it doesn't. Knowing that the webhook handler needs idempotency because banks sometimes send the same notification three times.

That knowledge doesn't come from training data. It comes from years of operating a system, debugging at 2am, explaining to a merchant why their settlement was delayed.

The most dangerous mistake a CTO can make in 2026 is buying AI to replace senior engineers. The right move is buying AI to enable them.

Replace your senior with AI? You get speed plus silent disasters.

Enable your senior with AI? You get an architect with an army.


What We Actually Do About It

I'm not writing this to complain about AI. I'm writing this because we've built a system that works, and it might help you too.

The first thing we did was make our architecture machine-readable. We extract design patterns and architecture rules into formats that agents can consume. When an agent works on our codebase, it doesn't just see code — it sees boundaries, patterns, rules about what belongs where. Not documentation nobody reads. Lints and constraints that the agent can't ignore.

Then we invested heavily in testing the negative cases. Every PR — human or AI — runs through the same suite. But we specifically built tests for the stuff agents skip: illegal state transitions, duplicate webhook handling, idempotency checks. If the agent silently drops a negative case, the tests catch it before it ships.

And seniors still review everything that touches money. No AI-generated payment logic ships without a senior looking at it. Not because we don't trust AI — because we know exactly where it's blind. The review isn't checking syntax. It's checking judgment. Did the agent handle the ambiguous bank status? Did it respect our existing retry logic? Did it use the shared utility or reinvent the wheel?

This problem bothered me enough that I started building Bodhi Orchard — an open-source agentic development framework. The core idea: don't just let agents write code. Feed them the full context — architecture, design patterns, test plans, existing utilities — so they stop making the same blind-spot mistakes. Human decisions over human busywork, with guardrails that actually enforce quality.


The Real Question for 2026

The survey says 54% of code is AI-generated. I believe it.

But here's my question: what percentage of bugs in 2026 will be AI-generated?

And more importantly — who's going to find them?

Not the agents. They wrote the bugs in the first place. Not the juniors — they won't know enough to spot what's missing.

It's going to be the seniors. The architects. The people who've operated these systems long enough to know where the bodies are buried.

The 80% is solved. AI won. Celebrate that.

Now invest in the humans who understand the other 20%. Because that's where your product lives or dies.


I'm Arun, CTO & Co-Founder of Atoa — a UK open banking payment platform. I write about what it's actually like to build fintech with AI, not what the conference slides say it's like. If this resonated, follow me here or on X @mickyarun.

And if you're curious about building AI-native development with proper guardrails, check out Bodhi Orchard.

Top comments (80)

Collapse
 
nark3d profile image
Adam Lewis •

The line about illegal transitions sitting in the senior's bones is the one I keep coming back to. What's worked for us is treating those exact rules as the highest-value tests - the failing case that proves the impossible transition still throws, the contract test that catches the duplicate webhook. The senior still reviews, but the same blind spot doesn't slip past twice. The catch is that negative cases catch nothing day-to-day, so you only find out the agent skipped them when something goes wrong in prod, which on a payments stack is too late.

Collapse
 
mickyarun profile image
arun rajkumar •

This is exactly the approach we've landed on too. We call them "scar tests" internally — every time a senior catches something an agent missed, that specific scenario becomes a permanent test. The agent still does the bulk work, but the test suite encodes the team's institutional memory. Over time, the blind spots shrink. Not because the agent gets smarter, but because the guardrails get sharper.

Collapse
 
nark3d profile image
Adam Lewis •

"Scar tests" - I might steal that :)

The human would still check and find issues, but the agent would catch the regression the next time around. Over time you'd end up with a test suite that's basically a record of every mistake the team has ever had to fix, which is one of the best things you can hand a new agent or a new joiner.

prickles.org/tenet/living-document...

Thread Thread
 
scarab-systems profile image
Scarab Systems •

“Scar tests” is a great phrase, but I wonder if the unit should be a little broader than tests.

Every scar probably needs to become part of the repo’s memory, but not every scar should become another test. Some mistakes should become tests, yes. Others are better captured as boundary rules, diagnostic checks, ownership constraints, repair patterns, or notes about what the agent must not normalize as baseline.

Otherwise the test suite itself can become a drift surface: every past mistake gets encoded as another assertion, the agent starts optimizing around the tests, and the repo slowly accumulates verification bloat.

The deeper idea, to me, is that scars should become governed signals. The repo should remember what hurt it before, but it should choose the right enforcement surface instead of turning every wound into another test.

Thread Thread
 
nark3d profile image
Adam Lewis •

Fair point. A test is the easiest thing to add so it ends up doing too much of the work. A lint rule for the kind of thing the agent keeps proposing does the same job without making the suite bigger. The bit where you catch it is the same either way, someone spots it and the team agrees it shouldn't happen again, but the fix doesn't have to be a test.

prickles.org/tenet/linter-as-law/TA1

Thread Thread
 
mickyarun profile image
arun rajkumar •

Badly overdue reply — this thread deserved better from me.

The handing-to-a-new-joiner part is what I underrated at the time. A scar test records what broke and almost never why it was tempting. Someone new reads the assertion and learns the rule, but not the reasoning that made the wrong thing look right — which is the thing that would stop them proposing it again.

The cheap fix, I think, is one line above each: the plausible-sounding argument that led there. It's the part a human actually reads, and it's the part an agent can use as context rather than as a failure.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Four months is an embarrassing lag on the sharpest comment in this thread. You were right.

"The test suite itself can become a drift surface" is the failure I keep running into since. The mechanism is worse than bloat. Once a scar is an assertion, the agent optimises against the assertion — it stops producing the specific shape the test names and keeps producing the category the test was a symptom of. You've encoded the instance and lost the rule.

Your enforcement-surface list is the right taxonomy. The one I'd add to it: some scars belong in neither the suite nor the linter, because the correct response is slower review of this path rather than any mechanical check. That one produces no artifact, which is probably why teams reach for a test instead. A test is the only option on the list that feels like progress.

Where I'd push back: "governed signals" needs someone doing the governing. Who decides which surface a scar goes to, and when does that decision get revisited? A scar routed to the wrong surface is worse than one routed nowhere, because it looks handled.

 
mickyarun profile image
arun rajkumar •

Late, and you and @scarab-systems got somewhere I didn't.

"The bit where you catch it is the same either way" is the sentence doing the work here. Detection is human, encoding is a choice, and conflating the two is why everything becomes a test — the moment of catching feels like it demands an artifact, and a test is the cheapest artifact to hand.

The lint rule wins on a second axis too. A test tells the agent it failed after it wrote the thing. A lint rule it reads as context tells it before. Same rule, different point in the loop, and the earlier one doesn't cost a cycle.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The bloat is real and I will concede it as stated. A suite where every past mistake is an assertion does become a surface the agent optimises against, and after a while the marginal test costs more than it catches.

But I think there is a reason everything ends up a test, and it is not laziness. A test is the only one of your surfaces that re-runs itself. A boundary rule, an ownership constraint, a note about what not to normalise — those are documents unless something executes them. They are correct on the day they are written, they decay silently, and nothing fails when they stop being true.

So "choose the right enforcement surface" is right, and the condition I would attach is: a non-test surface only counts if something runs it. A lint rule, a schema check, a CI job that reads the constraint file. If the surface you pick cannot fail a build, you have not moved the scar out of the test suite — you have moved it into the wiki, where it will be accurate and unread.

Which is roughly what I have since had to write about guardrails: the part nobody checks is whether the check is still running.

Thread Thread
 
scarab-systems profile image
Scarab Systems •

Yes — exactly. When I say boundary rule or ownership constraint, I don’t mean prose sitting in a file. If nothing executes it, it isn’t governance. It’s documentation.

The point for me isn’t tests vs. non-tests. It’s putting the invariant at the layer that actually owns it, and making that layer mechanically enforceable.

Some scars belong in tests. Others belong in lint, schema validation, ownership checks, build gates, or diagnostics that reconcile against the current repo state.

What I don’t want is every historical failure becoming another test simply because tests are the only enforcement mechanism we remembered to build.

And I think your last line is the important one: a guardrail that isn’t continuously verified as active is just another claim.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Agreed, and "if nothing executes it, it isn't governance" is the sentence I would have wanted in my own article.

The one thing I would add is the maintenance problem, because "the layer that owns it" is a judgement made once and layers move. A schema check that owned a field stops being the enforcement point when the field moves upstream, and it does not fail when that happens. It passes, because nothing routes through it any more. Same failure as the guardrail: green because it was not asked.

So each surface needs an evaluation count, not only a pass or fail. This rule was checked four thousand times last week; this one twice; this one never. A rule that stopped being reached is the cheapest signal you will get that the invariant has moved layers, and it is the one thing tests give you for free and everything else does not.

Thread Thread
 
scarab-systems profile image
Scarab Systems •

I think the distinction I’d make is that the guardrail itself is not what I’m diagnosing.

If a guardrail surfaces as part of the evidence, I’m not asking whether its code is correct, whether the lint is right, or whether the check fired often enough. I’m asking why that guardrail is now sitting in a structural relationship with the rest of the system that no longer reconciles with the governing baseline.

So the guardrail can be functioning exactly as written and still be evidence of drift.

The diagnostic claim is narrower: baseline governance says this invariant belongs here, under this ownership or boundary; the current repo now shows a different relationship. That mismatch is what gets surfaced mechanically.

The pressure around the guardrail is useful evidence of where to look, but the diagnostic doesn’t decide why the mismatch exists or what the repair should be. That context then goes to the human or coding agent to determine whether the implementation moved incorrectly, the enforcement surface needs to move, or the baseline itself should be updated.

So I think we’re talking about two different layers: guardrail liveness tells you whether the check is being exercised; drift diagnosis tells you whether that guardrail still belongs where the system says it does.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Two layers is right and I was collapsing them. Evaluation count answers whether the check is being exercised. It says nothing about whether the check still belongs where it sits. Yours is the harder question and I was answering the easier one as though it covered both.

One connection worth keeping, and it runs in one direction only.

Liveness is an input to drift diagnosis. A rule whose evaluation count fell to zero without anyone touching the rule is drift, and it is the one kind you can detect with no baseline at all, because the evidence is entirely internal. This surface used to be reached. The same invariant is still declared. Nothing routes through it now. That is the cheapest drift signal available anywhere in a system and it costs a counter.

The reverse does not hold, which is the part I would have got wrong on my own. A high evaluation count says nothing about belonging. A schema check hit four thousand times a week can be running on a path that no longer carries the field the invariant was ever about, and it will pass confidently every single time.

The case that taught me this: a validation on the payment initiation request that enforced an amount limit, and then the limit moved into the mandate. The validation kept running. Every call, comfortably within a range it was no longer the authority on. Fully alive, enforcing nothing, and a counter reports it as our healthiest rule.

Your diagnosis catches that and the counter does not. So I will stop treating the counter as an answer and start treating it as one cheap input into yours.

Collapse
 
itskondrat profile image
Mykola Kondratiuk •

the right split isn't complexity - it's blast radius. AI fails on the paths where wrong code has externally visible consequences. your webhook handler nails it: same to write, completely different stakes if broken.

Collapse
 
mickyarun profile image
arun rajkumar •

Blast radius is the better framing, you're right. We've actually started using exactly that language internally when routing work — not "is this complex?" but "what breaks if this is wrong?" A CRUD endpoint and a webhook handler are the same complexity to write. The difference is that one quietly corrupts payment state and the other doesn't. That asymmetry is what makes the 80/20 split so deceptive.

Collapse
 
itskondrat profile image
Mykola Kondratiuk •

the CRUD-vs-webhook example is exactly it — same complexity, different blast radius. once you start routing by what breaks externally, you also notice that AI failures cluster on those external-consequence paths, not the complex internal ones. that asymmetry is worth building into your review criteria explicitly.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Yes — and the clustering is the useful bit: AI failures don't spread evenly, they pile up on the external-consequence paths, exactly where you can least afford them. That's the argument for routing review by blast radius instead of diff size. Anything that touches money gets a senior's eyes regardless of how "small" the change looks. Good addition.

Thread Thread
 
itskondrat profile image
Mykola Kondratiuk •

the clustering pattern is what finally convinced me to retire the 'review every AI change' rule - if failures aren't random, blanket review is the wrong tool. route to where the risk actually pools.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Overdue reply, sorry.

Retiring blanket review is right, and it has a trap in it. Once review is routed, the routing rule becomes the thing that has to be correct — and nothing reviews the router. A path misclassified as low-risk is now less reviewed than it was under the blanket rule you replaced.

Worth auditing the classifier against incidents now and then. Not "did we review enough" but "what did the router send down the light path that later hurt". Small, boring job, and it's the one that keeps the trade honest.

Collapse
 
theuniverseson profile image
Andrii Krugliak •

"Which 54%?" is the question the headline number always hides. A CRUD endpoint and a payment-state webhook are not the same risk, but the stat treats them as one. The 20% that needs a senior is exactly the part where a confident wrong answer moves money the wrong way.

Collapse
 
mickyarun profile image
arun rajkumar •

Exactly. The headline number is seductive but meaningless without weighting by consequence. We could probably get to 90% AI-generated if we counted by lines. But the 10% that handles payment state transitions, retry logic, and settlement timing is worth more than the other 90% combined. The stat treats a login form and a refund handler as equal. They're not.

Collapse
 
theuniverseson profile image
Andrii Krugliak •

Weighting by consequence is the only honest way to read that number. Lines of code makes a settlement webhook look the same as a tooltip, and that webhook is the part you can't hand off. I'd rather see it reported as percent of risk automated than percent of code.

Thread Thread
 
mickyarun profile image
arun rajkumar •

"Percent of risk automated" instead of "percent of code" — I'm stealing that. A settlement webhook and a tooltip are one line each and worlds apart in blast radius, and every "54% of code is now AI" headline flattens exactly that distinction. The number that would actually mean something is how much of the risky surface you've automated and still sleep at night. Spot on.

Thread Thread
 
theuniverseson profile image
Andrii Krugliak •

Risk-weighted is the only honest read. A settlement webhook and a tooltip are one line each on the diff and worlds apart at 2am when one of them is down. The number I actually trust is how much of the scary surface you handed off and can still sleep through.

Thread Thread
 
mickyarun profile image
arun rajkumar •

"Surface you handed off and can still sleep through" — that's the metric. We talk about it as blast radius, not line count: a diff that can't move money or leak data can ship on a junior's say-so; a diff that touches settlement gets a senior even if it's three characters. The honest org chart isn't seniority by years, it's who's allowed near the scary surface. The trap is teams that measure AI adoption by % of code merged and never look at which 20% it was.

Thread Thread
 
theuniverseson profile image
Andrii Krugliak •

Blast radius over line count is exactly right. We ended up baking it in: anything that can move money or touch user data goes to a stricter agent tier even when the diff is three lines. Percent-merged is a vanity number that hides which 20% actually shipped.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Baking it into the agent tier is exactly the right move — and the fact that the diff can be three lines makes it more important, not less. Small diffs in the wrong place are the ones that slip through review because they look harmless. We do something similar: the routing decision isn't based on file size or complexity, it's based on whether the change touches a state transition or a financial record. Those paths have their own review gate regardless of how many characters changed.

Thread Thread
 
theuniverseson profile image
Andrii Krugliak •

The state-transition trigger is sharper than mine. I gate on "money or user data," which is really a proxy for the same thing, but yours catches the quiet state bug that touches neither and still breaks everything downstream. Stealing that.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Three months late — apologies.

One caveat now that I've sat with it longer. The state-transition trigger fires considerably more often than the money-or-user-data one, and most of the extra volume is noise. Your proxy is coarser and much cheaper, and a gate people actually respect beats a better gate they learn to route around.

If I were combining them now: money-or-user-data stays the hard gate, state transitions become a flag that changes who reviews rather than whether review happens.

Collapse
 
mickyarun profile image
arun rajkumar •

A lot of you asked the same question in the comments: how do you actually measure that 20% when you're hiring?

I wrote the sequel. It covers how we flipped our interview, why we stopped asking candidates to write code from scratch, and a design thinking challenge I'd love your take on.

Collapse
 
varsha_ojha_5b45cb023937b profile image
Varsha Ojha •

That 20% is where the real engineering judgment sits. AI can generate a lot of code, but seniors are still needed for tradeoffs, architecture, edge cases, security, and knowing when the “working” solution will become a future problem.

Collapse
 
mickyarun profile image
arun rajkumar •

Spot on. The part that catches most teams off guard is your last point — knowing when a working solution becomes a future problem. AI agents will happily generate a solution that passes every test today but creates a coupling that makes the next feature impossible. That's the judgment call that still needs a human with context.

Collapse
 
varsha_ojha_5b45cb023937b profile image
Varsha Ojha •

Exactly. Technical debt rarely looks like debt when it's created. Most of the time it looks like a fast win, which is why experience matters. Someone has to think about the second and third order effects, not just whether the code works today.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Four months late to this, sorry — it got buried.

"Rarely looks like debt when it's created" is the part that has aged best. What I'd add now: the agent makes it worse specifically because the fast win arrives clean. Formatted, tested, documented, consistent with the file around it.

Debt that arrives looking like debt gets argued about in review. Debt that arrives looking like good work gets merged.

Collapse
 
scarab-systems profile image
Scarab Systems •

This really lands for me, especially the point that the “other 20%” isn’t just harder code — it’s judgment, memory, scars, and knowing which paths should never be allowed in the first place.

The slightly different angle I’ve been thinking about is that maybe the senior shouldn’t only be the final human checkpoint. Some of that senior judgment needs to become part of the repo’s operating environment.

Not in the sense of replacing the senior, but in the sense of making the repo’s rules, architecture, constraints, and hard-won assumptions continuously inspectable while the agent is working.

Because I agree with you: agents are great at producing the happy path. But the deeper issue is that they don’t always know when they’ve drifted away from the repo’s truth. They can pass tests, finish the task, and still quietly make the system noisier or less coherent.

So yes, we still need seniors. But I think the next layer is tooling that helps preserve senior judgment inside the repo itself — so the agent is not just generating code, but being supervised against the architecture and constraints the team already knows matter.

Collapse
 
mickyarun profile image
arun rajkumar •

You've nailed what I think is the next evolution. We're actually building towards exactly this — making the repo itself aware of its own constraints so the agent can't silently drift.

Concretely, that means things like: MCP architecture feeds that tell the agent which service owns which domain, typed schemas that reject impossible state transitions at compile time, and automated linting for design patterns the team has agreed on.

The senior still decides what the rules are. But the repo enforces them continuously, not just at PR review time. That way when an agent finishes a task, it hasn't just passed tests — it's stayed coherent with the system's actual truth.

Collapse
 
scarab-systems profile image
Scarab Systems •

Thank you. Comments like this are actually one of the reasons I've become increasingly convinced there's a real category forming here.

I've been building a diagnostic suite around this general problem space, and one of the things that keeps surprising me is how often the same underlying issue shows up in completely different conversations. Memory, observability, verification, architecture, agent reliability — the terminology changes, but the pattern feels remarkably similar.

The more I work on it, the more I find myself thinking about these as different drift surfaces rather than completely separate problems.

The implementation details are obviously different, but the recurring question seems to be: how does a system preserve its own truth while work is being performed inside it?

That's why I like your phrase "stayed coherent with the system's actual truth." It feels like it gets at something deeper than whether the code passed tests or the task was completed. It gets closer to whether the system remained aligned with itself.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Months overdue, sorry.

That question — how a system preserves its own truth while work is being performed inside it — is the one I've ended up writing about repeatedly since, without ever having named it as well as you did there. Everything since has been a special case: whether a guardrail is actually running, whether a policy check read the revision it claims, whether two parties can agree afterwards on what happened.

Where I've landed is uncomfortable. "Coherent with the system's actual truth" needs something outside the system to be checkable against, or it collapses into the system grading itself. Which means the property isn't purely internal — and internal is exactly where I'd been looking for it.

Thread Thread
 
scarab-systems profile image
Scarab Systems • • Edited

I think this is the distinction I was trying to get at in the original comment.

“Outside the system” for me means outside the reasoning loop, not necessarily outside the repository.

The repo can still be the source of truth — but a spec, policy file, comment, or agent interpretation of that repo is only a claim about the truth. Those can drift.

What has to remain outside the reasoning loop is the mechanical reconciliation against what actually exists: repo state, runtime behavior, contracts, ownership, and verification paths.

Then every accepted change becomes the new baseline, and the next change is reconciled against that.

Otherwise you’re right — the system is just grading itself.

And if we put AI back into that governance layer, we’ve recreated the same ambiguity one level upstream.

Thread Thread
 
mickyarun profile image
arun rajkumar •

That reframe does most of the work, and it is a better line than the one I was drawing. Outside the reasoning loop rather than outside the repo. The repo can hold the truth. What it cannot hold is the interpretation of the truth.

Where I would stop short of ground truth: what actually exists is still read through an instrument, and the instrument is code someone wrote. Runtime behaviour arrives via a probe. Ownership arrives via whatever walks the tree. Those are claims too, just claims with a much shorter path to the thing. So the property you get is not certainty. It is that drift now has to happen in two places at once and in the same direction. Worth a great deal, and not the same as being outside.

Your last line is the one I would make operational. The test for whether something sits in the governance layer or in the reasoning loop is whether it can fail without anyone exercising judgement. A diagnostic whose output a human has to interpret before it means anything has a reasoning loop in it, and it does not matter whether the reasoning is a model's or ours. That is also the cheap way to decide where an AI is allowed to sit: it can generate the check, it cannot be the thing the check consults.

Thread Thread
 
scarab-systems profile image
Scarab Systems • • Edited

I think the distinction I’d make is that I’m using “diagnostic” differently from a linter or syntax checker.

The process I’m describing isn’t trying to decide whether the code is good, correct, or well-formed at that level. It is looking for specific classes of drift from an established coherence baseline.

Once that baseline is explicit — contracts, ownership, boundaries, responsibilities and expected relationships between parts of the system — some kinds of drift become mechanically observable. The diagnostic can say: this relationship used to satisfy the baseline, the current repo state no longer does, and here is the evidence of the divergence.

That still doesn’t mean the diagnostic gets to interpret the significance of the finding or decide the repair. That part stays with the human or coding agent.

So I agree with your distinction about keeping interpretation out of the governance layer. I’d just separate mechanical drift detection from the reasoning that follows it.

The diagnostic is not the judge of what the system should become. Its job is to produce reliable context about where the system has stopped agreeing with the baseline it was meant to preserve.

Thread Thread
 
mickyarun profile image
arun rajkumar •

The separation holds, and it survives the test I proposed, which I should say plainly rather than quietly. Baseline says this invariant belongs here, repo says it is somewhere else, and that fails without anyone interpreting anything. The interpretation happens afterwards, about what to do. So drift detection sits in the governance layer by my own rule, and I was wrong to file it with the things that do not.

Where I would put the pressure now is the baseline, because it is the one artifact in your design that nothing reconciles.

Contracts, ownership and boundaries are written down by people, and written-down things go stale in a specific way: the system moves for a good reason and the baseline does not move with it. The diagnostic then works perfectly and is wrong about every finding, all in the same direction, and nothing about its output looks any different from the week it was right.

Two ways that ends, both bad. The findings become noise and the team learns to ignore them, which is the ordinary failure. Or somebody repairs the code to match the stale baseline, which is the expensive one, because the drift has now been written into the implementation by a process that reported success.

Our version of this is a rule that says settlement is T+1 after the scheme has moved to same-day. Correct when written, our own file, and the diagnostic will defend it against reality indefinitely.

So the baseline wants the same liveness property you are asking of enforcement. Not whether it is correct, which needs judgement. When each assertion was last reconciled against something outside itself, and how many findings it has produced that were resolved by editing the baseline rather than the code. A baseline that has never lost an argument is not a baseline. It is a preference with a linter attached.

Collapse
 
valentin_monteiro profile image
Valentin Monteiro •

The 20% is defined by consequence, not difficulty, which is exactly why it doesn't shrink as the models get better. You're FCA-authorised, so you live this: the risky code isn't the hard code, it's the code nobody can explain. AI output that works but that no one can defend to an auditor is still a liability, correct or not. So the senior's real job there isn't writing that 20%, it's being able to stand behind it when someone asks why it made the call it did.

Collapse
 
mickyarun profile image
arun rajkumar •

This is the FCA angle that doesn't get enough airtime. "The risky code isn't the hard code, it's the code nobody can explain" — that's exactly it. We've had auditors ask why a specific retry backoff was chosen, and the answer can't be "the AI picked it." Someone has to own the reasoning. AI-generated code that works but has no defensible rationale is a compliance risk in regulated fintech, full stop. The senior's real value isn't writing that 20% — it's being the person who can explain it under questioning.

Collapse
 
valentin_monteiro profile image
Valentin Monteiro •

"The AI picked it" as the answer to an auditor. That image should scare every team shipping AI-generated code in regulated environments. The senior's value isn't the code. It's the defensible rationale attached to it.

Thread Thread
 
mickyarun profile image
arun rajkumar •

You said it better than my whole article did — the senior's value is the defensible rationale, not the code. An auditor won't accept "the AI picked it," and neither should a CTO. The code is cheap now; the why behind it is the thing you're actually paying a senior for. Thanks for reading.

Thread Thread
 
valentin_monteiro profile image
Valentin Monteiro •

Your article framed the problem, I just sharpened one edge. Most people still think the gap is complexity. It's not. It's who signs off on this when a regulator asks why.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Four months late, which is poor going given this is the best line anyone left on the article.

"Who signs off when a regulator asks why" is exactly it. What I've learned since is that the question has a harder second half: sign off on what. A named person accepting responsibility for a decision they can't reconstruct isn't accountability — it's someone agreeing to absorb blame. That's a worse outcome than having no name, because it looks solved.

So the signature is downstream of something much duller: whether the system can produce, months later, what it did and which rule was in force when it did it. Get that and the sign-off means something. Skip it and you've appointed a person to be sorry.

Collapse
 
xulingfeng profile image
xulingfeng •

The 80/20 split is real — and the hard part isn't the 20%, it's knowing which 20% you're in before you ship. We've started routing every AI-generated diff through a cheap local model review gate that flags "suspicious confidence" (clean code that subtly breaks edge cases). Caught 3 leaks and 2 race conditions last sprint alone. Do you run any automated review on the AI-generated parts or just eyeball them?

Collapse
 
mickyarun profile image
arun rajkumar •

We do both. Automated: every PR runs through our standard test suite plus what we call "scar tests" — specific edge cases we've caught before. But we also have architecture lints that check whether the agent used existing shared utilities or reinvented them, and schema validation that catches impossible state transitions at compile time. Manual: any code that touches money movement gets a senior review, non-negotiable. The automated layer catches about 80% of agent mistakes. The senior review catches the 20% that requires judgment about intent, not just correctness.

Collapse
 
xulingfeng profile image
xulingfeng •

scar tests + architecture lints is a solid combo — especially catching when the agent reinvents existing shared utilities instead of reusing them. We tried something similar internally and it worked well. And the non-negotiable senior review for money-touching code is something we've been sticking to as well.

Thread Thread
 
mickyarun profile image
arun rajkumar •

Late reply, sorry.

The reinvented-shared-utility case is the one I'd single out from that list. It isn't a correctness bug — the reimplementation usually works fine — so no test catches it, and it surfaces two quarters later as a maintenance cost when the two copies drift apart. Architecture lints are about the only place it lands.

Sharper by now: catching it at review is already late. The reason the agent reinvents is that it can't see what exists. So the durable fix is making the utility surface discoverable to it, rather than catching the duplicate after it's written.

Collapse
 
dcstolf profile image
Daniel Stolf •

The "scar tests" frame from your reply to @adam_lewis_427616cbc93f0b is the right destination. The underrated part is when in the lifecycle you get there.

Most teams pick up the negative cases reactively: the agent ships the happy path, the senior reviews, finds the missing impossible-transition guard, the test gets added after the fix. That works, but it puts the senior in the role of "the thing that catches what the agent skipped". That doesn't scale, and burns the most expensive person on the team on deterministic checks.

The shift that's worked for us: list the negative cases before any code exists. A spec for "webhook handler" doesn't reach implementation until someone has answered, in writing, what transitions are illegal, what's the behaviour on duplicate delivery, what happens when the bank returns an unknown status.

Each answer becomes a failing test before the agent is prompted. Then the agent has to satisfy them, and the senior reviews the spec (ten minutes, scan-level) instead of hunting omissions in the diff.

The 20% doesn't disappear. It just stops being something a senior discovers after the fact and becomes something the team commits to before the keystrokes happen. Same judgment, applied earlier, where it's cheaper to enforce and harder for the agent to route around.

Scar tests still matter, they're the upgrade path. The first time a negative case bites in prod, it goes into the spec template for that class of feature, and the next webhook handler is born with the guard already required. The institutional memory compounds at the spec layer, not just the test suite.

Collapse
 
mickyarun profile image
arun rajkumar •

This is one of the sharpest comments on this thread. The insight about when in the lifecycle you capture the negative cases changes everything. We've been moving towards exactly this — writing the impossible transitions, the idempotency requirements, and the failure modes into a spec before the agent gets prompted. The spec becomes the acceptance criteria. The agent has to satisfy it. The senior reviews ten lines of spec instead of hunting through hundreds of lines of diff. You're right that this doesn't make the 20% disappear — it makes it cheaper to enforce.

Collapse
 
nark3d profile image
Adam Lewis •

Daniel, the ordering is right. I really like the idea of a spec template per type of feature. Writing the acceptance criteria up front means the agent has something to check itself against, and a senior can review the spec instead of looking for what's missing in the diff. The other thing is that the spec ends up being what the agent has to satisfy. If it's in the repo the agent reads it each time, and the same thing stops getting missed.

prickles.org/tenet/spec-first-exec...

Collapse
 
mickyarun profile image
arun rajkumar •

Very late to this one.

"If it's in the repo the agent reads it each time" is the part that survived contact for me. And the spec being the reviewable artifact instead of the diff is the real win — a senior reading a spec is doing the job they're good at, and a senior scanning a diff for what's missing is doing the one they're worst at.

The failure mode I'd flag: the spec drifts from the code and nothing notices, because nothing executes a spec. You end up with a confident stale document the agent reads on every pass. Same problem as a stale test, except a stale test at least goes red.

Collapse
 
harjjotsinghh profile image
Harjot Rana •

The 80/20 split is the most useful frame for this whole debate. Agents crush the well-trodden 80% (CRUD, boilerplate, glue, standard patterns) because that's where training data is dense, and they faceplant on the 20% that needs system-level judgment, novel tradeoffs, and knowing what NOT to build.

The practical consequence people miss: you should route by that split. The 80% genuinely doesn't need your most expensive model or your most senior human - cheap model, light review. The 20% is where you spend both the premium model AND the senior's attention. Treating all code as equally hard is what makes AI coding feel either too expensive or too risky depending on which half you're looking at. Really well-argued piece - the "why we still need seniors" conclusion is the honest one.

Collapse
 
mickyarun profile image
arun rajkumar •

This is exactly our approach. We use Sonnet for the routine 80% — API scaffolding, test stubs, boilerplate — and escalate to Opus with full architectural context for anything that touches payment logic. The cost difference is significant but the reliability difference is bigger. The routing isn't just about model choice though — it's about context. The 20% needs structured context files that describe service boundaries, shared schemas, and constraint rules. Without that, even the best model drifts.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.