AI is making code cheaper to produce. That doesn't make architecture less important. It makes good architecture more valuable.
A few years ago, the limiting factor for many engineering teams was simple:
How quickly can we write the code?
Today, that question is changing.
AI coding tools can already generate features, write tests, fix bugs, refactor code, update documentation and work through GitHub pull requests with increasing autonomy. GitHub's coding agent, for example, can take an issue, work in an isolated development environment, modify the code, run tests and open a pull request for human review.
Tools such as Claude Code are also moving beyond autocomplete toward agents that can navigate a codebase, edit multiple files, execute commands, test and debug software.
So here's the question I think engineering leaders should be asking:
If producing code is no longer the primary bottleneck, what becomes the bottleneck?
My answer:
Context, architecture, verification and decision-making.
And that changes how we should design software teams.
1. The Old Software Engineering Loop
The traditional development loop looks roughly like this:
Requirement
↓
Engineer understands requirement
↓
Design
↓
Write code
↓
Write tests
↓
Code review
↓
CI/CD
↓
Production
A significant amount of engineering time is spent on implementation.
Now introduce an AI coding agent:
Requirement
↓
Engineer + AI
↓
Design
↓
AI implements
↓
AI writes tests
↓
AI runs validation
↓
Engineer reviews
↓
CI/CD
↓
Production
The implementation step becomes dramatically cheaper.
But notice something interesting.
The human still needs to answer:
- What are we actually building?
- What should the architecture look like?
- What constraints matter?
- What should the agent not touch?
- Is this API boundary correct?
- Is this database model correct?
- Is this security model acceptable?
- Does this code belong here?
- What happens when this fails?
- What happens at 10x scale?
- What happens during partial failure?
The human role is moving up the abstraction stack.
2. AI Doesn't Remove Engineering Work
This is where I think some AI discussions go wrong.
The assumption is:
AI writes code faster
↓
Software engineering becomes easier
The more realistic model is:
AI writes code faster
↓
More code gets produced
↓
More code needs validation
↓
More architectural decisions become important
↓
More system-level risks become possible
DORA's recent research describes AI as an amplifier of an organisation's existing strengths and weaknesses. Its 2026 analysis also found that the time saved during initial code creation is frequently reallocated to auditing and verification.
That's a really important signal.
The bottleneck may simply move.
3. The New Engineering Bottleneck
Imagine an engineer could previously produce:
100 lines/day
and an AI-assisted workflow can produce:
1,000 lines/day
That sounds like a 10x improvement.
But what if your team can only effectively review:
300 lines/day
?
You haven't created a 10x engineering organisation.
You've created a review queue.
AI
│
↓
┌──────────┐
│ Code │
│ generated│
└────┬─────┘
↓
┌───────────────────┐
│ Human verification│
└─────────┬─────────┘
↓
Production
This is one of the most important shifts I see coming:
Code generation is becoming cheaper. Code verification is becoming more valuable.
4. The New Division of Labour
I don't think the future is:
Humans → code
AI → code
It's closer to:
Humans own
Intent
Architecture
Boundaries
Business rules
Security decisions
Trade-offs
Risk
Quality bar
Production accountability
AI agents own more of
Implementation
Boilerplate
Tests
Refactoring
Documentation
Dependency upgrades
Migration scripts
Code exploration
Bug investigation
This doesn't mean AI should blindly own the second list.
It means those are areas where delegation can often provide high leverage.
The interesting engineering problem becomes:
How much autonomy should we give the agent for each type of change?
5. Not All Code Changes Should Have the Same Autonomy
This is one of the practical patterns I would introduce in an AI-native engineering team.
Classify changes by risk.
Level 1 — Low risk
Examples:
Documentation
Unit tests
Formatting
Simple refactoring
Boilerplate
Non-production tooling
AI can potentially operate with high autonomy.
Level 2 — Medium risk
Examples:
Business logic
API changes
Database queries
Performance changes
Dependency upgrades
AI can implement.
Human reviews carefully.
Level 3 — High risk
Examples:
Authentication
Authorisation
Payment logic
Data migrations
Encryption
Infrastructure
Production configuration
Security controls
AI can assist heavily.
But humans should retain strong approval and verification controls.
The workflow becomes:
AI autonomy
↑
│
Low risk ────────────┤ High autonomy
│
Medium risk ─────────┤ Human review
│
High risk ───────────┤ Strong human control
│
↓
Production
The goal isn't:
"Let AI do everything."
It's:
"Give AI as much autonomy as the risk model allows."
6. Architecture Becomes the Agent's Constraint System
Here's an interesting change.
Humans have always relied on architecture to constrain developers.
For example:
Controller
↓
Service
↓
Repository
↓
Database
Now that architecture also constrains AI agents.
This becomes extremely important.
If an agent sees:
services/
├── payments/
├── orders/
├── customers/
└── shared/
but there are no clear boundaries, it may make perfectly valid code changes that are architecturally wrong.
For example:
payments-service
↓
directly queries
↓
customers database
The code might compile.
Tests might pass.
The architecture is still wrong.
So architecture needs to become increasingly machine-readable.
7. Make Architecture Explicit
One practical pattern is to put architectural context directly into the repository.
For example:
repository/
│
├── README.md
├── ARCHITECTURE.md
├── ADR/
│ ├── 001-service-boundaries.md
│ ├── 002-database-choice.md
│ └── 003-event-driven-integration.md
│
├── services/
│
├── libraries/
│
├── tests/
│
└── docs/
Now the agent has something better than:
"Figure out how this system works."
It has explicit context.
For example:
Payments service
Owns:
- Payment lifecycle
- Payment state
- Payment reconciliation
Does not own:
- Customer identity
- Product catalogue
Integration:
- Customer information via Customer API
- Events published through Event Bus
Never:
- Access another service's database directly
This isn't documentation for humans only anymore.
It becomes part of the agent's operating context.
8. Your README Is Becoming Part of the Development Interface
We've historically treated documentation as something developers read.
With coding agents, documentation increasingly becomes something agents consume before making changes.
This changes what good documentation looks like.
Instead of:
This service handles payments.
we want:
# Payment Service
## Responsibilities
Owns payment lifecycle and reconciliation.
## Boundaries
DO:
- Create payment
- Authorise payment
- Capture payment
- Refund payment
DO NOT:
- Read customer database
- Modify order state directly
- Store authentication credentials
## Dependencies
Customer information:
Customer API
Order information:
Order API
## Testing
Run:
./gradlew test
Integration tests:
./gradlew integrationTest
This is useful to humans.
It is also much more useful to agents.
9. The Codebase Needs a "Constitution"
This is one of the patterns I think will become increasingly common.
Create a small set of explicit engineering rules.
For example:
ENGINEERING_RULES.md
Containing things like:
1. Services must not access another service's database.
2. Public APIs require backward compatibility.
3. Database migrations must be backward compatible.
4. All external calls require timeouts.
5. Retries must use bounded exponential backoff.
6. New APIs require contract tests.
7. Authentication logic must not be implemented inside
individual business services.
8. Production infrastructure changes require human approval.
Now an AI agent doesn't need to infer all of these rules from thousands of lines of code.
You've made the constraints explicit.
The industry is already moving toward this idea of giving agents structured procedural and organisational context. Anthropic, for example, describes "Agent Skills" as a way to package instructions, scripts and resources so agents can perform specialised work more reliably.
10. Tests Become More Important Than Ever
There's another major shift.
Previously:
Tests protect us from developer mistakes.
In an AI-native workflow:
Tests become executable specifications for agents.
Consider:
Requirement:
A payment must never be captured twice.
A human can understand that.
But a much stronger contract is:
@Test
void shouldNotCapturePaymentTwice() {
...
}
Now the agent has a deterministic boundary.
This creates a powerful loop:
Requirement
↓
Architecture
↓
Tests
↓
AI implementation
↓
Run tests
↓
Fix
↓
Run tests again
The better your tests are, the more autonomy you can safely give the agent.
This is one reason coding agents tend to work better in well-tested codebases. GitHub explicitly highlights well-tested repositories as a strong environment for coding-agent tasks.
11. CI/CD Becomes an AI Safety Boundary
In traditional development:
Developer
↓
PR
↓
CI
↓
Review
↓
Merge
With agents:
Human
↓
Task
↓
AI agent
↓
Code
↓
Tests
↓
Security checks
↓
Policy checks
↓
PR
↓
Human approval
CI is no longer just:
"Does the code compile?"
It becomes an executable governance layer.
For example:
PR
│
├── Compile
├── Unit tests
├── Integration tests
├── Contract tests
├── Static analysis
├── Dependency scan
├── Secret scan
├── Security tests
├── Architecture rules
└── Policy checks
This is how you allow agents to move faster without removing engineering controls.
12. Don't Give Agents Production Access by Default
This sounds obvious.
But agentic workflows change the threat model.
An agent that can:
read code
write code
run commands
access network
modify infrastructure
deploy software
has a very different risk profile from autocomplete.
Modern agent tooling is therefore increasingly introducing sandboxing, scoped credentials and filesystem/network boundaries. Anthropic, for example, describes sandboxing Claude Code with filesystem and network isolation as part of making autonomous coding safer.
A practical model is:
Developer Agent
│
├── Read repository
├── Modify branch
├── Run tests
└── Build artifacts
│
↓
Human approval
│
↓
Deployment
Not:
AI
↓
Production
13. Small PRs Become Even More Important
Imagine an agent generates:
27 files
4,000 lines changed
3 new dependencies
2 database migrations
Technically impressive.
Practically horrible to review.
The human reviewer is now the bottleneck.
Instead:
Task
↓
Small change
↓
Tests
↓
PR
↓
Review
↓
Next change
Small changes make:
- reviews easier
- failures easier to diagnose
- rollbacks safer
- AI mistakes easier to detect
- context easier to maintain
AI gives us the ability to generate huge changes.
Engineering discipline should prevent us from automatically accepting them.
14. Context Is Becoming an Engineering Resource
Here's another subtle change.
An AI agent is only as effective as the context it can access.
Consider:
Task:
"Add customer notifications."
That's almost useless.
Compare:
Goal:
Send an email when a payment succeeds.
Architecture:
Notifications are owned by notification-service.
Constraints:
- Payment service publishes PaymentCaptured event.
- Notification service consumes the event.
- Payment service must not call the notification service synchronously.
- Email delivery must be idempotent.
Acceptance criteria:
- Duplicate events must not send duplicate emails.
- Failed delivery must retry.
- Retry must be bounded.
The second task is dramatically easier for both humans and agents.
This leads to a new engineering skill:
Context engineering for software development.
Not prompt tricks.
Actual engineering context:
- architecture
- domain vocabulary
- APIs
- constraints
- tests
- examples
- ADRs
- operational knowledge
- security rules
15. The New Staff Engineer Job
This is where I think the role of Staff and Principal Engineers changes significantly.
The question becomes less:
"How do I write this code?"
And more:
"How do I create a system where humans and agents can safely produce good software?"
That means investing in:
Architecture
Clear boundaries.
Developer experience
Easy local setup.
Documentation
Machine-readable context.
Testing
Strong executable specifications.
CI/CD
Automated verification.
Security
Least privilege and isolation.
Observability
Fast feedback from production.
Engineering standards
Explicit constraints.
Platform engineering
Reusable capabilities for every team.
In other words:
Staff engineering becomes increasingly about designing the environment in which software gets produced.
16. The Architecture of an AI-Native Team
Putting this together:
HUMAN
│
│
Intent / Architecture
│
↓
┌─────────────────┐
│ AI Agent(s) │
└────────┬────────┘
│
┌─────────────┼──────────────┐
↓ ↓ ↓
Code Tests Documentation
│ │ │
└─────────────┼──────────────┘
↓
CI / Policy
│
┌─────────────┼─────────────┐
↓ ↓ ↓
Security Architecture Quality
checks checks checks
│ │ │
└─────────────┼─────────────┘
↓
Human Approval
│
↓
Production
│
↓
Observability
│
↓
Feedback
Notice what's missing?
A human manually writing every line.
The human is instead controlling:
intent → architecture → constraints → verification → risk → outcome
17. What Happens to Junior Engineers?
This is an uncomfortable question.
If AI can handle:
- boilerplate
- basic CRUD
- straightforward tests
- simple refactoring
- documentation
- basic debugging
then some traditional entry-level tasks become less valuable as learning mechanisms.
That doesn't mean junior engineers become unnecessary.
It means the learning path has to change.
A junior engineer shouldn't only learn:
How to write CRUD APIs
They should increasingly learn:
Why this API exists
Why this boundary exists
How failure happens
How data flows
How systems behave under load
How to test assumptions
How to review generated code
How to reason about security
How to operate software
The ability to judge code becomes more important than the ability to produce large amounts of code.
18. Measure the Right Things
One of the easiest mistakes is measuring AI adoption by:
Lines of code generated
or:
Number of AI prompts
Those aren't meaningful engineering outcomes.
I'd rather measure:
Lead time for change
Change failure rate
Review time
PR size
Defect rate
Rollback rate
Incident rate
Test coverage
Developer satisfaction
Time from idea → production
And perhaps one of the most interesting metrics:
How much human attention does each production change require?
Because that's potentially the new scarce resource.
19. A Practical AI-Native Development Workflow
If I were introducing agentic development into an engineering team tomorrow, I'd start relatively conservatively.
Step 1 — Give the agent context
README
ARCHITECTURE
ADRs
ENGINEERING_RULES
API contracts
Step 2 — Give it a bounded task
Bad:
"Improve the payment system."
Better:
"Add idempotency handling to POST /payments.
Do not change the API contract.
Use the existing IdempotencyStore.
Add unit and integration tests.
Do not modify database schema."
Step 3 — Let the agent implement
Agent
↓
Code
↓
Tests
Step 4 — Run automated verification
Build
Tests
Security
Architecture
Contracts
Step 5 — Human reviews the important things
Not:
"Did the AI write valid Java?"
But:
"Did it implement the correct behaviour?"
and:
"Did it preserve the architecture?"
Step 6 — Merge small changes
Keep the feedback loop tight.
20. The Biggest Anti-Pattern: AI + Bad Architecture
Here's the scenario I'm most concerned about.
You have:
Poor boundaries
+
Poor tests
+
Poor documentation
+
AI coding agent
What happens?
The AI doesn't magically fix the architecture.
It can actually make the problem worse faster.
Bad architecture
↓
AI generates more code
↓
More coupling
↓
More complexity
↓
Harder reviews
↓
More technical debt
This is exactly why the "AI makes developers 10x" narrative needs context.
AI can amplify good engineering.
It can also amplify bad engineering.
DORA's research makes essentially this broader point: AI tends to amplify the underlying capabilities and weaknesses of the organisation using it.
21. So What Actually Changes?
I don't think the fundamental principles of software engineering disappear.
They become more important.
We still need:
Good boundaries
Good APIs
Good data models
Good tests
Good observability
Good security
Good failure handling
Good operational practices
But now we also need:
Machine-readable architecture
Explicit engineering rules
Agent-friendly documentation
Risk-based autonomy
Automated verification
Strong CI/CD guardrails
Scoped agent permissions
Small change sets
That's what I would call an AI-native engineering environment.
22. My Rule of Thumb
If I had to summarise the whole article in one sentence:
When AI makes implementation cheap, invest more heavily in the things AI cannot safely decide for you: intent, architecture, constraints, verification and accountability.
Or, even more simply:
Before AI:
Code was expensive.
Review was cheaper.
With AI:
Code is cheaper.
Good judgment is expensive.
And that changes where engineering organisations should invest.
Final Thought
I don't think the future of software engineering is:
Humans stop coding and AI writes everything.
I think it looks more like:
Human
↓
Defines intent
↓
Designs boundaries
↓
Sets constraints
↓
AI implements
↓
AI validates
↓
Automated systems verify
↓
Human judges
↓
Production learns
↓
Architecture evolves
The interesting competitive advantage won't simply be who has access to the best coding model.
Most teams will have access to very capable models.
The advantage will increasingly come from:
Who has the best architecture, context, engineering constraints, feedback loops and ability to safely turn AI-generated code into production outcomes.
That's the real architecture of an AI-native software team.
Top comments (0)