We've reached an interesting point in AI development.
A coding agent can now do much more than generate a code snippet.
It can:
Read repositories
Modify multiple files
Execute shell commands
Install dependencies
Run tests
Call APIs
Interact with browsers
Create pull requests
Potentially access cloud resources
So the architecture of an AI application is changing.
Old approach
User
↓
LLM
↓
Response
Agentic approach
┌── Tools
├── APIs
├── Filesystem
User → Agent ─┼── Database
├── Browser
└── Shell
↓
Runtime
↓
Production Systems
And this introduces a completely different class of engineering problems.
What happens if the agent:
❌ Reads a secret it shouldn't access?
❌ Executes a dangerous command?
❌ Gets manipulated by prompt injection?
❌ Modifies the wrong repository?
❌ Uses an API with excessive permissions?
❌ Deploys something without approval?
The solution can't simply be:
“Let's make the model smarter.”
The model should not be the final security boundary.
We need infrastructure around the agent.
The agent runtime needs:
- Sandboxing
Limit filesystem, network and process access.
- Permissions
Give agents only the capabilities required for the task.
- Identity
Every agent action should have an identifiable actor and scope.
- Credential isolation
Agents shouldn't receive unrestricted API keys or cloud credentials.
- Policy enforcement
Define what an agent can and cannot do.
- Observability
Log tool calls, decisions, failures and resource usage.
- Human approval
Require approval before sensitive operations.
- Kill switches
Be able to immediately stop an agent.
This is why recent work around agent runtimes and safety infrastructure is so interesting.
We're moving from:
LLM Engineering
toward:
Agent Systems Engineering
And the architecture starts looking more like:
┌──────────────┐
│ Model │
└──────┬───────┘
↓
┌──────────────┐
│ Agent │
└──────┬───────┘
↓
┌────────────────────────┐
│ Agent Runtime │
│ │
│ Permissions │
│ Sandbox │
│ Credentials │
│ Policies │
│ Observability │
└───────────┬────────────┘
↓
Tools / APIs / Systems
The interesting part?
The model may be probabilistic.
The environment around it shouldn't be.
As agents become more autonomous, I think the next major engineering challenge isn't simply making them more capable.
It's making them capable, constrained, observable and trustworthy.
That's where AI engineering starts looking a lot like traditional distributed-systems and security engineering.
Would you build an AI agent with direct production access, or always put a runtime/security layer between the agent and production?
Top comments (0)