Claude Code, Codex, and other agent runtimes are becoming capable enough that we increasingly don't need to build the agent loop ourselves.
Yet in client work, the engineering hasn't disappeared. It has moved toward MCP servers, APIs, DSLs, domain tools, state management, authorization, execution guarantees, and the systems agents operate.
This made me rethink what we are actually building.
I find it useful to separate two responsibilities: Agent Runtime Engineering and what I'll call Domain Runtime Engineering.
The Agent Runtime manages how the agent pursues a goal.
The Domain Runtime gives its operations precise meaning and enforceable guarantees.
"Domain Runtime" is not a standardized industry term here.
I'm using it as a convenient name for the application services, workflow engines, compilers, validators, rules, calculations, and transaction logic that execute operations according to domain semantics.
Two runtimes, two responsibilities
An Agent Runtime enables an AI agent to continue working toward a goal.
It manages model calls, context, tool selection, observations, execution progress, interruption, resumption, and sometimes runtime-level approvals.
A Domain Runtime executes operations according to domain semantics.
It determines whether inputs are valid, how specialized calculations behave, which invariants apply, and—where applicable—whether an operation may produce an authoritative state change.
| Agent Runtime | Domain Runtime | |
|---|---|---|
| Central question | What should I do next to achieve the goal? | What does this operation mean, and how should it execute? |
| Typical state | Conversation, plans, observations, run progress | Domain data, versions, approvals, operation results |
| Success | The requested task is satisfied | Domain postconditions hold |
| Failure handling | Replan, investigate, try another path | Validation, rollback, idempotency, compensation, reconciliation |
| Changes when | Models, prompts, tools, harnesses change | Domain rules, contracts, algorithms, operating requirements change |
I would not draw this boundary as probabilistic versus deterministic.
Agent runtimes contain deterministic control logic, and domain processing may itself use AI. The stronger distinction is who owns domain semantics, invariants, and authoritative outcomes.
Domain Runtime does not only mean state changes
Consider a facility analytics system.
The user asks:
Compare occupancy rates across these buildings and investigate the outliers.
The Agent Runtime decides what data to inspect, which analysis to run, whether the result is sufficient, and what additional question to investigate.
The Domain Runtime resolves occupancy_rate against an accepted domain definition and executes it according to that definition.
That definition might ultimately evaluate something like:
occupied_area / leasable_area
but the formula alone is not the domain meaning.
The definition may also specify aggregation granularity, missing-value behavior, unit conversion, eligibility rules, and what happens when leasable_area is zero.
No database mutation is required.
The Domain Runtime's job is to produce a valid result according to accepted domain semantics. The Agent Runtime uses that result to decide what to do next.
Agent Runtime
│
│ "calculate occupancy by building"
▼
Domain Runtime
│
│ resolves and applies
│ versioned domain semantics
▼
Valid domain result
│
▼
Agent Runtime
│
└─ decides what to investigate next
Those domain definitions do not necessarily have to be handwritten by developers.
An agent may propose a new metric or modify an existing definition, but that proposal does not automatically make it authoritative. Validation, review, activation, and execution remain governed application responsibilities.
The runtime can validate and execute an explicit domain definition.
That alone does not establish that the definition correctly represents the real business.
This distinction matters because an LLM can calculate a ratio itself.
What makes the calculation part of the application is not arithmetic. It is the contract around what the calculation means.
State changes make the boundary even clearer
Now suppose the user says:
Analyze this dataset and publish the result.
The Agent Runtime can inspect the data, run analysis tools, identify problems, revise its approach, and eventually determine that it is ready to publish.
But publish is not merely another reasoning step.
The domain may need to determine which version is being published, whether it is valid, whether the caller is authorized, whether business approval exists, and whether the same version has already been published.
User / Product
│
▼
┌────────────────────┐
│ Agent Runtime │
│ │
│ goal │
│ context │
│ planning │
│ tool selection │
│ work continuation │
└─────────┬──────────┘
│
│ MCP / API / CLI
▼
┌────────────────────┐
│ Domain Runtime │
│ │
│ domain semantics │
│ authorization │
│ validation │
│ business approval │
│ state transition │
│ transaction logic │
└─────────┬──────────┘
│
▼
Authoritative State
This is a logical responsibility boundary.
Both runtimes may live in one process. The Domain Runtime may simply be the backend application that existed years before anyone added an agent.
The publish succeeded, but the response disappeared
This is where the distinction becomes operationally important.
Suppose the Agent Runtime requests publication of version 42. The Domain Runtime validates the request and commits the publication successfully.
Then the connection fails before the response reaches the agent.
Agent Runtime
│
│ publish(version=42)
▼
Domain Runtime
│
│ commit succeeds
▼
State: PUBLISHED
X response lost
From the Agent Runtime's perspective, the outcome is unknown.
Retrying is a reasonable agent behavior, but safe retry is a contract between both sides. The domain defines the retry guarantees; the agent integration must preserve operation identity and follow that contract.
A robust contract might assign a stable operation identity to the approved publication request.
1. User approves publish intent P
2. Operation O is durably bound to payload P
3. Domain Runtime executes O
4. Domain state is committed
5. Response is lost
6. Agent retries or queries operation O
7. Domain Runtime recognizes O
8. Existing outcome is returned
Retries preserve both the operation identity and the business payload.
The domain system durably records the outcome and detects duplicate execution.
External effects may complete independently.
Publishing a record and sending a notification, for example, can fail at different points. The system therefore needs to track which effects completed and which still require delivery or reconciliation.
Not every recovery path has to transparently return the original success result.
For some operations, preventing duplicate execution, querying durable state, and marking an unresolved attempt as interrupted may be the safer contract.
This is not something conversation history can resolve.
The agent's memory of the world is not the world.
Runtime approval is not business approval
Approval has a similar ambiguity.
An Agent Runtime may ask:
Allow this shell command?
That controls whether the agent may perform an action in its execution environment.
The business application may separately require:
Approve version 42 for publication to the production portal?
That approval has domain meaning.
Agent Runtime
│
├─ Runtime approval
│ "May the agent attempt this action?"
│
▼
Domain Runtime
│
└─ Domain authorization / approval
"May this domain operation take effect?"
The first approval cannot silently imply the second.
Business policy needs an authoritative owner and consistent enforcement across UI, API, scheduled jobs, and agent paths. Runtime approvals complement domain authorization; they do not replace it.
Agent Runtime and Domain Runtime need a contract
As existing Agent Runtimes improve, this boundary becomes one of the most important custom engineering surfaces.
Agent Runtime
│
│ chooses and requests operations
▼
┌─────────────────────────┐
│ Agent-facing Interface │
│ │
│ Tools │
│ MCP │
│ API │
│ CLI │
│ Schema / DSL │
└────────────┬────────────┘
│
▼
Domain Runtime
A backend interface designed for a human-facing application is not automatically a good agent interface.
A frontend can orchestrate many low-level endpoints because application code already knows the workflow. Giving those endpoints directly to an agent pushes that orchestration back into model reasoning.
For example, this:
create_revision(...)
update_rows(...)
set_status(...)
create_permission(...)
send_notification(...)
may be a worse agent surface than:
prepare_import_plan(...)
validate_import_plan(...)
publish_version(...)
The deeper question is not only which protocol to use.
What is the smallest, safest, and most legible operation we can expose to the agent?
MCP is one way to expose that boundary
MCP fits naturally into this architecture.
It can expose domain capabilities to an Agent Runtime through a standard tool interface.
Agent Runtime
│
│ MCP
▼
Domain Tool
│
▼
Domain Runtime
But MCP does not determine the business meaning behind the operation.
publish_report could be exposed through MCP, REST, an internal function, a dynamic tool, or a CLI. The same Domain Runtime should still enforce the invariants that must hold regardless of the caller.
MCP exposes a capability. The Domain Runtime determines its domain effect.
Sometimes the MCP server and Domain Runtime are implemented together.
Sometimes MCP is only an adapter around an existing service. In other systems, the agent runtime may expose its own dynamic tool mechanism instead.
The logical responsibility matters more than the transport.
DSLs are another way to design the boundary
Some domains benefit from something more expressive than individual tool calls.
Suppose an analysis can be represented as:
analysis Occupancy {
source facilities
metric occupancy_rate
group_by building
}
Here, occupancy_rate refers to a versioned domain definition rather than redefining the metric inside every agent-authored program.
The Agent Runtime can read this program, propose one, modify one, or explain one. That does not mean the Agent Runtime determines what the referenced metric means.
Agent Runtime
│
│ proposes DSL
▼
Domain Program
│
▼
Compiler / Validator
│
▼
Domain Runtime
│
▼
Valid result / state
A compiler may translate it. A validator may reject it. A workflow engine may evaluate it. Domain services may execute the resulting operations.
The agent can propose a domain program. The Domain Runtime determines whether and how it takes effect.
An agent can also propose a new metric definition.
That is a different operation from using an existing metric. The proposed definition must be validated and separately accepted before other programs can rely on it as an authoritative domain definition.
The same distinction applies to mandatory policy.
An agent might generate:
workflow PublishAnalysis {
analyze
review
publish
}
but a mandatory approval requirement cannot depend on whether the agent remembered to encode:
publish requires approval
If approval is a domain invariant, deleting such a line cannot remove the requirement.
Agent-authored programs cannot weaken mandatory domain policies. Policy changes are separately authorized and validated.
The Domain Runtime therefore evaluates the proposal against constraints the proposal itself does not control.
Agent-authored program
│
│ proposed intent
▼
Domain Runtime
│
├─ resolves domain definitions
├─ validates program
├─ applies mandatory policy
└─ executes if valid
Not every system needs a DSL.
Typed functions, JSON Schema, or an ordinary API may already express the domain clearly enough. The important property is the separation between proposing an operation and determining its domain meaning and effect.
Runtime control is a different boundary
There is also an interface on the other side of the Agent Runtime.
A product may need to start work, resume it, send new input, receive events, or participate in runtime-level approvals.
Product UI / IDE
│
│ runtime control
▼
Agent Runtime
│
│ capability access
▼
Domain Runtime
Codex App Server is one concrete example of a control interface for integrating Codex into another product.
It exposes an interface for rich clients to interact with Codex conversations, approvals, and streamed agent events. This is a logical boundary, not a requirement for a separate runtime service.
That is a different concern from MCP.
One interface controls the agent's execution. The other gives the agent access to domain capabilities.
Structured errors become part of the contract
An agent-facing Domain Runtime should make domain failures actionable.
Compare:
400 Invalid configuration
with:
{
"code": "STALE_VERSION",
"expected_version": 43,
"provided_version": 42,
"recoverable": true,
"suggested_action": "reload_and_revalidate"
}
The second response does not tell the agent what its goal should be.
It makes the domain state explicit enough for the Agent Runtime to choose its next action.
The Domain Runtime should not decide what the agent wants to do next.
It should make domain consequences and constraints explicit enough that the agent can decide.
Agent-native does not mean agent-owned
Making a system easier for agents to use does not mean moving business logic into the agent.
Often the opposite architecture is safer.
An agent-facing interface should be optimized for model reasoning:
clear operations
typed inputs
structured outputs
stable identifiers
explicit preconditions
actionable failures
idempotent mutations
queryable operation status
But the meaning behind those operations should remain authoritative regardless of whether the caller is Claude Code, Codex, a custom agent, a human UI, a scheduled job, or another service.
Agent-native interfaces should make the domain easier to reason about, not make the agent the owner of the domain.
Better Agent Runtimes move the build boundary
This changes how I think about build versus buy.
If Claude Code, Codex, an Agents SDK, or another runtime already gives us a capable agent loop, context management, tool execution, resumption, and sandbox integration, rebuilding those pieces may create little value.
Instead, I increasingly ask:
- Are the domain operations well defined?
- Can the agent invoke them at the right abstraction level?
- Are domain definitions versioned and governed?
- Are mandatory policies independent of agent-authored programs?
- Are authorization and business approvals authoritative?
- Are operations safely retryable?
- Can uncertain outcomes be reconciled?
- Are errors structured enough for the agent to recover?
- Can agents and conventional application paths share the same semantics?
That is where much of our custom engineering now lives.
Engineering the boundary
This started for me as a naming problem.
We were building MCP servers, DSLs, APIs, CLI tools, domain services, sandboxes, databases, and agent integrations, and none of those labels described the whole job.
The more useful realization is that many of these components exist to engineer one boundary:
User / Product
│
▼
Agent Runtime
goal / context / next action
│
MCP / API / CLI / DSL
▼
Domain Runtime
semantics / guarantees / results
│
▼
Data / External Systems
The Agent Runtime should be free to explore, replan, retry, and choose different paths toward a goal.
The Domain Runtime should make domain operations precise, enforce what must be enforced, and produce results the rest of the system can trust.
Existing Agent Runtimes will keep getting better, reducing how much harness infrastructure we need to build ourselves.
But someone still has to define what occupancy_rate and publish mean, which guarantees apply, and which invariants can never be delegated to model reasoning.
The Agent Runtime manages how the agent pursues a goal.
The Domain Runtime gives its operations precise meaning and enforceable guarantees.
For me, the interesting engineering is increasingly the boundary between the two.
References
The Agent Runtime / Domain Runtime distinction in this article is my own responsibility model. The following official documentation is referenced for the products and protocols discussed above.
Top comments (0)