Everyone is shipping agents. Far fewer teams have shipped the thing agents need underneath them, which is a system that can be handed work and trusted with the consequences.
Agent frameworks have converged quickly on a set of primitives: tool use via MCP, progressive disclosure via skills, instructions from a repo file, shell execution, patch-based file edits. Those let an agent finish a task in a sandbox. They do not solve the operating problems of a live business.
Here is the unglamorous list of what an execution layer has to provide.
Idempotency, because agents retry
Agents retry. Networks fail. Containers get rescheduled. A workflow that resumes mid-flight will eventually resume in a state where an action may or may not have already happened.
That forces a decision early: every externally visible operation needs a stable key so a repeat is detectable, and the check has to live in the system that performs the action rather than in the agent's memory. An agent saying "I don't think I sent that" is not a safety mechanism.
The newer MCP revisions lean into this. The 2026-07-28 specification dropped the initialization handshake and the session header, so any request can land on any instance behind a round-robin load balancer. That is much easier to scale, and it also means state that used to be implicit in a session now has to be explicit. The spec is direct about the workaround: mint a handle from a tool and have the model thread it back as an argument, rather than hiding state in the transport.
Permissions that survive an audit
"The agent was allowed to" is not an answer a security review accepts. It wants to know which principal acted, under which policy, with what scope, and where the record lives.
Which is why permission decisions belong in the execution layer rather than in the prompt. A prompt instruction is a suggestion to a probabilistic system. A permission check is a decision made by code the model cannot talk its way past.
Rate and blast-radius controls
Any system that can send messages needs limits that hold when the agent has a bad day. Daily caps, per-domain caps, timing windows, exclusion lists, suppression on bounce. None of these should be left to a model's memory.
Durable workflow state
Multi-step work spans hours or days and has to survive restarts, deploys, and expired sandboxes. The interesting question is not how to store state but what the correct resumption semantics are. Re-running a step is usually wrong. Skipping it silently is worse. Somewhere between those two is a policy, and the policy is business-specific.
An audit trail a human can read
The output of an autonomous system has to be reviewable by someone who was not watching. Not a token log. A record of what was decided, what was done, what the result was, and what happened next.
A null result
Most pipelines have no output that means "nothing." Every detected pattern becomes a task and the queue becomes noise. An execution layer needs a first-class way to do nothing, and to explain why.
Failure semantics, including handing back
Agents fail in two directions: doing the wrong thing, and quietly stopping. The second is more dangerous in production because nobody notices. Durable execution needs explicit failure states and an escalation path to a person, which is a product decision as much as an engineering one.
Why this stays specialized
A frontier model can write all of this from scratch, badly, once. Operational infrastructure is not a code-generation problem, it is a maintenance problem: deliverability reputation, data hygiene, permission models, compliance controls, and the accumulated rules of a specific business.
That is the argument for a division of labour rather than a merger. General agents orchestrate. Specialized infrastructure executes and takes responsibility for the outcome.
We build this at SalesRuns: contact records, sending infrastructure, workflow state, permissions, and audit history across email, WhatsApp, Telegram, LINE, Slack, Discord, and WeCom. The full argument for why this is where AI sales is heading is in the main article: The Next Generation of AI Agents Will Not Just Answer Questions. They Will Do the Work..
Originally published on SalesRuns.com.
Top comments (0)