DEV Community

shizhe Lim
shizhe Lim

Posted on Fully Autonomous

Managed Agent Runtimes Still Need a Control Plane

Hosted agent runtimes are making a lot of infrastructure feel like a solved problem.

Long-running sessions, context compaction, tool execution, sandboxes, and subagent coordination can now be consumed as platform features. That is useful — especially for teams that would rather spend their time on a product workflow than on maintaining an agent loop.

But I think we should separate the agent runtime from the agent control plane.

A runtime is where work executes. A control plane is where a team can answer operational questions after something goes wrong.

For a production agent, I would want the control plane to own at least:

  • Execution journal: inputs, tool calls, outputs, and decisions that can be inspected later.
  • Replay and recovery: a failed run should be resumable or reproducible without guessing what happened.
  • Approval gates: side-effecting actions need policy and a clear human-in-the-loop path.
  • Evaluation corpus: production failures should become regression cases.
  • Model-routing policy: different steps may need different models, latency budgets, or cost ceilings.

A hosted runtime can still be the right choice. It may reduce operational burden dramatically. But “managed” should not mean “opaque.”

My rule of thumb: let a platform operate the commodity infrastructure; keep evidence, governance, and decision policy portable enough that your team can investigate and improve the system.

How are you drawing this boundary? Is there a component you would never hand to a hosted agent platform?

Top comments (0)