DEV Community

MishabuildingAI
MishabuildingAI

Posted on Originally published at gobare.dev Fully Autonomous

What is an agent runtime? Six questions that tell you if you need one

An agent runtime is the layer that keeps an agent's work alive between turns: the session's state and event history, when the machine sleeps and wakes, pauses for approval, and the credentials the agent must never read.

That's the short version. The longer one matters because the four words people use for this stack (framework, harness, runtime, sandbox) mean different things depending on who is writing. LangChain, which coined a lot of this vocabulary, says in its own post that there is no clear definition of framework versus runtime versus harness.

Disclosure: I work on Gobare, a hosted agent runtime. We run the three layers as separate processes, so the lines below are ones we had to draw in code.

Four words, four jobs

Layer What it is What it owns Examples
Framework A library you write an agent with Your control flow LangChain, CrewAI, OpenAI Agents SDK
Harness A finished agent loop around a model Prompting, tool calls, context within a turn pi, Claude Agent SDK, Codex
Runtime The service that runs that loop over time State, events, lifecycle, waits, credentials Bedrock AgentCore Runtime, Google Agent Runtime, Gobare
Sandbox An isolated machine Where commands, files and browsers execute E2B, Daytona, Modal

Two things make this messier in practice. OpenAI's Agents API and Claude Managed Agents bundle all of it behind one API, so comparing them with a sandbox provider goes nowhere. And some roundups file sandboxes under "runtimes", which is true only in the sense that code runs there.

One test that separates them

Ask what is still there when the agent's process ends: the turn finished, the machine was reclaimed, or something crashed.

  • A framework leaves nothing. It was a library in your process.
  • A harness leaves whatever it wrote to disk.
  • A sandbox leaves its filesystem until it is destroyed.
  • A runtime is the thing that decides the answer.

The hyperscalers answer it differently, and say so. AWS's AgentCore Runtime gives each session a microVM, stops it after 15 minutes idle by default or 8 hours of compute, and documents session state as ephemeral, with AgentCore Memory as the place for anything durable. Google's Agent Runtime, formerly Vertex AI Agent Engine, supports long-running operations of up to seven days. Both are coherent. The mistake is picking one without reading its answer.

What a runtime has to do

If you build one yourself, this is the list you end up writing:

  1. Keep the session. Transcript and events stored outside the machine, replayable from the last event a client saw.
  2. Sleep and wake the machine. And tell "idle" from "running a long build", which is where naive timers break.
  3. Wait for people and for your code. A function only your backend can run, an approval, a question. Durable, so a restart doesn't lose the question.
  4. Hold secrets the agent can't read. The model key should never enter the sandbox, which means something outside it makes the model call.
  5. Notify your systems. Signed webhooks, retried, delivered at least once, so your backend doesn't hold a connection open for an hour.
  6. Enforce limits. Concurrency, rate limits, a maximum lifetime per machine.
  7. Let results outlive the machine. OpenAI's hosted sandboxes and Gobare both publish files under /workspace/outputs as artifacts, which suggests this is settling into a shared convention.

Do you need one? Six questions

Count how many are true for your agent:

  • [ ] A task takes longer than one HTTP request will wait.
  • [ ] It has to stop and wait for a person or your system.
  • [ ] Many sessions run at once, for many users.
  • [ ] A crash or a deploy must not lose the work.
  • [ ] It holds credentials the agent itself must not read.
  • [ ] Someone other than the person who started it reads the result later.

Zero or one: a model and a harness are enough. A chatbot that answers in one request doesn't need a runtime. Two or more: you are building one whether you call it that or not. It shows up as a queue, a state table, a reaper for idle machines, a webhook sender and a secret boundary.

Build it or use one

Build on a durable execution engine (Temporal, Inngest), a sandbox provider and a harness. You get every choice and own every failure mode. It is the right call when the runtime is your product.

Use one, and pick by what it binds you to:

  • AgentCore Runtime and Google Agent Runtime bind you to a cloud you may already be on.
  • OpenAI's Agents API and Claude Managed Agents bind you to a model provider.
  • Gobare runs the open-source pi harness on any model you connect with your own key. It runs on a single sandbox provider and you can't bring your own compute, and a sandbox is reclaimed two hours after it was created, paused time included. Those are real limits to weigh.

Which of the six questions was the one that pushed you to build a runtime, or away from it?

The full version, with a layer diagram and an FAQ on AgentCore and Google's Agent Runtime, is on gobare.dev.

Top comments (1)

Collapse
 
bhavin-allinonetools profile image
Bhavin Sheth •

Really clear way to explain the runtime layer. The six-question checklist is especially useful because it turns an abstract architecture concept into something you can actually apply to a project. The distinction between project state, lifecycle, and sandbox execution is also easy to miss.