Originally published at AI Frontier Post
For the past year, shipping an AI agent meant writing the same loop as everyone else: call the model, parse its tool calls, run them, append the results, repeat — then bolt on context management, retries, and recovery when the demo had to become a product. At DevDay 2026, OpenAI productized the loop itself. The Agents API, now in public beta, exposes the managed Codex harness behind the company's own agents as an API: OpenAI runs sessions, orchestration, context compaction, and recovery, while your application supplies the tools and picks the execution environment.
This is the same engine that powers Dots, the always-on agents OpenAI announced on the DevDay stage. And with the Assistants API wound down in August, the Agents API is the stateful path forward on OpenAI's platform. In this tutorial you will build a working agent on it end to end: your first session, the three environment options, tools including MCP, multi-turn sessions, subagents, and cleanup. Every call below is verified against OpenAI's official documentation and the Python SDK itself.
What you'll need
An OpenAI platform account and an application API key with three scopes:
api.agents.readandapi.agents.writefor session operations, plusapi.responses.writefor model inference. Create the key in your OpenAI platform project and export it asOPENAI_API_KEY.Python 3.10 or newer with the OpenAI Python SDK at version 3.13.0 or later. Install or upgrade with
pip install --upgrade openai. The snippets below were verified against SDK 3.22.0, including thebeta.agentsnamespace.A budget. There is no separate Agents API fee during the beta: you pay the selected model's API rates, standard rates for OpenAI tools, and standard container rates for OpenAI-hosted sandboxes.
Eyes open on the caveats. Public beta means the surface can still change — pin your SDK version and watch the changelog. Data residency is US-only, and Zero Data Retention is not supported, even if you bring your own sandbox.
Step 1 — Create a scoped API key
In your OpenAI platform project, create an application API key and grant it exactly three scopes: api.agents.read, api.agents.write, and api.responses.write. Then export it:
export OPENAI_API_KEY="your-api-key-here"
Two things the docs stress that are easy to get wrong. First, keep this key outside the agent's sandbox — the agent runs code, so treat anything inside its environment as reachable by it. Second, every Agents API request needs the OpenAI-Beta: agents=v1 header. The official SDKs add it automatically; if you ever call the API with cURL or another HTTP client, include it explicitly or your requests will fail.
Step 2 — Run your first session
The API is organized around four concepts: the agent (model, instructions, tools, MCP servers), the environment (the sandbox or computer where it works), the session (a durable instance of the agent), and events and items (the live and saved records of what happened). The fastest way to feel this shape is the official quickstart: an agent that writes a tree.py script, runs it, and reports a directory tree.
from openai import OpenAI
with OpenAI() as client:
with client.beta.agents.sessions.create(
agent={
"model": "gpt-6-astra",
"instructions": "Write clean code, run it, and report the actual output.",
},
environment={"type": "openai_hosted"},
input="Create tree.py, a Python script that prints a readable tree "
"of the files in the current directory. Run it and show me the output.",
stream=True,
) as events:
for event in events:
print(event.to_json(indent=None), flush=True)
Save it as quickstart.py and run it with python quickstart.py. That single call creates a session, provisions an OpenAI-hosted sandbox, starts a turn of work, and streams progress as JSON events. On a successful run, the agent creates tree.py inside the sandbox, executes it, and reports a directory tree containing that file.
Learn to read the stream — it is your primary debugging surface:
agent.session.turn.completedmeans the turn finished. Heed the docs' warning: a completed turn does not guarantee every tool call succeeded — always check the agent's reported result.Events ending in
turn.failed,turn.cancelled, orsession.failedmean failure or cancellation.agent.session.idlealone does not mean success.If the stream disconnects early, retrieve the session and its saved items before retrying — the work is durable server-side.
One habit to build now: save the session_id from the events. You will need it for everything in Step 5.
Diagram generated for AI Frontier Post.
Step 3 — Pick the right environment
The environment decides where the agent's commands run and where its files live. There are three options, and this is the biggest architectural decision in the tutorial:
openai_hosted— OpenAI provisions and manages a Linux sandbox for the session. This is the default for anything that runs code or touches files. You can configure it with packages, input files, and network policy, and download artifacts the agent produces.none— no sandbox at all. Use it for agents that answer questions or call external tools without running commands or working with local files:environment={"type": "none"}.self_hosted— the agent runs on your infrastructure. Your application runscodex exec-serverinside its own compute, passing it the returnedsession.environment.idandsession.environment.remote_url. The executor authenticates with a separate, restrictedOPENAI_EXECUTOR_API_KEYcarrying theapi.agents.environments.connectscope, and its IP restrictions must allow requests from your sandbox's outbound network.
session = client.beta.agents.sessions.create(
agent={"model": "gpt-6-astra"},
environment={"type": "self_hosted", "workspace_directory": "/workspace"},
)
Diagram generated for AI Frontier Post.
Treat the environment as a security, cost, network, artifact, and data-governance decision — not just a config flag. The hosted sandbox is the fastest way to start; self-hosted is the answer when the agent must touch your VPC, your data, or your toolchain.
Step 4 — Give the agent tools
Tools are declared on the agent, alongside the model and instructions. The API supports custom functions, built-in tools like web_search and programmatic_tool_calling, and MCP servers. This example, adapted from the official docs, builds a research assistant wired to OpenAI's own documentation MCP server plus web search:
session = client.beta.agents.sessions.create(
agent={
"model": "gpt-6-astra",
"instructions": (
"Use the OpenAI documentation MCP and web search to answer "
"technical questions accurately. Delegate independent research "
"tasks to subagents when useful."
),
"tools": [
{"type": "programmatic_tool_calling"},
{
"type": "mcp",
"server_label": "openai_docs",
"transport": {
"type": "http",
"server_url": "https://developers.openai.com/mcp",
},
},
{"type": "web_search"},
],
"multi_agent": {"enabled": True, "max_concurrent_subagents": 4},
},
environment={"type": "none"},
input="Research how to connect an MCP server to an OpenAI agent, "
"check for recent updates, and summarize the recommended setup.",
)
A few things worth noticing. The MCP tool needs a server_label and an HTTP transport with the server URL — that is the entire wiring. environment={"type": "none"} is deliberate here: this agent reads the web; it does not need a sandbox. And multi_agent opts into subagents, which Step 6 covers.
Step 5 — Think in sessions, not requests
The session is the durable unit: agent configuration, conversation, and saved work persist across turns, so you never rebuild conversation context by hand. To keep working, stream new input into the existing session with its ID:
with client.beta.agents.sessions.stream(
session_id, # saved from Step 2
input="Add a maximum-depth option to tree.py, run it, and show me the output.",
) as events:
for event in events:
print(event.to_json(indent=None), flush=True)
Open the event stream before sending follow-up input so you don't miss early events. Because the harness manages context compaction — summarizing previous work as sessions approach context limits — long sessions keep working instead of dying at the context edge, and a session can be resumed where it left off.
Everything the session did is inspectable after the fact:
turns = client.beta.agents.sessions.turns.list(session_id)
items = client.beta.agents.sessions.items.list(session_id)
subs = client.beta.agents.sessions.subagents.list(session_id)
files = client.beta.agents.sessions.artifacts.list(session_id)
Turns carry per-turn status and usage; items are the saved outputs (messages, reasoning, function calls, tool outputs); artifacts are the files the agent produced. When the work is done, download anything you need first — then delete the session, since it persists until you do:
client.beta.agents.sessions.delete(session_id)
Step 6 — Fan out with subagents
When work splits into independent chunks — researching three APIs, triaging three alerts — one agent doing them in sequence wastes the parallelism the harness offers. Setting multi_agent.enabled lets the main agent break work into subtasks and delegate to subagents, each with its own context:
session = client.beta.agents.sessions.create(
agent={
"model": "gpt-6-astra",
"instructions": "Research the topic and delegate independent subtasks.",
"multi_agent": {"enabled": True, "max_concurrent_subagents": 4},
},
environment={"type": "none"},
input="Compare the release notes of the last three OpenAI model launches.",
)
The official "compare release notes with subagents" example is the canonical pattern: each subagent investigates independently, then the main agent combines the findings into one answer. Two rules of thumb: only delegate work that is truly independent, and coordinate around shared state carefully — give each subagent its own working directory or output file so their writes don't collide. Four concurrent subagents is a sensible ceiling, and it happens to be the value in the docs' own example. Inspect subagent activity anytime with sessions.subagents.list(session_id) from Step 5.
Which approach should you use?
Four ways to ship an agent on OpenAI's platform, and when each wins:
ApproachBest whenWhat you give up
Agents API (this tutorial)You want long-running, tool-using agents without building the loop, the session store, compaction, or recovery yourself.Some control over orchestration internals; beta stability; US-only data residency.
Responses API + your own loopYou already have orchestration, or you need full control over every retry, compaction, and tool-execution decision.Development speed — you own every edge case the managed harness handles.
Agents SDK (openai-agents package)You want a lightweight, in-process agent with handoffs and guardrails inside your own application.Managed sessions, hosted sandboxes, and server-side recovery.
Dots (the product)You are the end user and want an always-on assistant rather than a developer surface.Everything programmatic — it is a product, not an API.
Default to the Agents API when the task is long-running and tool-heavy: incident response, research, data analysis, repo work. Drop to the Responses API when the harness's opinions chafe. The SDK sits between the two for single-process agents.
Costs, limits, and guardrails
Model choice is your biggest cost lever. The quickstart uses
gpt-6-astra; for price-sensitive workloads, GPT-6.1 Sol — announced at DevDay at $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens — is positioned as near-Astra intelligence at roughly a fifth of the price. Swap themodelfield and re-run; nothing else changes.Watch the metered extras. Model tokens are only part of the bill: OpenAI tools bill at their standard rates and hosted sandboxes at standard container rates. Heavy tool use in long sessions adds up — the per-turn
usageon the session is where you audit it.Design for the beta. Pin
openaiin your requirements and re-read the changelog before upgrading. Keep secrets out of the sandbox, scope keys to the three permissions from Step 1, and remember the US-only residency and no-ZDR constraints if you handle regulated data.
The takeaway
The Agents API collapses months of agent-infrastructure work — the loop, the session store, compaction, recovery, sandbox provisioning — into sessions.create. What remains yours, and what was always the hard part: the task contract, the tools, the environment choice, and the judgment about when the agent is done. Start with the hosted sandbox and the quickstart above, graduate to none or self_hosted when your security posture demands it, and let sessions — not requests — be the unit you build around.
Sources
OpenAI — Agents API quickstart (official API docs)
OpenAI — Agents API overview (official API docs)
openai/openai-cookbook — Agents API examples (official examples)


Top comments (0)