TL;DR
- An agent harness is the loop around the model: it assembles context, calls the model, runs tools, keeps the session and enforces what the agent may touch.
- OpenAI's own documentation splits the stack three ways, and the split is the useful part: harness, environment, application server. The harness can run without an environment; the Bash tools, workspace files and executor MCPs cannot.
- Own the harness and the model becomes a component you can replace. Let a provider own both and the model becomes the product, with its retirement dates attached.
- Provider timelines explain the urgency better than any argument: OpenAI gives at least six months' notice for generally available models, three months for specialised variants, and as little as two weeks for preview models.
- The work is unglamorous and small: one internal interface for model calls, the model id in configuration, evaluations in the repository, and a swap drill on twenty real tasks.
What a harness is made of
OpenAI's architecture page names the pieces plainly. The harness is the loop: it runs the model and the tools and maintains the session. The environment is where commands run, and it can be an OpenAI-hosted sandbox, a self-hosted executor, or nothing at all. The application server is your code, which submits tasks, receives events and handles function tools.
That separation is worth reading twice, because most teams collapse it into one word: agent. An agent that answers questions from a knowledge base needs a session and tools and little else. An agent that edits files needs an environment, and once it has one, someone owns provisioning, reconnection, shutdown and the files that survive between runs. OpenAI states the boundary directly: with no environment, the built-in Bash and apply-patch tools, the workspace files and the executor MCPs are simply unavailable.
Official OpenAI figure from the Agents API architecture page (developers.openai.com/api/docs/guides/agents-api/architecture, retrieved 2026-10-04): the application talks to the harness, the harness talks to the model and the tools, and the execution environment is a separate decision.
The loop is the asset, and the model is a component
Here is the distinction that decides how much a team controls. The model produces tokens; the harness decides what those tokens can do. Permissions, tool allow-lists, human approval gates, retries, context compaction, logging, cost accounting and the record of what the agent actually did all live in the loop. None of them live inside the model.
Take a concrete case. Imagine a ten-person logistics company in Casablanca whose back-office agent reconciles supplier invoices against purchase orders and flags the mismatches for a human. The valuable engineering is not the model call. It is the retrieval of the right PO, the rule that the agent may read the accounting system and may write only a draft note, the retry when the ERP times out, and the log that proves which invoice, which document and which model produced each flag. That work survives a model change. A prompt does not, and neither does an unrecorded session.
Two ways to own the harness
OpenAI's own guide lays out three runtimes and the trade is explicit. With the Agents API, OpenAI runs the managed Codex harness, session state is saved by the provider, and integration effort is low. With the Agents SDK, the runner executes the loop inside your application, state lives in your storage or in the provider's conversation objects, and effort is medium. With the Responses API, you call models directly and manage history yourself, at the highest integration cost and the most control.
The managed path buys real things: automatic context compaction, multi-agent orchestration, programmatic tool calling and support for MCP servers arrive without you building them. Compaction alone is a project most teams underestimate. The trade is symmetric, though, and it is the one clients ask about: the provider sees and stores the session, the loop's behaviour changes when the provider changes it, and swapping the managed runtime for your own later means rebuilding the parts you outsourced.
Netics editorial comparison built on the runtime table in OpenAI's Agents guide: the managed harness moves session state and maintenance to the provider, while the in-application harness keeps the swap and the failure surface inside your own release cycle.
Netics' position, after wiring agents into client systems, is that the choice is less important than making it deliberately. A managed harness for a support assistant on a twelve-month horizon is a good trade. A managed harness under a workflow that touches payroll or invoicing is a conversation about data residency, audit and exit, and it should be written down before the first tool is connected rather than after the pilot succeeds.
The model layer ages on a published schedule
Model churn is the argument for a seam, and both major providers publish the calendar. OpenAI's deprecation page sets notice periods: at least six months for generally available models, at least three months for specialised variants such as the Codex line, and, in its own words, preview models "may be retired with much shorter notice, such as 2 weeks." The same page states that preview models are not recommended for business-critical production workloads unless you can migrate on short notice. On 2026-10-01 OpenAI announced deprecations including GPT-5.3-Codex, GPT-5.1 and GPT-5.4-Nano, scheduled for removal from the API on 2027-04-01 with six months' notice, and text-to-speech models removed on 2027-01-06.
Anthropic runs a four-state lifecycle — active, legacy, deprecated, retired — and is blunt about the last one: requests to retired models fail. Its table carries dates such as claude-opus-4-1-20250805, deprecated on 2026-06-05 with retirement on 2026-08-05, and on 2026-09-30 it notified developers using Claude Sonnet 4.5 that the model is being retired on the Claude API. The same page contains the detail that matters most for architecture: partner-operated platforms such as Amazon Bedrock and Google Cloud set their own retirement schedules, so the same model can be alive on one route and gone on another.
Netics editorial checklist: the five practices that turn a model dependency into a configuration change, drawn from the deprecation policies OpenAI and Anthropic publish rather than from any vendor's architecture diagram.
Read that alongside the two-week window for preview models and the picture is clear. The application can be a five-year asset; the model id inside it is a dependency with an expiry date that somebody else sets. A team that treats that as an architectural fact, rather than as bad news to handle when the email arrives, spends its maintenance budget in planned swaps instead of emergency migrations.
Tools and the MCP seam
Tool calling is where portability is usually lost. Bind an agent to a vendor's proprietary tool format and the model swap stops being a routing change and becomes a rewrite of every integration.
The Model Context Protocol exists to remove that particular coupling. It is an open standard for connecting AI applications to external systems — data sources, tools and workflows — and its documentation uses the analogy of a USB-C port for AI applications. Its specification version on the page we read is 2026-07-28, and the client list includes Claude, ChatGPT, Visual Studio Code and Cursor, which is the practical argument: a tool server you publish once is reachable from several runtimes. OpenAI's harness, for its part, "can call remote MCP tools directly," while function tools route back through your code, which is exactly where your business rules belong.
What a six-person team can do this quarter
The seam costs less to build than the migration costs to avoid. Put every model call behind one internal interface, so the provider is chosen in one module. Keep the model id in configuration and log it next to every output, so a quality regression is traceable to a swap rather than debated. Keep twenty to thirty real tasks with known answers in the repository, and run them against the current model and the candidate before anything changes in production. Then run one swap on purpose, in staging, and write down what broke.
That is the shape of the platform Netics builds when a client wants agents on infrastructure they control: the loop, the tool servers, the evaluation set and the permissions live in the client's estate, and the model is a component with a configuration value. The page on the self-hosted AI agent platform describes that offer, and our earlier analysis of the OpenAI Agents API covers what changes when a provider takes the loop over.
Official MCP figure from the Model Context Protocol documentation (modelcontextprotocol.io/docs/getting-started/intro, specification version 2026-07-28, retrieved 2026-10-04): the host, clients, servers and data sources, which is the seam that keeps tools reusable across runtimes.
The uncomfortable part is worth stating. Owning the harness means owning compaction, retries, session storage and the migration calendar, and none of that is where the demo excitement lives. It is, however, the difference between an agent that runs for a year and an agent that has to be rebuilt when the model behind it is retired.
Originally published on the Netics blog.




Top comments (0)