On September 9, I wrote that an AI agent is just a loop.
And that the interesting part is everything we have to build around it.
This month, that "everything" got a name almost everywhere I looked.
The harness.
- On September 10, OpenAI launched the Agents API as a way to "build and run cloud agents with the Codex harness."
- On September 17, a team of researchers from UMass Amherst, Emory, UNC Charlotte and Zoom published An Empirical Study of Harness Design for Coding Agents.
- On September 21, Strands released Strands harness, which it describes as "a fully assembled state-of-the-art agent harness."
- The same day, the paper behind Google Research's RRSI came out. Its README opens with: "An LLM agent's capability is largely set by its harness."
Different teams.
Same word.
So what is an agent harness, exactly?
The short answer:
The model decides what to do next. The harness decides what actually happens.
The longer answer is easier to build than to explain.
So let's build a tiny one.
By the end, you will run one command:
npx tsx harness.ts
And watch a harness call tools, ask for approval, block a dangerous action, trim its own context, and stop a model that never finishes.
No API key.
No framework.
No real model.
Just TypeScript and one idea:
Everything around the model is engineering you can see, test and control.
Table of Contents
- What Is an Agent Harness?
- What We Are Building
- Project Setup
- Step 1: Define the Pieces
- Step 2: Register the Tools
- Step 3: Check Permissions
- Step 4: Trim the Context
- Step 5: Write the Loop
- Step 6: Mock the Model
- Run the Harness
- What This Harness Gets Right
- Where It Breaks Down
- The Bigger Idea
What Is an Agent Harness?
The cleanest definition I've found comes from LangChain.
In The Anatomy of an Agent Harness (March 10, 2026), Vivek Trivedy writes:
Agent = Model + Harness
If you're not the model, you're the harness.
He defines a harness as "every piece of code, configuration, and execution logic that isn't the model itself."
That's a useful line to draw.
The term isn't brand new either.
When Anthropic renamed the Claude Code SDK to the Claude Agent SDK on September 29, 2025, it called it "the agent harness that powers Claude Code."
And OpenAI's Agents API post says it plainly:
Useful agents need a powerful harness that manages context, uses tools efficiently, and coordinates subagents.
Here is how I think about the split:
| Model | Harness | |
|---|---|---|
| What it does | Reasons and proposes the next step | Runs the loop and executes the step |
| What it sees | Whatever context it is given | Everything that happened |
| Tools | Asks for them | Owns them |
| Permissions | Can't enforce them | Enforces them |
| When to stop | Suggests it | Decides it |
| Behavior | Probabilistic | Deterministic |
That last row is the one I care about most.
On September 9, I wrote that models are probabilistic but your infrastructure shouldn't be.
The harness is where that infrastructure lives.
It also sits in a specific place in the stack:
Model
↓
Loop (observe → decide → act → repeat)
↓
Harness (tools, permissions, context, budgets, logs)
↓
Lane (the job the agent owns)
On September 20, in You Can't Give an AI a Job Until It Has a Lane, I wrote: An agent is a loop. A job is a lane.
The harness is the part in between.
It's what makes the loop safe enough to be given a lane.
What We Are Building
Our harness will have six parts:
| Part | Job |
|---|---|
| Loop | Calls the model until it finishes or runs out of steps |
| Tool registry | The only actions the model is allowed to request |
| Permission check | Decides allow, ask, or deny for every tool call |
| Context trimming | Keeps what the model sees inside a budget |
| Step budget | Stops a model that never says it's done |
| Event log | Records what actually happened |
The model will be a mock.
That's intentional.
If you can't test your harness without a model, you can't really test your harness.
The example is a small release assistant. The changelog, file names and email address are made up.
Project Setup
You will need Node.js 18 or newer.
Create the project:
mkdir tiny-harness
cd tiny-harness
npm init -y
npm install --save-dev typescript tsx @types/node
Save the following TypeScript blocks, in order, as harness.ts.
Every block below is part of the same file.
Step 1: Define the Pieces
type Role = "system" | "user" | "assistant" | "tool";
type Message = {
role: Role;
content: string;
};
type ModelTurn =
| { type: "tool"; tool: string; input: string }
| { type: "finish"; answer: string };
type Model = (context: Message[]) => ModelTurn;
type Permission = "allow" | "ask" | "deny";
type Tool = {
name: string;
permission: Permission;
run: (input: string) => string;
};
type HarnessEvent = {
step: number;
kind: "call" | "approved" | "blocked" | "result" | "trim" | "stop";
detail: string;
};
There are two important ideas here.
First, the model has exactly one job:
type Model = (context: Message[]) => ModelTurn;
It receives context.
It returns a proposal.
Either "call this tool with this input" or "I'm finished."
It never runs anything.
Second, every tool carries its own permission.
The risk level belongs to the tool, not to the prompt.
A model can be talked into things.
A type can't.
Step 2: Register the Tools
const files = new Map<string, string>([
["CHANGELOG.md", "1.4.0: fix retry backoff, add CSV export"],
]);
const tools: Tool[] = [
{
name: "read_file",
permission: "allow",
run: (path) => files.get(path) ?? `not found: ${path}`,
},
{
name: "write_file",
permission: "allow",
run: (input) => {
const [path, body = ""] = input.split("::");
files.set(path, body);
return `wrote ${path}`;
},
},
{
name: "send_email",
permission: "ask",
run: (to) => `sent to ${to}`,
},
{
name: "delete_branch",
permission: "deny",
run: (branch) => `deleted ${branch}`,
},
];
const registry = new Map(tools.map((tool) => [tool.name, tool]));
These tools are deliberately boring.
The interesting part is the registry at the bottom.
That Map is the entire world the model can touch.
If a tool isn't in the registry, it doesn't exist.
Notice that delete_branch is in the registry with permission: "deny".
Why register a tool you never want to run?
Because models will ask for things anyway.
I would rather see the request in the log than pretend it could never happen.
Step 3: Check Permissions
type Approver = (tool: string, input: string) => boolean;
type Verdict = { ok: boolean; reason: string };
function checkPermission(
tool: Tool,
input: string,
approve: Approver,
): Verdict {
if (tool.permission === "allow") {
return { ok: true, reason: "allowed" };
}
if (tool.permission === "deny") {
return { ok: false, reason: "denied by policy" };
}
return approve(tool.name, input)
? { ok: true, reason: "approver said yes" }
: { ok: false, reason: "approver said no" };
}
This is the smallest function in the harness.
It might be the most important one.
allow → run it
ask → a human decides
deny → never, no matter what the model says
In a real product, approve would be a Slack button, an email, or a review queue.
Here it's a plain function so the example runs on its own.
The empirical study I mentioned uses the same three words. Its harness has a permission layer that "classifies each action as allow, ask, or deny."
That's a nice sign the pattern is converging.
Step 4: Trim the Context
const estimateTokens = (message: Message) =>
Math.ceil(message.content.length / 4);
function trimContext(messages: Message[], budget: number) {
const [system, task, ...rest] = messages;
let used = estimateTokens(system) + estimateTokens(task);
const recent: Message[] = [];
for (let i = rest.length - 1; i >= 0; i--) {
const cost = estimateTokens(rest[i]);
if (used + cost > budget) {
break;
}
recent.unshift(rest[i]);
used += cost;
}
return {
context: [system, task, ...recent],
dropped: rest.length - recent.length,
};
}
The model never sees the whole history.
It sees the system prompt, the task, and as many recent messages as fit inside the budget.
Newest first.
Oldest dropped.
Dividing characters by four is a rough token estimate, not a real tokenizer. It's fine for a demo and wrong for billing.
The important part is where this lives.
Context is something the harness manages, not something the model manages.
Production harnesses go much further. Strands says its harness truncates tool results over about 1,500 tokens and triggers compaction when the context window passes 85%. OpenAI says its Agents API handles earlier context as a session approaches its limit, so developers don't have to write their own compaction logic.
And the empirical study found that context management matters most when the context budget is tight, mostly because it prevents runs from failing on overflow.
Ours just drops old messages.
It's the simplest version of the same idea.
Step 5: Write the Loop
type HarnessConfig = {
model: Model;
approve: Approver;
maxSteps: number;
contextBudget: number;
};
function runHarness(task: string, config: HarnessConfig) {
const log: HarnessEvent[] = [];
const history: Message[] = [
{ role: "system", content: "You are a release assistant." },
{ role: "user", content: task },
];
for (let step = 1; step <= config.maxSteps; step++) {
const { context, dropped } = trimContext(
history,
config.contextBudget,
);
if (dropped > 0) {
log.push({ step, kind: "trim", detail: `dropped ${dropped} old messages` });
}
const turn = config.model(context);
if (turn.type === "finish") {
log.push({ step, kind: "stop", detail: turn.answer });
return log;
}
log.push({ step, kind: "call", detail: `${turn.tool}(${turn.input})` });
history.push({ role: "assistant", content: `${turn.tool} ${turn.input}` });
const tool = registry.get(turn.tool);
const verdict: Verdict = tool
? checkPermission(tool, turn.input, config.approve)
: { ok: false, reason: "unknown tool" };
if (!tool || !verdict.ok) {
log.push({ step, kind: "blocked", detail: verdict.reason });
history.push({ role: "tool", content: `blocked: ${verdict.reason}` });
continue;
}
if (tool.permission === "ask") {
log.push({ step, kind: "approved", detail: verdict.reason });
}
const result = tool.run(turn.input);
log.push({ step, kind: "result", detail: result });
history.push({ role: "tool", content: result });
}
log.push({
step: config.maxSteps,
kind: "stop",
detail: "step budget reached",
});
return log;
}
There is a lot happening here, so let's walk through the important parts.
Every Step Starts With Trimming
The model is called with context, not history.
The harness keeps the full history for itself.
The model gets the view.
Blocked Calls Go Back to the Model
When a call is denied, we don't throw.
We push blocked: denied by policy into the history and keep going.
The model gets to see what happened and choose something else.
The empirical study does the same thing: "Tool errors are returned to the model as observations rather than raised."
A failed action shouldn't crash the loop.
The Budget Is Outside the Model
This line matters more than it looks:
for (let step = 1; step <= config.maxSteps; step++) {
The model can say it's done.
The harness decides it's done.
A model that never finishes isn't ambitious. It's expensive.
Everything Gets Logged
Every call, approval, block, result, trim and stop becomes a HarnessEvent.
On September 19, in The Computer Is Becoming an API for AI, I wrote that computer use without observability is chaos wearing a robot costume.
That's true for any agent.
The event log is the smallest possible answer.
Step 6: Mock the Model
function scriptedModel(script: ModelTurn[]): Model {
let turn = 0;
return () =>
script[turn++] ?? { type: "finish", answer: "nothing left to do" };
}
const releaseModel = scriptedModel([
{ type: "tool", tool: "read_file", input: "CHANGELOG.md" },
{
type: "tool",
tool: "write_file",
input: "RELEASE.md::Fixed retry backoff. Added CSV export.",
},
{ type: "tool", tool: "send_email", input: "team@example.com" },
{ type: "tool", tool: "delete_branch", input: "release/1.4.0" },
{ type: "finish", answer: "Release notes written and sent." },
]);
const stuckModel: Model = () => ({
type: "tool",
tool: "read_file",
input: "CHANGELOG.md",
});
const approve: Approver = (tool, input) =>
tool === "send_email" && input.endsWith("@example.com");
scriptedModel returns one planned turn per call.
A real model would read the context and decide.
Ours ignores it and follows a script.
That sounds like cheating.
It's actually the point.
The harness doesn't care how the decision was made. It only cares what the model proposed, and whether that proposal is allowed.
stuckModel is the model we all eventually meet. It asks for the same file forever.
The approve function stands in for a person. It only says yes to email sent to our own domain.
Run the Harness
Add the last block:
function print(title: string, log: HarnessEvent[]) {
console.log(title);
for (const event of log) {
console.log(` ${event.step} ${event.kind.padEnd(9)}${event.detail}`);
}
console.log("");
}
print(
"Run 1: release notes",
runHarness("Write release notes for 1.4.0 and email the team.", {
model: releaseModel,
approve,
maxSteps: 10,
contextBudget: 200,
}),
);
print(
"Run 2: a model that never finishes",
runHarness("Write release notes for 1.4.0 and email the team.", {
model: stuckModel,
approve,
maxSteps: 4,
contextBudget: 40,
}),
);
Run it:
npx tsx harness.ts
You should see:
Run 1: release notes
1 call read_file(CHANGELOG.md)
1 result 1.4.0: fix retry backoff, add CSV export
2 call write_file(RELEASE.md::Fixed retry backoff. Added CSV export.)
2 result wrote RELEASE.md
3 call send_email(team@example.com)
3 approved approver said yes
3 result sent to team@example.com
4 call delete_branch(release/1.4.0)
4 blocked denied by policy
5 stop Release notes written and sent.
Run 2: a model that never finishes
1 call read_file(CHANGELOG.md)
1 result 1.4.0: fix retry backoff, add CSV export
2 call read_file(CHANGELOG.md)
2 result 1.4.0: fix retry backoff, add CSV export
3 trim dropped 2 old messages
3 call read_file(CHANGELOG.md)
3 result 1.4.0: fix retry backoff, add CSV export
4 trim dropped 4 old messages
4 call read_file(CHANGELOG.md)
4 result 1.4.0: fix retry backoff, add CSV export
4 stop step budget reached
Two runs.
Same harness.
Different model behavior.
In run 1:
- Reading and writing files is allowed, so it just happens.
- Sending email is
ask, so the approver decides. - Deleting a branch is
deny, so it never runs, even though the model asked. - The model finishes on step 5.
In run 2:
- The model never finishes.
- The context budget is small, so the harness starts dropping old messages on step 3.
- The step budget stops the run on step 4.
Nothing in the model changed its mind.
The harness did the stopping.
What This Harness Gets Right
The Model Only Proposes
Every action goes through code we wrote.
The model can suggest delete_branch.
It can't run it.
Permissions Live With the Tools
The rule for send_email isn't buried in a prompt.
It's a field on the tool.
You can read it, test it and review it in a pull request.
Limits Are Explicit
Steps and context have hard budgets.
Nobody has to hope the model notices it's been running for an hour.
The Model Is Swappable
We ran two very different "models" through the same harness.
That's the whole point of the separation.
Swap the mock for a real provider and the permissions, budgets and log stay exactly the same.
Where It Breaks Down
This is a teaching harness.
Here is what a real one would need.
The Model Is a Script
A real model reads the context and can surprise you.
That's where tool descriptions, system prompts and planning start to matter.
Trimming Forgets Instead of Summarizing
Dropping old messages is lossy.
If the important fact was in message 3, it's gone.
Real harnesses compact: they summarize older context, offload large tool results to files, and keep the task visible.
It Doesn't Notice It's Stuck
Run 2 burned its whole budget reading the same file.
The empirical study's harness watches for streaks of identical tool calls, reminds the model to change approach, and ends the run early if identical failing calls keep piling up.
Ours just waits for the budget.
Approval Is a Function, Not a Person
A real ask has to pause the run, notify someone, wait, and resume later.
That means durable state.
That means the harness is slowly becoming a distributed system.
There Is No Sandbox
Our tools run in the same process as the harness.
Real tools touch files, shells and browsers, and they should run somewhere isolated.
OpenAI's Agents API draws this line explicitly: "OpenAI hosts and maintains the harness. You choose the agent's compute environment."
The Log Is a List, Not a Graph
Our event log records what happened.
It doesn't record why.
Which result led to which call? Which approval unlocked which action?
Those are relationships.
That's the graph engineering idea again: once events are connected, you can ask better questions about them.
There Are No Evals
If you change the harness, how do you know you made it better?
The empirical study held the loop fixed and varied planning, tools and context management across 176 settings to find out.
The RRSI README goes further and describes evolving the harness itself, while noting that it can overfit to the tasks it was tuned on.
Harnesses need tests too.
These limitations aren't reasons to skip the prototype.
They're a map of where a real harness has to go next.
The Bigger Idea
When people compare agents, they usually compare models.
I think that's only half the comparison.
The same model inside a different harness is a different agent. The empirical study puts it carefully: changing the harness while holding the model fixed "can substantially change model performance."
A real agent system starts looking like this:
┌─────────────────────────────────────┐
│ Harness │
│ │
│ Context ──→ Model ──→ Proposal │
│ ↑ ↓ │
│ │ Permissions │
│ │ ↓ │
│ Event log ←── Result ←── Tool │
│ │
│ Step budget: stop when it's out │
└─────────────────────────────────────┘
The model provides reasoning.
The loop provides persistence.
The tool registry provides capabilities.
Permissions provide boundaries.
The context budget provides focus.
The step budget provides a stopping point.
The event log provides evidence.
And the lane provides a reason for all of it to exist.
On September 9, I wrote that the loop may be 30 lines of code, and everything around the loop is the product.
Now I know what to call that part.
The model thinks. The harness works.
Try Roster
I'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access, running inside a harness that decides what they're allowed to do with all of it.
If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.



Top comments (0)