DEV Community

Cover image for What Is an Agent Harness? Build a Tiny One in TypeScript
Bobby Hall Jr
Bobby Hall Jr

Posted on

What Is an Agent Harness? Build a Tiny One in TypeScript

On September 9, I wrote that an AI agent is just a loop.

And that the interesting part is everything we have to build around it.

This month, that "everything" got a name almost everywhere I looked.

The harness.

  • On September 10, OpenAI launched the Agents API as a way to "build and run cloud agents with the Codex harness."
  • On September 17, a team of researchers from UMass Amherst, Emory, UNC Charlotte and Zoom published An Empirical Study of Harness Design for Coding Agents.
  • On September 21, Strands released Strands harness, which it describes as "a fully assembled state-of-the-art agent harness."
  • The same day, the paper behind Google Research's RRSI came out. Its README opens with: "An LLM agent's capability is largely set by its harness."

Different teams.

Same word.

So what is an agent harness, exactly?

The short answer:

The model decides what to do next. The harness decides what actually happens.

The longer answer is easier to build than to explain.

So let's build a tiny one.

By the end, you will run one command:

npx tsx harness.ts
Enter fullscreen mode Exit fullscreen mode

And watch a harness call tools, ask for approval, block a dangerous action, trim its own context, and stop a model that never finishes.

No API key.

No framework.

No real model.

Just TypeScript and one idea:

Everything around the model is engineering you can see, test and control.

Table of Contents

  1. What Is an Agent Harness?
  2. What We Are Building
  3. Project Setup
  4. Step 1: Define the Pieces
  5. Step 2: Register the Tools
  6. Step 3: Check Permissions
  7. Step 4: Trim the Context
  8. Step 5: Write the Loop
  9. Step 6: Mock the Model
  10. Run the Harness
  11. What This Harness Gets Right
  12. Where It Breaks Down
  13. The Bigger Idea

What Is an Agent Harness?

The cleanest definition I've found comes from LangChain.

In The Anatomy of an Agent Harness (March 10, 2026), Vivek Trivedy writes:

Agent = Model + Harness

If you're not the model, you're the harness.

He defines a harness as "every piece of code, configuration, and execution logic that isn't the model itself."

That's a useful line to draw.

The term isn't brand new either.

When Anthropic renamed the Claude Code SDK to the Claude Agent SDK on September 29, 2025, it called it "the agent harness that powers Claude Code."

And OpenAI's Agents API post says it plainly:

Useful agents need a powerful harness that manages context, uses tools efficiently, and coordinates subagents.

Here is how I think about the split:

Model Harness
What it does Reasons and proposes the next step Runs the loop and executes the step
What it sees Whatever context it is given Everything that happened
Tools Asks for them Owns them
Permissions Can't enforce them Enforces them
When to stop Suggests it Decides it
Behavior Probabilistic Deterministic

Model vs Harness

That last row is the one I care about most.

On September 9, I wrote that models are probabilistic but your infrastructure shouldn't be.

The harness is where that infrastructure lives.

It also sits in a specific place in the stack:

Model
  ↓
Loop        (observe → decide → act → repeat)
  ↓
Harness     (tools, permissions, context, budgets, logs)
  ↓
Lane        (the job the agent owns)
Enter fullscreen mode Exit fullscreen mode

Model to Loop to Harness to Lane

On September 20, in You Can't Give an AI a Job Until It Has a Lane, I wrote: An agent is a loop. A job is a lane.

The harness is the part in between.

It's what makes the loop safe enough to be given a lane.

What We Are Building

Our harness will have six parts:

Part Job
Loop Calls the model until it finishes or runs out of steps
Tool registry The only actions the model is allowed to request
Permission check Decides allow, ask, or deny for every tool call
Context trimming Keeps what the model sees inside a budget
Step budget Stops a model that never says it's done
Event log Records what actually happened

The model will be a mock.

That's intentional.

If you can't test your harness without a model, you can't really test your harness.

The example is a small release assistant. The changelog, file names and email address are made up.

Project Setup

You will need Node.js 18 or newer.

Create the project:

mkdir tiny-harness
cd tiny-harness

npm init -y
npm install --save-dev typescript tsx @types/node
Enter fullscreen mode Exit fullscreen mode

Save the following TypeScript blocks, in order, as harness.ts.

Every block below is part of the same file.

Step 1: Define the Pieces

type Role = "system" | "user" | "assistant" | "tool";

type Message = {
  role: Role;
  content: string;
};

type ModelTurn =
  | { type: "tool"; tool: string; input: string }
  | { type: "finish"; answer: string };

type Model = (context: Message[]) => ModelTurn;

type Permission = "allow" | "ask" | "deny";

type Tool = {
  name: string;
  permission: Permission;
  run: (input: string) => string;
};

type HarnessEvent = {
  step: number;
  kind: "call" | "approved" | "blocked" | "result" | "trim" | "stop";
  detail: string;
};
Enter fullscreen mode Exit fullscreen mode

There are two important ideas here.

First, the model has exactly one job:

type Model = (context: Message[]) => ModelTurn;
Enter fullscreen mode Exit fullscreen mode

It receives context.

It returns a proposal.

Either "call this tool with this input" or "I'm finished."

It never runs anything.

Second, every tool carries its own permission.

The risk level belongs to the tool, not to the prompt.

A model can be talked into things.

A type can't.

Step 2: Register the Tools

const files = new Map<string, string>([
  ["CHANGELOG.md", "1.4.0: fix retry backoff, add CSV export"],
]);

const tools: Tool[] = [
  {
    name: "read_file",
    permission: "allow",
    run: (path) => files.get(path) ?? `not found: ${path}`,
  },
  {
    name: "write_file",
    permission: "allow",
    run: (input) => {
      const [path, body = ""] = input.split("::");
      files.set(path, body);
      return `wrote ${path}`;
    },
  },
  {
    name: "send_email",
    permission: "ask",
    run: (to) => `sent to ${to}`,
  },
  {
    name: "delete_branch",
    permission: "deny",
    run: (branch) => `deleted ${branch}`,
  },
];

const registry = new Map(tools.map((tool) => [tool.name, tool]));
Enter fullscreen mode Exit fullscreen mode

These tools are deliberately boring.

The interesting part is the registry at the bottom.

That Map is the entire world the model can touch.

If a tool isn't in the registry, it doesn't exist.

Notice that delete_branch is in the registry with permission: "deny".

Why register a tool you never want to run?

Because models will ask for things anyway.

I would rather see the request in the log than pretend it could never happen.

Step 3: Check Permissions

type Approver = (tool: string, input: string) => boolean;

type Verdict = { ok: boolean; reason: string };

function checkPermission(
  tool: Tool,
  input: string,
  approve: Approver,
): Verdict {
  if (tool.permission === "allow") {
    return { ok: true, reason: "allowed" };
  }

  if (tool.permission === "deny") {
    return { ok: false, reason: "denied by policy" };
  }

  return approve(tool.name, input)
    ? { ok: true, reason: "approver said yes" }
    : { ok: false, reason: "approver said no" };
}
Enter fullscreen mode Exit fullscreen mode

This is the smallest function in the harness.

It might be the most important one.

allow → run it
ask   → a human decides
deny  → never, no matter what the model says
Enter fullscreen mode Exit fullscreen mode

Allow, ask, deny: every tool carries its own permission

In a real product, approve would be a Slack button, an email, or a review queue.

Here it's a plain function so the example runs on its own.

The empirical study I mentioned uses the same three words. Its harness has a permission layer that "classifies each action as allow, ask, or deny."

That's a nice sign the pattern is converging.

Step 4: Trim the Context

const estimateTokens = (message: Message) =>
  Math.ceil(message.content.length / 4);

function trimContext(messages: Message[], budget: number) {
  const [system, task, ...rest] = messages;

  let used = estimateTokens(system) + estimateTokens(task);
  const recent: Message[] = [];

  for (let i = rest.length - 1; i >= 0; i--) {
    const cost = estimateTokens(rest[i]);

    if (used + cost > budget) {
      break;
    }

    recent.unshift(rest[i]);
    used += cost;
  }

  return {
    context: [system, task, ...recent],
    dropped: rest.length - recent.length,
  };
}
Enter fullscreen mode Exit fullscreen mode

The model never sees the whole history.

It sees the system prompt, the task, and as many recent messages as fit inside the budget.

Newest first.

Oldest dropped.

Dividing characters by four is a rough token estimate, not a real tokenizer. It's fine for a demo and wrong for billing.

The important part is where this lives.

Context is something the harness manages, not something the model manages.

Production harnesses go much further. Strands says its harness truncates tool results over about 1,500 tokens and triggers compaction when the context window passes 85%. OpenAI says its Agents API handles earlier context as a session approaches its limit, so developers don't have to write their own compaction logic.

And the empirical study found that context management matters most when the context budget is tight, mostly because it prevents runs from failing on overflow.

Ours just drops old messages.

It's the simplest version of the same idea.

Step 5: Write the Loop

type HarnessConfig = {
  model: Model;
  approve: Approver;
  maxSteps: number;
  contextBudget: number;
};

function runHarness(task: string, config: HarnessConfig) {
  const log: HarnessEvent[] = [];

  const history: Message[] = [
    { role: "system", content: "You are a release assistant." },
    { role: "user", content: task },
  ];

  for (let step = 1; step <= config.maxSteps; step++) {
    const { context, dropped } = trimContext(
      history,
      config.contextBudget,
    );

    if (dropped > 0) {
      log.push({ step, kind: "trim", detail: `dropped ${dropped} old messages` });
    }

    const turn = config.model(context);

    if (turn.type === "finish") {
      log.push({ step, kind: "stop", detail: turn.answer });
      return log;
    }

    log.push({ step, kind: "call", detail: `${turn.tool}(${turn.input})` });
    history.push({ role: "assistant", content: `${turn.tool} ${turn.input}` });

    const tool = registry.get(turn.tool);

    const verdict: Verdict = tool
      ? checkPermission(tool, turn.input, config.approve)
      : { ok: false, reason: "unknown tool" };

    if (!tool || !verdict.ok) {
      log.push({ step, kind: "blocked", detail: verdict.reason });
      history.push({ role: "tool", content: `blocked: ${verdict.reason}` });
      continue;
    }

    if (tool.permission === "ask") {
      log.push({ step, kind: "approved", detail: verdict.reason });
    }

    const result = tool.run(turn.input);

    log.push({ step, kind: "result", detail: result });
    history.push({ role: "tool", content: result });
  }

  log.push({
    step: config.maxSteps,
    kind: "stop",
    detail: "step budget reached",
  });

  return log;
}
Enter fullscreen mode Exit fullscreen mode

There is a lot happening here, so let's walk through the important parts.

Every Step Starts With Trimming

The model is called with context, not history.

The harness keeps the full history for itself.

The model gets the view.

Blocked Calls Go Back to the Model

When a call is denied, we don't throw.

We push blocked: denied by policy into the history and keep going.

The model gets to see what happened and choose something else.

The empirical study does the same thing: "Tool errors are returned to the model as observations rather than raised."

A failed action shouldn't crash the loop.

The Budget Is Outside the Model

This line matters more than it looks:

for (let step = 1; step <= config.maxSteps; step++) {
Enter fullscreen mode Exit fullscreen mode

The model can say it's done.

The harness decides it's done.

A model that never finishes isn't ambitious. It's expensive.

Everything Gets Logged

Every call, approval, block, result, trim and stop becomes a HarnessEvent.

On September 19, in The Computer Is Becoming an API for AI, I wrote that computer use without observability is chaos wearing a robot costume.

That's true for any agent.

The event log is the smallest possible answer.

Step 6: Mock the Model

function scriptedModel(script: ModelTurn[]): Model {
  let turn = 0;

  return () =>
    script[turn++] ?? { type: "finish", answer: "nothing left to do" };
}

const releaseModel = scriptedModel([
  { type: "tool", tool: "read_file", input: "CHANGELOG.md" },
  {
    type: "tool",
    tool: "write_file",
    input: "RELEASE.md::Fixed retry backoff. Added CSV export.",
  },
  { type: "tool", tool: "send_email", input: "team@example.com" },
  { type: "tool", tool: "delete_branch", input: "release/1.4.0" },
  { type: "finish", answer: "Release notes written and sent." },
]);

const stuckModel: Model = () => ({
  type: "tool",
  tool: "read_file",
  input: "CHANGELOG.md",
});

const approve: Approver = (tool, input) =>
  tool === "send_email" && input.endsWith("@example.com");
Enter fullscreen mode Exit fullscreen mode

scriptedModel returns one planned turn per call.

A real model would read the context and decide.

Ours ignores it and follows a script.

That sounds like cheating.

It's actually the point.

The harness doesn't care how the decision was made. It only cares what the model proposed, and whether that proposal is allowed.

stuckModel is the model we all eventually meet. It asks for the same file forever.

The approve function stands in for a person. It only says yes to email sent to our own domain.

Run the Harness

Add the last block:

function print(title: string, log: HarnessEvent[]) {
  console.log(title);

  for (const event of log) {
    console.log(`  ${event.step}  ${event.kind.padEnd(9)}${event.detail}`);
  }

  console.log("");
}

print(
  "Run 1: release notes",
  runHarness("Write release notes for 1.4.0 and email the team.", {
    model: releaseModel,
    approve,
    maxSteps: 10,
    contextBudget: 200,
  }),
);

print(
  "Run 2: a model that never finishes",
  runHarness("Write release notes for 1.4.0 and email the team.", {
    model: stuckModel,
    approve,
    maxSteps: 4,
    contextBudget: 40,
  }),
);
Enter fullscreen mode Exit fullscreen mode

Run it:

npx tsx harness.ts
Enter fullscreen mode Exit fullscreen mode

You should see:

Run 1: release notes
  1  call     read_file(CHANGELOG.md)
  1  result   1.4.0: fix retry backoff, add CSV export
  2  call     write_file(RELEASE.md::Fixed retry backoff. Added CSV export.)
  2  result   wrote RELEASE.md
  3  call     send_email(team@example.com)
  3  approved approver said yes
  3  result   sent to team@example.com
  4  call     delete_branch(release/1.4.0)
  4  blocked  denied by policy
  5  stop     Release notes written and sent.

Run 2: a model that never finishes
  1  call     read_file(CHANGELOG.md)
  1  result   1.4.0: fix retry backoff, add CSV export
  2  call     read_file(CHANGELOG.md)
  2  result   1.4.0: fix retry backoff, add CSV export
  3  trim     dropped 2 old messages
  3  call     read_file(CHANGELOG.md)
  3  result   1.4.0: fix retry backoff, add CSV export
  4  trim     dropped 4 old messages
  4  call     read_file(CHANGELOG.md)
  4  result   1.4.0: fix retry backoff, add CSV export
  4  stop     step budget reached
Enter fullscreen mode Exit fullscreen mode

Two runs.

Same harness.

Different model behavior.

In run 1:

  • Reading and writing files is allowed, so it just happens.
  • Sending email is ask, so the approver decides.
  • Deleting a branch is deny, so it never runs, even though the model asked.
  • The model finishes on step 5.

In run 2:

  • The model never finishes.
  • The context budget is small, so the harness starts dropping old messages on step 3.
  • The step budget stops the run on step 4.

Nothing in the model changed its mind.

The harness did the stopping.

What This Harness Gets Right

The Model Only Proposes

Every action goes through code we wrote.

The model can suggest delete_branch.

It can't run it.

Permissions Live With the Tools

The rule for send_email isn't buried in a prompt.

It's a field on the tool.

You can read it, test it and review it in a pull request.

Limits Are Explicit

Steps and context have hard budgets.

Nobody has to hope the model notices it's been running for an hour.

The Model Is Swappable

We ran two very different "models" through the same harness.

That's the whole point of the separation.

Swap the mock for a real provider and the permissions, budgets and log stay exactly the same.

Where It Breaks Down

This is a teaching harness.

Here is what a real one would need.

The Model Is a Script

A real model reads the context and can surprise you.

That's where tool descriptions, system prompts and planning start to matter.

Trimming Forgets Instead of Summarizing

Dropping old messages is lossy.

If the important fact was in message 3, it's gone.

Real harnesses compact: they summarize older context, offload large tool results to files, and keep the task visible.

It Doesn't Notice It's Stuck

Run 2 burned its whole budget reading the same file.

The empirical study's harness watches for streaks of identical tool calls, reminds the model to change approach, and ends the run early if identical failing calls keep piling up.

Ours just waits for the budget.

Approval Is a Function, Not a Person

A real ask has to pause the run, notify someone, wait, and resume later.

That means durable state.

That means the harness is slowly becoming a distributed system.

There Is No Sandbox

Our tools run in the same process as the harness.

Real tools touch files, shells and browsers, and they should run somewhere isolated.

OpenAI's Agents API draws this line explicitly: "OpenAI hosts and maintains the harness. You choose the agent's compute environment."

The Log Is a List, Not a Graph

Our event log records what happened.

It doesn't record why.

Which result led to which call? Which approval unlocked which action?

Those are relationships.

That's the graph engineering idea again: once events are connected, you can ask better questions about them.

There Are No Evals

If you change the harness, how do you know you made it better?

The empirical study held the loop fixed and varied planning, tools and context management across 176 settings to find out.

The RRSI README goes further and describes evolving the harness itself, while noting that it can overfit to the tasks it was tuned on.

Harnesses need tests too.

These limitations aren't reasons to skip the prototype.

They're a map of where a real harness has to go next.

The Bigger Idea

When people compare agents, they usually compare models.

I think that's only half the comparison.

The same model inside a different harness is a different agent. The empirical study puts it carefully: changing the harness while holding the model fixed "can substantially change model performance."

A real agent system starts looking like this:

┌─────────────────────────────────────┐
│               Harness               │
│                                     │
│   Context ──→ Model ──→ Proposal    │
│      ↑                     ↓        │
│      │                Permissions   │
│      │                     ↓        │
│   Event log ←── Result ←── Tool     │
│                                     │
│   Step budget: stop when it's out   │
└─────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The model provides reasoning.

The loop provides persistence.

The tool registry provides capabilities.

Permissions provide boundaries.

The context budget provides focus.

The step budget provides a stopping point.

The event log provides evidence.

And the lane provides a reason for all of it to exist.

On September 9, I wrote that the loop may be 30 lines of code, and everything around the loop is the product.

Now I know what to call that part.

The model thinks. The harness works.


Try Roster

I'm building Roster around this idea: AI employees with real responsibilities, tools, memory, schedules and computer access, running inside a harness that decides what they're allowed to do with all of it.

If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.

Try Roster →

Top comments (0)