DEV Community

Takashi Abe
Takashi Abe

Posted on Originally published at zenn.dev

"Is That a Harness Too?" Making Sense of Harness Engineering for AI Agents

When sales says "everything is a harness," they're not wrong. It's just too broad

Lately I keep hearing the same lines from sales folks around me: "Oh, that's harness engineering," or "Are you really thinking about harness engineering?" What exactly do they mean by it? For some people, building an MCP server is harness engineering. So is wiring up RAG. So is adding guardrails. Everything is harness engineering.

Let's get on the same page first.

By harness engineering, I don't mean test harnesses, and I don't mean Harness.io, the CI/CD product. I mean Harness Engineering in the sense of what it takes to turn an AI agent into a system that actually works in practice.

With that settled: in the broad sense, the sales instinct isn't wrong. But taken at face value, it's far too broad to be useful as an IT architecture term. It ends up meaning roughly "we're doing something around AI agents."

So I tried defining it again:

Harness Engineering is the design of the execution environment and feedback loops around a probabilistic AI model, so that it becomes a "work-executing agent" that is "controllable," "verifiable," and "observable."

This isn't a technique for making AI smarter. It's a technique for taking an AI that is admittedly smart but unreliable, and containing it so it can do real work. Below is how I arrived at that one sentence, plus a quick test for when sales tells you "that's a harness."

A harness is the whole loop outside the model

As of 2026, there's a formula you see everywhere:

$$
Agent = Model + Harness
$$

LangChain uses almost exactly this phrasing. Microsoft defines an Agent Harness as "the execution foundation that makes a language model function as an AI agent that can actually get work done." It's made up of the pieces responsible for conversation state, context, tool calls, approval policies, multi-step progress, and so on.

Roughly sketched, it looks like this:

flowchart TD
    E[User / Event] --> C
    M([Model])
    subgraph Harness
        C[Context]
        A[Action / Tool]
        R[Result / Observation]
        V{Validate / Evaluate}
        RT[Retry]
        F[Finish]
        ES[Escalate]
    end
    C --> M
    M --> A
    A --> R
    R --> V
    V -->|Failed| RT
    V -->|Done| F
    V -->|Risky| ES
    RT --> C

This entire outer loop is the harness.

The model's job is to figure out "what's probably the right next move." The harness's job is to decide what to show it, what to let it do, what to forbid, how to check the results, what to do on failure, and when to stop. The automation I've been building lately by combining Jev with n8n is, I'd argue, exactly what you'd call harness engineering.

The core is the control loop, not the tool connections

Get this wrong and Harness Engineering turns into "every AI-adjacent technology, bundled."

Google Cloud describes building a harness as "wrapping a language model in scripts that manage its inputs, outputs, and tool execution." Tools are executed via MCP and the like, the results are appended to the conversation history, and everything is fed back to the model. For enterprise use, it's positioned as the layer that manages state, permissions, and tool execution, and because it owns all input and output, it also makes monitoring and cost tracking easier.

So if you think about it in stages, the boundary comes into view. This alone is just "API integration":

flowchart LR
    LLM --> API

Once you get here, you're gradually entering harness territory:

flowchart TD
    LLM --> EX[Call API]
    EX --> CK[Check result]
    CK --> Q{Goal achieved?}
    Q -->|No| NX[Decide next action]
    Q -->|Yes| FIN[Finish]

And if you've designed all the way to the third stage below, I think you can pretty clearly call it Harness Engineering:

flowchart TD
    P[Check permissions] --> E[Execute]
    E --> V[Verify result]
    V --> FX[Fix if failed]
    FX --> ES[Escalate to a human if risky]
    ES --> LG[Store logs and eval data]

It doesn't end once things are connected. You look at what came back, decide the next action, and take care of retries and cutoffs. That "taking care of it" part is the core.

MCP, RAG, and guardrails are all just parts of a harness

In practice, it helps to break down what a harness decides like this:

Area What the harness decides
Instructions What goal the agent acts toward
Context What to show the model this time
Retrieval What to pull from internal docs and DBs
Tools What it can execute
Identity / Permission Whose permissions it acts with, and how far
Orchestration In what order and loops it runs
State / Memory How state from previous runs is kept
Guardrails What is forbidden
Verification How to check the result is correct
Retry / Recovery How to roll back on failure
Human approval Where a human steps in
Observability What gets kept as logs and traces
Evaluation Whether the agent behaves as expected
Termination When the job counts as done

Microsoft's current Agent Harness likewise groups planning, file memory, tool approval, OpenTelemetry, web search, skills, background agents, shell execution, looping, and more as harness capabilities.

Looking at this table, it's clear that MCP is not the harness itself. MCP is one mechanism for connecting the tools a harness uses. RAG is the same: one part of Context / Retrieval. Guardrails are one part. Memory is one part.

Whether you're talking with sales or with customers, keeping these separate makes life easier as an architect. "We connected MCP" and "we designed a harness" simply don't belong in the same sentence.

Prompt Engineering and Context Engineering both live inside the harness

The relationship among the three is easiest to understand as nesting:

flowchart TD
    H[Harness Engineering] --> P[Prompt / Instruction Engineering]
    H --> C[Context Engineering]
    C --> RAG
    C --> MEM[Memory]
    C --> PD[Progressive Disclosure]
    H --> T[Tool Engineering]
    T --> API
    T --> MCP
    H --> O[Orchestration]
    H --> G[Guardrails / Permissions]
    H --> EV[Evaluation / Verification]
    H --> OB[Observability / Recovery]

Birgitta Böckeler's breakdown, published on everyone's favorite Martin Fowler site, splits Harness Engineering into Guides and Sensors. Her framing is scoped specifically to working with coding agents.

A Guide is feedforward: it steers the agent in the right direction before it fails. Rules, documentation, system prompts, Skills, and tool specs all fall here.

A Sensor is feedback: it observes what happened and flags mistakes. Tests, linters, logs, metrics, reviews by another AI, and so on.

flowchart LR
    G["Guide (feedforward)"] --> AG[Agent]
    GO[Goal] --> AG
    AG --> AC[Action]
    AC --> S["Sensor (feedback)"]
    S --> AG

Giving instructions is only half a harness. It becomes a harness only once you've built the part where the agent acts, you observe, and it corrects itself.

Polishing prompts strengthens the Guide. If that's all you do, nobody will notice when the agent gets something wrong.

Why harnesses, why now? Because LLMs got smarter

It's a bit paradoxical, but the reason is that models got smarter.

Back when models were weak, it looked like this:

flowchart LR
    HU[Human] --> AI
    AI --> AN[Answer]

A human asks, the AI answers, and if it's wrong, the human catches it. That was enough.

Once models are strong enough, it looks like this:

flowchart TD
    AI --> S[Investigate]
    S --> J[Decide]
    J --> E[Execute]
    E --> C[Check result]
    C --> N[Next decision]

Now that they can act this autonomously, the question has shifted from "can the AI answer?" to "how much can we hand over to the AI?"

The Harness Engineering post OpenAI published in February 2026 is about exactly this. When they had Codex write code, early progress was slow not because Codex lacked the ability, but because the environment wasn't described well enough. In the end, they built testing, verification, review, feedback handling, and failure recovery into the system side as part of the development loop.

In other words, the problem moved from model capability to system engineering. That, I think, is why Harness Engineering is suddenly getting so much attention. And it's good news for architects: we can't touch the model's internals, but system engineering is our home turf.

Sales support, refunds, and incident response: where the boundary shows up

Let's walk through some concrete examples.

Sales support agent

"Prep me for the next meeting with Company X."

On its own, that's just using an LLM. Build a harness, and you get a flow like this:

flowchart TD
    A[Fetch Account from CRM] --> B[Fetch Contacts / Opportunities]
    B --> C[Fetch past meetings]
    C --> D[Check internal access rights]
    D --> E[Search external sources if needed]
    E --> F[Draft meeting hypotheses]
    F --> G[Cross-check against sources]
    G --> H[Draft CRM updates]
    H --> I[Human approves]
    I --> J[Update CRM]

The CRM itself is not the harness. The harness is the mechanism that controls when the CRM is accessed, under whose permissions, what gets fetched, and how much the agent is allowed to update.

Customer support agent

Say a customer writes, "I was double-charged, refund me."

A bare model would probably fire back "We're so sorry. We'll refund you" right away. Anyone can say that. An agent that can actually issue a refund looks like this:

flowchart TD
    A[Verify identity] --> B[Fetch order]
    B --> C[Check payment history]
    C --> D[Fetch refund policy]
    D --> E[Calculate refundable amount]
    E --> F[Threshold check]
    F --> G[Human approval if needed]
    G --> H[Call payment API]
    H --> I[Re-fetch to confirm the refund actually happened]
    I --> J[Update CRM]
    J --> K[Reply to customer]

That step near the end, "re-fetch to confirm the refund actually happened," matters enormously. "We called the refund API" and "the refund went through" are two different things.

In Harness Engineering, verification matters more than action.

Incident response agent

For "The website is slow. Find out why," the following flow is relatively safe:

flowchart TD
    A[Fetch metrics] --> B[Fetch logs]
    B --> C[Inspect traces]
    C --> D[Form hypotheses]
    D --> E[Investigate further]
    E --> F[Identify likely causes]
    F --> G[Draft a fix]

Because it only reads and thinks.

But once you allow actions that affect running processes, the story changes:

kubectl rollout restart
Enter fullscreen mode Exit fullscreen mode

The harness needs to check:

  1. Is this a production change?
  2. What's the blast radius?
  3. Can it be rolled back?
  4. Does it need approval?
  5. Did things return to normal afterward?

Only when you can answer all five does it hold up as an enterprise system.

What all three examples share is that between the agent's "read" phase and its "change" phase, there's a layer of permissions, verification, and human approval. That layer in between is the harness itself.

A harness is never "done"

Here are the pros and cons of a harness side by side:

Pros Cons
Improves performance independently of model changes The system gets considerably more complex
Lets you raise agent autonomy safely More to operate: traces, evals, permissions, state, and more
Improves reproducibility, auditability, explainability Model improvements can turn an old harness into a hindrance
Reduces human review work Excessive rules can kill the model's capabilities
Feeds failures back into the next Guide / Sensor The harness itself needs testing

The third con is especially easy to miss.

Anthropic pointed out in April 2026 that a harness encodes assumptions about "what Claude can't do on its own," so as models improve, those assumptions need to be revisited.

They also gave a real example. They had added context resets to deal with Sonnet 4.5's habit of sensing the context limit and wrapping up work early. But with Opus 4.5, that habit disappeared, and the resets became dead weight.

That's why a harness is never finished once built. You need to trim it down or rebuild it alongside model updates. What you're maintaining is "the parts built on the assumption that the AI will get things wrong."

The problem with "everything is a harness" is that the word stops being useful

In conversations with sales, this is probably the most important part.

Böckeler's article itself says that "Agent = Model + Harness is a very broad definition, so it's worth narrowing it down by type of agent." She then explicitly limits her own definition to the bounded context of working with coding agents. I think that's the right approach.

For example, calling "I built an MCP server" Harness Engineering isn't wrong.
But the moment you label every item below as "Harness Engineering," the term becomes nearly synonymous with "we're doing something around AI agents." A word that doesn't tell you what it refers to is useless in a design conversation.

  • REST API development
  • IAM
  • RAG
  • Prompt
  • MCP
  • Workflow
  • Logging
  • Monitoring
  • Evaluation

So how about distinguishing them like this?

What you're doing What to call it
Improving instructions Prompt / Instruction Engineering
Selecting information, RAG, Memory Context Engineering
MCP / API connections Tool Integration
Procedures, routing Agent Orchestration / Workflow
Permissions, prohibitions Guardrails / Governance
Evaluation Agent Evaluation
Traces, metrics Agent Observability / AgentOps
Combining all of these to control autonomous execution Harness Engineering

Harness Engineering isn't so much the name of an individual technology as a higher-level concept: designing the control system that integrates all of these to make an agent's behavior work. Framed that way, everything falls neatly into place.

It's fine for sales to tell a customer "you need Harness Engineering." The architect's job is to break that one line down into the rows of the table above and show what's covered and what's missing.

Eight questions to answer "Is that a harness?"

When sales says "this is Harness Engineering," checking these in order seems to sort things out:

Priority Question
1 What does the agent have to achieve to be "done"?
2 What can the agent see?
3 What can the agent execute?
4 What must it not execute?
5 How are execution results verified?
6 How does it retry / recover when it gets things wrong?
7 Where does it escalate to a human?
8 Can its decisions and actions be traced afterward?

If all eight are designed, then "I see, that is Harness Engineering" is a fair verdict.

If, on the other hand, it's just "we connected the CRM via MCP," that's Tool Integration. It's one of the parts that make up a harness, but it isn't Harness Engineering as a whole. That's how I'd sort it.

The order of these eight questions matters, too. For an agent with no defined "done" condition (question 1), there's no way to design "verification" (question 5) or "traceability" (question 8). That's why you always start with question 1.

Harness Engineering is really just good old system design

Harness Engineering is a new name, but from an architect's point of view, it's not all that foreign. Isn't it what we've been doing all along?

Let's list what we've always done:

  • Define boundaries
  • Define permissions
  • Manage state
  • Detect failures
  • Retry
  • Roll back
  • Monitor
  • Test
  • Escalate to humans

That's system design, plain and simple. Designing batch reruns, compensating transactions, alerting for operations monitoring: all of it fits on this list.

There is, however, one big difference.

Traditionally, humans defined a program's behavior as code. In an agentic system, part of that behavior is delegated to a non-deterministic model. That's why you need a strong control system around it. That's the harness.

I've been writing programs for over 40 years, and I believe that turning routine work into programs is what computers are fundamentally for. That belief hasn't changed in the age of AI. If anything, it's become clearer.

Whatever can be made routine, you lock down with deterministic, definitive programs, just as before. Only the judgments that can't be made routine get handed to a probabilistic model. And the results of those delegated judgments get verified, again, by deterministic programs. A harness is the design of that "deterministic part." Let computers do what they're good at, and don't let uncertain things run loose as they are. Drawing that line has always been the architect's job.

Seen this way, the difference among the three terms fits in a single line.

If Prompt Engineering is "how you instruct the AI" and Context Engineering is "what you show the AI," then Harness Engineering is the craft of designing "what kind of world you put the AI to work in."

With this definition, I think we can finally push back on the "everything is a harness" problem.

Top comments (0)