When an AI agent fails, the first reaction is usually: "We need a better model."
But most of the time, the model isn't the problem. The problem is everything around it.
That "everything around it" has a name: the harness.
What is a harness?
A language model on its own does one thing: it takes in text and gives back text. That's it.
It can't open a file. It can't call an API. It doesn't remember what happened ten minutes ago. It has no idea whether its own answer is right.
So how do tools like coding agents actually do things, like reading your codebase, running tests, and fixing bugs?
Someone built a system around the model that lets it act. That system is the harness.
A simple way to remember it:
Agent = Model + Harness
The model is the brain. The harness is the body, the rules, and the safety checks.
What's inside a harness?
There's no single standard, but most harnesses are made of the same building blocks:
| Part | What it does |
|---|---|
| 📋 Instructions | The rules, goals, and context the agent starts with |
| 🔧 Tools | What it's allowed to use: APIs, files, search, a code runner |
| 🧠 Memory | What it keeps between steps or sessions |
| 🔁 Control loop | Plan → act → look at the result → decide the next step → stop |
| 🛡️ Guardrails | What it's not allowed to do without permission |
| ✅ Checks | Tests and validation that verify its output |
| 📊 Logs | A record of every step it took |
| 🙋 Human review | A person approves risky actions before they happen |
None of these are the model. All of them decide how well the model performs.
A simple example
Imagine you ask an agent: "Fix the failing test in this project."
Without a harness, the model can only guess. It writes some code that looks right and hopes for the best.
With a harness, the flow looks like this:
- The instructions tell it how the project is structured and which commands to use.
- A tool lets it read the test file and the code being tested.
- Another tool lets it run the test and see the actual error.
- It edits the code, then the check runs the test again.
- If it still fails, the loop sends it back to try again, up to a limit.
- Before it changes anything risky, a guardrail asks you for approval.
- Every step is saved in the logs, so you can see exactly what happened.
Same model. Completely different result.
Why the harness matters so much
Here's what surprised me most when I started reading about this: two teams can use the exact same model and get very different results.
That's because the harness controls:
- What the model sees. Good context in, good decisions out.
- What the model can do. The right tools, nothing more.
- How mistakes get caught. Checks catch errors before they reach you.
- When it stops. A clear end point keeps time and cost under control.
Models are getting better and more similar every few months. The harness is where teams can really make their agent stand out.
What happens if you ignore it
If you plug a powerful model into a weak harness, you get a powerful but unreliable agent. Common problems:
- ❌ Endless loops: the agent repeats the same step because nothing tells it to stop
- ❌ Lost context: it forgets earlier decisions halfway through a long task
- ❌ Unchecked errors: wrong answers go straight through with nothing verifying them
- ❌ Risky actions: it deletes, sends, or changes things with no approval step
- ❌ Silent failures: something breaks, and with no logs you can't tell why
- ❌ Rising costs: every extra loop is another paid model call
A better model can make an agent smarter. Only a better harness makes it reliable.
If you're building an agent, start here
You don't need a complex setup to begin. Ask yourself these five questions:
- Does my agent know the rules? Write clear instructions and project context.
- Does it have only the tools it needs? Fewer, well-defined tools beat many vague ones.
- How does it check its own work? Add at least one test or validation step.
- When does it stop? Set a step limit or a clear finish condition.
- Can I see what it did? Log every action.
Get these right, and even a smaller model can do surprisingly well.
Final thought
We spend a lot of time comparing models: which one is smarter, faster, cheaper. That matters.
But the next time an agent lets you down, look at the harness before you blame the brain.
What's the biggest problem you've hit while building or using an AI agent? Tell me in the comments 👇
I write about how tech actually works at Compiled Thoughts on Substack. If this was useful, follow me here or subscribe there for more breakdowns like this.
Further reading:
Top comments (0)