DEV Community

What Is an Agent Harness? Harness Engineering Explained

Tejas Kumar on September 14, 2026

An agent harness is everything around the model that gives it grounding in reality: the tools it can call, the context it sees, the guardrails that...
Collapse
 
raju_dandigam profile image
Raju Dandigam •

@tejas_kumar_83c520d6bef27, the six-part split is helpful, especially treating verification as a loop around the agent loop. I’d add an explicit outcome state between tool execution and verification—confirmed, failed, or ambiguous—so a retry cannot duplicate a side effect merely because the trace lacks proof. In agent-inspect I’m working on that evidence boundary for TypeScript runs; do you expect the harness to own idempotency keys, or should each tool adapter expose its own retry-safe contract?

Collapse
 
tejas_kumar_83c520d6bef27 profile image
Tejas Kumar •

I’m not sure, I think it depends on the use case. This post was more to orient developers on a basic harness and understand the fundamentals. From here, there are so many places we can take this!

Collapse
 
kaziava profile image
Hardcore Engineer •

This resonates hard — the "model claims success" problem hit us exactly this way,
just quieter. Our GraphRAG pipeline would return entity graphs that looked valid
but were actually dumber because Ollama silently truncated 6k-token extraction
prompts against a 2-4k default window. No exception, no HTTP error, no contract
violation I could catch programmatically: the model just said "done" and returned
a graph missing 40% of the relationships.

The fix wasn't a better prompt. It was a boring guardrail: an independent
verifier (separate process, its own timeout) that logged token counts in/out on
every call and alerted when context headroom crossed 80%. The verifier never saw
the generator's reasoning — only output metadata. Rule since then: any limit a
service enforces silently is a bug in its docs, not in your client.

Your verify step reading tool history is interesting. We struggled with cases
where "done" isn't observable from trace alone (like "did the model actually
understand the table structure?"). Do you keep verification purely deterministic,
or do you have cases where you escalate to a stronger model to verify the weaker
one's work?

Collapse
 
tejas_kumar_83c520d6bef27 profile image
Tejas Kumar •

i wish at least 1 comment is human written here

Collapse
 
kaziava profile image
Hardcore Engineer •

ha, fair call 😅 I did let an LLM tidy up my comment — which is painfully
on-theme for an article about not trusting model output. The incident itself
is mine though: prod GraphRAG pipeline, Ollama silently truncating our
6k-token extraction prompts into a 4k window for two days, no errors, just
dumber entity graphs. We blamed "model quality" until someone finally logged
token counts. Point taken: next comment I write myself, typos included.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

A model saying it's done is a claim, not evidence. That's basically a rule my own agent work already runs on - never log an action as finished off a proxy signal. Check the real state after. One thing worth pushing on: the verify step in branch 3 is string matching against event.result. Still a heuristic, just written by a human instead of a model. Break the "now at" match with a UI copy change on a run that actually succeeded, and does the harness read that as a false failure, or just retry blind?

Collapse
 
jo-do profile image
Jo Do •

'Everything around the model that gives it grounding in reality' is the definition I'd defend. The reported-success-from-a-login-page story is exactly the failure that makes the harness a discipline: the model will narrate completion whether or not reality agreed. The harness is where 'the tool returned 200' gets separated from 'the thing happened'. Once burned, you stop writing agent loops and start writing verification loops that happen to contain a model.

Collapse
 
tejas_kumar_83c520d6bef27 profile image
Tejas Kumar •

:(

Collapse
 
unitbuilds profile image
UnitBuilds •

Actually have had quite alot of fun experimenting with harnesses. I did a series a few months ago that started with a basic harness to use Cloudflare workers ai in vs code as an extension, then built the guardrails and quality assurance layers, then expanded into a minimal egui IDE, then just went full out and turned it into a standalone OS 😂 IDE is live on GitHub, though it's definitely not a toy anymore... OS is getting there, who'd have thought putting a custom quantization and a homebrew harness directly into the kernel of a Ring-0 SASOS would be difficult to get working right?

Worth reading Pascal's post that started it all, he had tried various AI against a benchmark he made, then he tried Kimi K2.7, fastest, best, cheapest... Completely unusable in production, because it followed the task to the letter, not considering production safety like not sending plain-text passwords... That's the exact case where a harness shows it's value, when it can course correct a LLM to making 'smarter' choices, instead of just 'optimal' choices.

Collapse
 
raknaos profile image
Raknaos •

Splitting evaluation harness from agent harness is genuinely useful, because the two communities say "harness" to mean a benchmark runner and a runtime and keep designing for the wrong one. Your definition by job — grounding a model you don't control in an environment you do — is the one I'd keep, and the tool registry / context / guardrails / verify decomposition maps cleanly onto what actually breaks in production.

The verify step is where I'd ask a follow-up: in your demo it reads the tool history and confirms the upvote happened, which is code doing the checking. I keep seeing people implement the verification as another model call, and then the checker inherits exactly the same confidence problem as the checked. Do you treat the verify step as strictly deterministic, and if so how do you handle jobs where "done" isn't observable from tool history alone?

Collapse
 
tejas_kumar_83c520d6bef27 profile image
Tejas Kumar •

did claude write this comment or did you

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

Six parts: tool registry, model, context management, guardrails, agent loop, verify step. That verify step is where most tutorials stop short.