DEV Community

Cover image for Secret Agent Overmind: diagnose an AI agent with code, traces, evals and datasets
Tyler Edwards
Tyler Edwards

Posted on Originally published at overmindlab.ai

Secret Agent Overmind: diagnose an AI agent with code, traces, evals and datasets

To diagnose an AI agent you need four records joined, not four dashboards. Read the codebase for what the agent can do, traces for what it actually did, dataset rows for what it will be trained and tested on, and eval scores for how well it scored. The answer to "why did it do that" lives in the gaps between them.

Originally published at overmindlab.ai.

Your agent is the most documented system in your engineering org and yet you can't answer simple questions.

No, Christopher McQuarrie has not asked me to direct the next instalment of Mission: Impossible, though I am keeping a keen eye on my inbox. Today, we are talking state secrets!

That is also not true. But I did spend many years building stuff and things (the OFFICIAL term) for British Intelligence, and the single most valuable thing I took away was not state secrets. Instead, it was a deeply troubling obsession with data.

"Is this person a credible threat?" is the canonical example, and a very difficult one to answer. Intelligence analysts are basically superhumans, capable of finding links between vast amounts of varied data, and one of the ways they do that is by layering context.

Data, data, everywhere.

That data is vast: signals intelligence, human intelligence, financial records, travel manifests, satellite imagery, and a mountain of open-source material. Some of it is a tidy database, some of it is a photograph, and some of it is words scrawled on a piece of paper. The scale is hard to picture: by the NSA's own figures, it "touches" around 29 petabytes of data a day, which is equivalent to the entire Library of Congress being inspected nearly 3,000 times over, daily.

I hear you shouting, "Well Tyler, big tech companies have access to lots of data, so why is MI5 so special?" Beyond the legal implications and the fact that we don't (yet) live in a Blade Runner-style dystopia, the reasons are structural. Broadly speaking, any one organisation only ever holds a single thin slice of the picture: what you searched, what you spent, or where you travelled. You can know someone's entire spending history in perfect detail and still have no idea what any of it means. An intelligence agency exists precisely to sit above all those slices and join them, which is the part almost nobody else is mandated, or built, to do. The fusion, not the volume, is the whole point.

This is the core function of an intelligence agency: not simply to collect the data, but to surface the hidden correlations in vast pools of information. Which brings us, as everything eventually does, to AI agents.

Context really is king

Your agent is the most thoroughly documented system in your engineering org. Every prompt, every completion, every tool call, every retry is captured somewhere: traces in your observability tool, scores in your eval suite, rows in your dataset. And yet you still can't answer simple questions. Why did it do that? Where did the behaviour come from? Who gave it a gun? The records are all there, sitting in separate systems, and nobody joins them up.

Each of those systems holds one thin slice, and each slice has a blind spot the next one covers.

Layer What it holds What it can tell you What it can't
Codebase The tools the agent can call and their contracts, its input and output schema, the prompts, the expected output What the agent could do Whether it ever did
Production traces Every prompt, completion, tool call and retry What the agent actually did What it was built to be able to do
Dataset rows Format, schema, per-column distribution, label space, cohorts, hygiene problems What you are about to evaluate or train on Whether that covers the agent's real surface
Eval scores Scores against your eval suite That a behaviour was wrong Why it was wrong

So we did exactly that. We built a context layer that joins all these things together: one grounded picture of how your agent is put together and what it actually does, assembled from your codebase and traces, then turned loose to power Overmind. Overmind is the model training platform for AI teams. It turns your production traces into specialised models you own, automatically trained, benchmarked and served. Getting there starts with knowing what the agent actually is.

What is a Capability Card?

We start at the codebase, because that's where the ground truth lives. Before Overmind touches a single trace, it reads the code the agent runs on and compiles a Capability Card: the tools the agent can call and their real contracts, its input and output schema, the prompts, and the expected output. That's the agent at its raw logical level, taken from the source rather than inferred from behaviour. Skip it, as most tooling does, and a trace is just a behaviour with nothing behind it, leaving you to reverse-engineer intent from whatever the trace happened to capture.

What Overmind reads from the code What it pins down
The tools the agent can call, with their real contracts The full set of actions available to it, not only the ones a trace happened to capture
Input and output schema The shape the agent is written to accept and return
The prompts The instructions in the source, not a reconstruction from behaviour
The expected output The result the code is written to expect

Then we bring in the data and the traces. Every record of what the agent did is a context source: production traces, dataset rows, eval scores, all built into the same shape. For a dataset, that layer is the spine: format, schema, per-column distribution, label space, cohorts, hygiene problems, derived straight from the data. If you have not got traces flowing yet, start there, and if the word "agent" is doing a lot of undefined work in your org, we pulled it apart here.

How do you know whether your data covers your agent?

The codebase context is what makes that data useful. With the Capability Card in hand, Overmind knows the full surface of the agent and can measure the data against it: how much of that surface the data actually covers, which behaviours are well represented, which tools never appear, and which paths the agent takes in production that no dataset row touches. A pile of traces tells you what happened. The codebase tells you what could happen, and the gap between the two is the thing you need to see. That delta is what decides whether the data is good enough to evaluate against or train on, or whether you're about to write evals for a fraction of the agent or fine-tune on the same narrow slice of its total behaviours.

The Overmind console showing the PR Approval Agent (StampHog), discovered from the posthog/posthog repository. The page lists the agent's prompt, its ci, github and security tags, and the seven models it calls, including Claude Sonnet 4.6, GPT 4.1 Mini, o3, Gemini 2.5 Flash Lite and Claude Haiku 4.5. A score panel reads 50 percent, with 1 of 2 runs traced and a badge that says eval metrics not approved. Below, a flow graph links Input, Agent and Read nodes to conversation_summarizer, title_generator, session_summarizer and llm_traces_summarizer, ending at Output.

Overmind's analysis of a PostHog Agent. The same view in text:

What the console shows For the PostHog agent
Agent PR Approval Agent (StampHog)
Discovered from The posthog/posthog repository
Tags ci, github, security
Models it calls Seven, including Claude Sonnet 4.6, GPT 4.1 Mini, o3, Gemini 2.5 Flash Lite and Claude Haiku 4.5
Score 50%
Trace coverage 1 of 2 runs traced
Eval status Eval metrics not approved
Flow graph Input, Agent and Read into conversation_summarizer, title_generator, session_summarizer and llm_traces_summarizer, ending at Output

Nothing in that card was filled in by hand. Seven models, three tags and a flow graph, read out of a public repository, with the coverage number sitting next to it so you can see how much of that surface anyone has actually watched run.

Codebase tied to data tied to traces is the working surface for everything else we do. The data preprocessing reads from it, the evals are written against it, the fine-tuning sets are curated out of it, and the agent optimisation is steered by it. That's the line between a tool that can tell you a behaviour was wrong and one that can tell you why, as well as how to fix it.

So, what do you know about your agent?

So much, maybe too much. But not in a useful way.

That is what Overmind is built for: not more raw evidence, but the relationships between it. Code, traces, datasets, evals, and training runs joined into one grounded view of the agent, so you can see what is really happening and where the gaps are.

FAQ

Why can't I tell why my AI agent did something?

Because the evidence is split across systems that never meet. Traces live in your observability tool, scores in your eval suite, rows in your dataset, and the agent's actual definition in your repository. Each one answers a different question, and no single one of them answers "why".

Are production traces enough to debug an agent?

No. A trace tells you what happened, not what could have happened. Without the codebase you cannot tell whether a behaviour was one of many available paths or the only one, which tools the agent never reached for, or which parts of its surface nobody has ever watched run.

What is agent coverage analysis?

Measuring your traces and dataset rows against the agent's full capability surface, taken from the code. It tells you which behaviours are well represented, which tools never appear, and which production paths no dataset row touches. That delta decides whether your data is fit to evaluate against or train on.

How is this different from an observability tool?

Observability records behaviour. Joining that record to the codebase, the datasets and the eval scores is what turns a wrong behaviour into an explained one, and then into a fix you can ship.

Tyler Edwards is Co-Founder and CEO of Overmind. He writes about agent infrastructure, fine-tuning, and what it takes to ship AI that actually improves in production. Connect on LinkedIn.

Overmind is the model training platform for AI teams. It turns your production traces into specialised models you own. Get started.

Top comments (0)