DEV Community

Cover image for We built a context system for our coding agents. They ignored it and used grep.
Kent Alstad
Kent Alstad

Posted on AI-assisted

We built a context system for our coding agents. They ignored it and used grep.

This is the first of a set of write-ups behind my Build Stuff 2026 talk, Control Your Fate at the Metal, Vilnius, 2 December. Details at the end.

A year ago we started giving our AI coding agents something better than search. The idea is the obvious one, and probably the same one you've had: a coding agent wastes an enormous amount of its context window reading files to find out what's in them. So index the repository ahead of time. Write a short description of every file. Build a call graph. Give the agent a tool that answers "which files matter for this task?" in one call instead of twenty.

We built all of it. Then we read 467 of our own agent sessions to see whether it was being used.

What the transcripts said

Over thirty days:

What we counted (30 days) Count
Times agents ran grep or rg through a shell 46,177
Times agents called our retrieval tool ~600
Sessions that called it even once 77 of 467 (16%)
Agents reached for grep about 75 times as often. Five sessions in six never touched the thing we'd built at all.

Three reasons, and none of them were "the model is dumb"

It only got used when something told it to. Nearly every call we could trace back to a cause came from a prompt, a hook or a workflow that named the tool explicitly. Tools that nothing named got almost no calls, regardless of how good they were. Adoption didn't track quality. It tracked whether some upstream instruction said the tool's name out loud.

If you've built an MCP server and wondered why nobody calls it, this is probably why. The agent sees a one-line description and chooses, right then, between a tool it has to reason about and a grep it already knows works.

Our matching rule couldn't fire on real input. Each file had a short pre-written summary, and we matched a request against it by word overlap — roughly a third of the words in the request had to appear in a 280-character summary. Real requests don't look like that. Somebody types "the login redirect is broken on mobile." The summary says "Handles OAuth callback routing and session establishment." Zero overlap, zero results. On one repository, our index put the file the agent went on to actually edit in the top five results zero times out of fifty-five.

We ranked on the summary instead of the thing it summarised. We had full descriptions of every file and were matching against the 280-character abstract of each one. When we scored both ways against real past tasks, ranking on the full text found the right file 17 times to 10.

The part that should worry you

An agent that greps isn't failing. It gets there. It opens four files instead of one, spends more of its window, takes a bit longer, and produces a reasonable answer. No error is thrown. Nothing in your logs looks wrong.

That's the trap. A retrieval system doesn't fail loudly — it gets quietly skipped while every dashboard stays green. Ours reported 100% acceptance and "14× better than baseline" during the month I just described. Both were artifacts of how we'd computed them. The honest acceptance number was 6.7%.

The "tokens saved" graph was worse. It was computed as what a naive agent would have read, minus what we returned. An empty result therefore booked almost the entire hypothetical saving. One tool claimed 16.9 million tokens saved in a month where three-quarters of its answers were empty. We were measuring an imaginary agent. The real one was grepping.

What I'd do differently

Build the thing that records what people actually ask before you build the index.

We spent a year on the supply side — descriptions, graphs, ranking — and never looked at the demand side. We had no record of what our agents were actually asked to do, so we had nothing to test a change against. Every improvement was a guess.

The fix was unglamorous: log each request alongside the files that session went on to read or edit. A backfill over ninety days gave us 5,415 requests and 1,578 request-to-file pairs. That's a test set. Now a change either moves the right file into the top five for real past tasks, or it doesn't ship.

The recorder took a fraction of the effort the index did. We should have built it first.

Talking about this at Build Stuff 2026

The context system in this article is one of the subsystems we chose to own rather than rent. What it took to make it actually work is part of the argument in my talk at Build Stuff in Vilnius this December, the 15th edition, under the theme The Age of Agency.

Control Your Fate at the Metal Wednesday 2 December, 15:00, Human OS Lab room. 45 minutes.

Your tooling depends on somebody's pricing page, rate limiter, and product roadmap. That used to be unavoidable. We went the other way — our own models, our own inference, our own registry, our own protocol implementations — until the only thing under us was hardware. Here's what it costs, what it buys, and where the metal pushes back.

If you're coming, say hello. If you're not yet, 2026SPEAKER_20 takes 20% off a ticket.

BuildStuff15 @buildstuffconf https://buildstuff.events

We used agents to draft this article and document this project as we built.

Top comments (1)

Collapse
 
indiainfranotes profile image
IndiaInfraNotes •

The 100% acceptance vs 6.7% honest number is the most important line here. Silent skipping is the agent version of a green dashboard over a dead service, and it only shows up when you log every tool choice per session, not just tool results. Did you try renaming the retrieval tool in the system prompt as the default first step, and did adoption move without changing the index at all?

iin1005h1228