For about a year I thought observability was the finish line. I had traced everything. Every model call, every tool call, every retrieval showed up as a span, and I could open any request and see exactly what my agent did, in order. It felt like control.
Then I looked at what all that tracing had actually changed, and the answer was close to nothing. I was capturing roughly a hundred failing conversations a week and turning maybe three of them into something that stopped the bug from coming back. The other ninety-seven sat in storage with a nicer interface. My agent was fully observable and getting no better.
A trace records what happened, not whether it was any good
This was the part I had backwards. A trace is a faithful recording of behavior. It tells me the model was called, the tool looped three times, the retriever returned these documents, the whole turn took 800 milliseconds and came back a 200. What it does not tell me is whether the answer was right. A wrong answer returns a 200 too. It gets filed next to a thousand correct runs with no label on it, and nothing about having recorded it makes the next version of my agent avoid it.
Monitoring told me the agent responded in 800 milliseconds with a 200. Observability told me it responded by retrieving a stale document and summarizing it wrong, with total confidence. Both were true. Neither one scored the answer, and that missing score was the whole problem.
The missing piece was a score, and then a loop
The thing that turns an inert trace into a signal is a score. The moment I attach a quality label to a trace, from a rule, a judge, or a human review, I can do something with it. So I stopped treating observability as the end and started treating it as step one of a loop.
The loop has six steps, and it is a circle, not a line. Observe: capture every request as a trace. Evaluate: score each trace against a rubric, a deterministic rule, or an LLM judge, so every run gets a quality label. Identify: filter the ones that scored below my threshold. Curate: turn those failing traces into labeled eval cases, each with the input, the output, and why it failed. Improve: sharpen the prompt, repair the context, fix the harness, or change the model. Deploy behind a regression gate: re-run that dataset on every change and block the deploy if the pass rate drops. Then step six feeds back into step one, and the failure I caught once is what stops it recurring.
The step everyone skips is the gate
For a long time I did the first five steps and skipped the sixth. I would find a failure, fix it, feel productive, and move on. A few weeks later the same class of bug was back, because nothing in my pipeline remembered that I had already fixed it.
The regression gate is what remembers. I keep the failing cases I have already fixed in a dataset, re-run the whole set on every change, and fail the build if the pass rate drops below a bar. The bar is rarely 100 percent, which is too brittle to live with. I set it around 95, which still means a change that breaks five cases I already fixed cannot merge without me seeing it. It is the cheapest insurance I have, and it is the piece I ignored the longest.
Real failures beat the ones you invent
The best eval cases I have are not hand-written. They are production traces that already failed. A synthetic test set is my guess about what might go wrong. A failing trace is a thing that actually went wrong, on a real request, with the exact input that broke it. Converting those into eval cases is the highest-value move in the whole loop.
The trap is volume. If I promote every single failure, my dataset turns into copy after copy of the same bug and I drown in it. So I group near-identical failures and keep a representative few of each kind. The dataset should cover distinct failure modes, not count incidents. One good example of a failure beats a pile of duplicates of it.
I track two numbers, and only together
A loop you cannot measure is a loop you will quietly abandon, so I watch two numbers. The first is loop yield: of all the failing traces I captured, how many actually became eval cases. If I am logging a hundred failures a week and converting three, the loop is decorative, which was exactly where I started. The second is regression-gate pass rate: of the cases in my dataset, how many the current agent passes.
Read either one alone and it will lie to you. High yield with a falling pass rate means I am taking on debt faster than I am fixing it. A high pass rate with near-zero yield means I am passing a tiny stale dataset and calling it confidence. Read together, they tell me whether the agent is actually learning from what it got wrong.
The mistakes that kept my agent from improving
- I treated capture as the goal. Tracing everything felt like progress. It was storage until I scored it.
- I skipped the regression gate. Fixing a bug without a gate just means fixing it again next month.
- I promoted every failure. The dataset bloated with duplicates and I stopped trusting it. Dedupe to distinct modes.
- I set the gate at 100 percent. Too brittle, so it blocked everything, so I turned it off, which was worse. 95 held.
- I measured nothing. Without loop yield I had no idea that ninety-seven of every hundred failures were evaporating.
The lesson I keep coming back to is that observability is necessary and nowhere near sufficient. Capturing what the agent did is step one. The value is in the five steps after it, the ones that turn a recorded failure into a test the next deploy has to pass. If you want the full version, with the span-capture code and the exact curation step I skipped for too long, I wrote it all up in this piece.
If you run an agent in production, I am curious which step you skip. For me it was the regression gate, and I paid for it by fixing the same bug more than once before it stuck.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.