DEV Community

Michael Hairetis
Michael Hairetis

Posted on

I spent a week building an optimization, A/B tested it, and deleted it

I had a 25,000-token system-prompt overhead on every agent call. Obvious problem, obvious solution: build a session-reuse layer so the overhead is paid once and amortised across many calls.

So I built it. Session pooling, lifecycle management, identity tracking, the works. It was the most intricate thing in that part of the codebase and I was quietly pleased with it.

Then I A/B tested it against not having it.

No difference.

The provider's prompt cache hits on prefix content regardless of session identity. Two completely independent calls inside the cache window already got the benefit. The optimisation I had carefully engineered was buying something I already had for free, and had been getting for free the entire time I was building the thing to get it.

I deleted the whole layer.

The lesson is not "measure before optimising." Everyone says that and everyone nods and nobody does it, because the optimisation is obviously going to help.

The sharper version is this: when your semantics are stateless, write stateless code. I had reached for a stateful design to solve a problem the platform had already solved statelessly. The complexity was not buying performance, it was buying a mental model that did not match the system underneath.

Three other things from the same project, all in the same shape:

Structure beats trust. A single agent given a complex task fails unpredictably. It loses subgoals and declares victory early. Rather than trying to make one agent reliable, the pipeline does the managing and catches failures at phase boundaries.

It is a routing problem. I had one heavyweight backend answering both "which of these three pipelines" and "go do two minutes of multi-step research." Making a heavyweight agent answer a classification question costs heavyweight money for a featherweight answer. Splitting them made both paths faster and cheaper.

Old heuristics still earn their keep. Asking a model "is this task finished" costs about five cents a check. Pattern matching on the agent's own activity signals, with the model as fallback for genuinely ambiguous cases, cut average per-request cost roughly tenfold.

None of these were model problems. None were solved by a better prompt.

The full write-up is on my site, including why I think the infrastructure half is the half that actually decides whether the thing works in production.

https://openred.space/blog/the-boring-parts-of-agentic-ai.html

Top comments (0)