DEV Community

Cover image for Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation
Karthik Bommineni
Karthik Bommineni

Posted on

Jev as a Tool Router: Cutting Agent Cost Without Killing the Investigation

By now, you have probably heard about Jev, a System 1 model that has been getting a lot of attention lately. In simple terms, a System 1 model is built for fast, cheap, bounded decisions, while a System 2 model is the slower, heavier LLM that reasons through open-ended work (the big 3: ChatGPT, Claude, Gemini (or even Grok)). I will not dig into Jev's architecture in this post; that is probably a topic for another day. This blog is about one practical use case for Jev in your agents: routing the next tool call so you can save tokens, context window, and cost, while keeping latency low.

Most AI agents pick the next tool the expensive way

Every turn, they send the full tool catalog into a large language model: every name, every description, every input field. The model looks at the menu, picks a tool, and the loop continues. That works when you have twenty tools. It gets painful when you have fifty, a hundred, or two hundred.

Three things go wrong as the menu grows.

  1. Cost. The same giant menu is re-sent on every hop. You are not paying once for the catalog. You are paying for it again and again.
  2. Latency. More tokens in means a slower decision, every time.
  3. Accuracy. Published benchmarks and a lot of practitioner lore say tool-picking gets worse once the menu crosses roughly forty or fifty tools, and worse again at hundreds.

So I ran a small, honest experiment.

Question: if I pull the "which tool next?" decision out of the big LLM and hand it to Jev, can I keep the investigation quality while cutting cost as the catalog scales?

I used Moonshot Kimi K3 for the first pass because it is OpenAI-compatible on OpenRouter, so the same tool-calling loop works without a special integration, and it was cheap enough to iterate on. Many enterprise agents sit on OpenAI models, so I also ran the same full-menu loop with GPT-6 Astra (openai/gpt-6-astra) on the same 50 / 100 / 200 tasks.

The code, catalogs, tasks, and raw result JSON are here:

https://github.com/karthik-bommineni/tool-routing-experiment-with-jev

What about using another agent as the router?

A lot of teams already do a version of this with multi-agent setups: one agent decides which tools (or which specialist agent) should run next, and another agent does the work. That can save context on the main worker, because the worker no longer has to carry the full tool menu every turn.

But it only helps if the routing agent is not also eating the whole catalog. If your "router" is still a full System 2 LLM that sees every tool schema on every hop, you mostly moved the same bill to a different call. You did not remove it.

That is why a System 1 model like Jev is interesting in that slot. Keep the expensive LLM for argument filling and writing. Use something cheap and fast for the bounded choice of "what comes next?"

What I measured

I did not build a production agent. I measured one decision inside a multi-hop investigation loop: given the task and the facts gathered so far, which tool should run next?

Three conditions. Same tasks. Same catalogs. Same fake tool results. Same golden checks.

Condition A — Kimi full-menu router

Moonshot Kimi K3 (via OpenRouter) sees the entire tool menu on every hop, with tool_choice set to auto. It can call several tools in one reply, or stop and write the incident note.

Condition B — Jev router

TypeSafe Jev 1.13 (via OpenRouter's Decisions API) sees the task plus the facts so far, plus a Choice menu of every tool and a special finish option. It returns exactly one name. Then Kimi sees only that one tool's schema and fills the arguments. When Jev picks finish, Kimi writes the note with no tools.

Condition C — Astra full-menu router

OpenAI GPT-6 Astra (via OpenRouter) uses the same full-menu loop as Condition A. This is the enterprise-shaped baseline: same prompts, same tools, more expensive model.

I did not call real GitHub, Kubernetes, or Grafana. Every tool result is fake but consistent, so every router plays in the same world.

The catalogs and the tasks

The menus are real MCP tool definitions, nested so the smaller sets sit inside the larger ones.

Size Catalog
50 tools GitHub MCP tools
100 tools those 50 plus Kubernetes MCP tools
200 tools those 100 plus Grafana MCP tools

Each catalog size has one investigation-style task, written like something an on-call engineer would actually ask.

  • 50: checkout latency after a recent change. Find the open checkout issue, the commit that changed the timeout, the file contents, the open PR, CI status, and secret scanning. Comment suspect abc1234 on the checkout issue only. Write an incident note.
  • 100: continue in the payments namespace. List checkout pods, read logs of the not-ready pod, list warning events, get the Deployment. Decide whether the cluster symptom matches a 200ms client timeout. Write an incident note.
  • 200: close with metrics. Query Prometheus for checkout p95, Loki for deadline-exceeded logs, list the firing alert rule and OnCall alert group. Decide true positive and whether to roll back the timeout. Write an incident note.

Golden scoring checks the final note and important side effects (comments posted, dangerous tools avoided). Hop count is logged but is not part of the score. That matters, because Jev is forced to one tool per lap while full-menu LLMs can batch.

The numbers

Scoreboard

Catalog Kimi full menu Jev + Kimi Astra full menu
50 Strict golden pass Strict golden fail (extra comment); note and routing OK Strict golden pass
100 Strict golden pass Strict golden pass Strict golden pass
200 Strict golden fail (checker only); note OK Strict golden fail (checker only); note OK Strict golden fail (checker only); note OK

Cost and hops

Catalog Metric Kimi Jev + Kimi Astra
50 Hops 6 10 4
50 Prompt tokens (filler LLM) 55,734 18,866 27,638
50 Total cost $0.086 $0.086 $0.144
50 Jev cost — $0.0034 —
100 Hops 3 5 3
100 Prompt tokens (filler LLM) 55,032 5,035 43,108
100 Total cost $0.078 $0.031 $0.228
100 Jev cost — $0.0034 —
200 Hops 2 5 2
200 Prompt tokens (filler LLM) 44,215 4,197 33,162
200 Total cost $0.139 $0.031 $0.244
200 Jev cost — $0.0040 —

What that means in plain words

At 50 tools, Kimi and Jev+Kimi cost about the same: roughly nine cents. Astra also solved the task and matched golden, but cost about $0.144. Jev still routed through a sensible path; its strict golden fail was Kimi posting a second "internal note" comment after Jev chose add_issue_comment. The required suspect abc1234 comment was there. So for Jev at 50: routing pass, golden fail on a side effect.

At 100 tools, all three matched the golden answer. Kimi full menu: about $0.078. Astra full menu: about $0.228. Jev plus Kimi: about $0.031. With Jev, Kimi's prompt tokens dropped from about 55k to about 5k, because it only ever saw one tool schema at a time.

At 200 tools, the pattern is the same and clearer. Kimi: about $0.139. Astra: about $0.244. Jev plus Kimi: about $0.031, roughly a 4.5x cut versus Kimi and about an 8x cut versus Astra. Almost all of that saving comes from not stuffing 200 tool schemas into the expensive LLM on every hop. Jev's own bill stayed around $0.004.

All three 200 runs wrote a correct incident note: p95 of 4.2 seconds, 1,200 Loki matches, firing alerts, rollback yes. All three scored matched_golden false for the same boring reason. The checker looks for the substring 1200. The models wrote 1,200. That is a scoring bug in my harness, not a wrong investigation.

Astra also shows why the enterprise baseline matters. On these tasks it was accurate and fast to finish, but the per-token price made the full-menu design much more expensive than Kimi, even when Astra used fewer prompt tokens. If your production agent already sits on OpenAI, the cost of sending the whole catalog every turn is not a theoretical problem.

A few details that matter for reading this honestly

Jev returns one option per call. Full-menu Kimi and Astra often batched several tools in one reply. That is why hop counts are higher for Jev. Fewer hops does not mean "smarter." It often means "parallel tool calls."

OpenRouter can route the same model slug to different hosts at different prices. Some of my early Kimi hops looked like a $3 / $15 per million host; later hops were cheaper. Astra's OpenRouter list price is much higher. If you want a clean cost-vs-catalog curve in a follow-up, pin a provider.

I am not claiming Jev replaces the LLM. In this design, Jev only picks the next tool name. The LLM still fills arguments and writes the final note. That split is the point. Routing is a bounded choice. Argument filling and writing are not.

Observations

Everything I ran was only three tasks, one each for 50, 100, and 200 tools, across Kimi, Jev+Kimi, and Astra. I cannot claim a real accuracy number or publish proper metrics yet. That would need a much larger eval set. I would love help building that eval set. If you have some free time, feel free to contribute to the GitHub repo. There is a CONTRIBUTING.md with a starting point.

I have not finished the embedding-router condition. An early dry run showed the obvious failure mode: with a finish option in the menu and no facts gathered yet, nearest neighbor picked finish on hop one because the task text talks about writing an incident note. That experiment is paused.

A natural follow-up is Jev + Astra as the argument filler, so the enterprise model only sees one tool schema per lap instead of the full menu.

The short version

Sending the whole tool catalog to a big LLM on every hop is the simple default. It is also where a lot of the money goes once the menu gets large, especially on OpenAI-priced models.

On the same multi-hop SRE-style investigations, handing "which tool next?" to Jev and letting Kimi fill only that one tool's arguments cut total spend by about 2.5x versus full-menu Kimi at 100 tools and about 4.5x at 200 tools, while keeping the investigation quality. Against full-menu Astra, the gap is larger still. At 50 tools the Kimi vs Jev cost was a wash, and the only clear Jev miss was an extra comment, not a wrong route.

If you are building agents with growing MCP catalogs, the useful question is not "can the LLM see 200 tools?" It is "does the LLM need to see 200 tools on every turn?"

Code and results: https://github.com/karthik-bommineni/tool-routing-experiment-with-jev

Top comments (3)

Collapse
 
reidmarlow profile image
Reid Marlow •

The token drop from 55k to 5k shows the real weight of MCP schemas, but the hidden tax with single-choice routers is serialization latency. In on-call investigations, three independent reads like listing pods, tailing logs, and checking alerts can run concurrently in a full-menu model. Forcing the loop into single-tool hops means each round waits on network I/O. Letting the router emit a small batch of read-only tools while strictly serializing mutative commands gives you the schema savings without turning a three-second query into fifteen seconds of round-trips.

Collapse
 
karthikbommineni profile image
Karthik Bommineni •

That's a fair catch. This run assumed a serial loop: one tool per hop. That shows the cost win from not sending the full MCP menu every turn. For parallel reads, a good follow-up is using Jev’s probability scores to batch a few high-scoring read-only tools in one lap, while keeping writes strictly serial. Haven’t tested that yet but it's worth trying next.

Collapse
 
manogna_yamani profile image
Manogna Yamani •

This was such a cool read!!