DEV Community

gj0xv
gj0xv

Posted on

How I cut my LLM bill by 54.9%: a local router that costs $0 per decision (open source and measured)


This week, two of the biggest names in AI shipped the same idea within days of each other. On September 15, TypeSafe launched Jev, the first "System One model": an AI that outputs typed decisions instead of text. On September 29, OpenAI answered at DevDay with the Decisions API built on Luna, their fast, cheap model for real-time decision-making.

The premise is right. The pricing usually isn't.

That is where open source matters. When the decision layer runs on your machine, with your data, at $0 per call, the economics flip: every prompt gets the right-sized model instead of the one that maximises someone else's margin. You can audit it, fork it, and turn it off when it does something stupid. No vendor slide can promise you that.

So here is mine: laya-router, the first of four open-source tools that put local, measured decisions between your app and your LLMs, built on the laya decision engine. Benchmarked with Z.ai's new GLM-5.3-flashX as the cheap tier, which is itself a fitting story: a fast, cheap model deciding which prompts it can handle itself.

54.9% estimated cost reduction
80.6% of prompts routed to the cheap tier
$0 per routing decision
460 ms warm latency (p50)
Enter fullscreen mode Exit fullscreen mode

The idea

You keep your code exactly as it is. You only change the base_url. The proxy ignores the model you asked for and picks one itself:

client = OpenAI(base_url="http://localhost:8000/v1", api_key="...")
# from here on, everything is standard
Enter fullscreen mode Exit fullscreen mode

Three layers decide:

  1. Regex fast path: trivial prompts ("hi", "translate: merci") skip the model entirely
  2. A local decision model (the laya engine, Apache 2.0, runs on CPU) classifies the prompt as simple, standard or complex, with calibrated confidence
  3. A confidence gate: uncertain prompts escalate to the frontier tier. The system is allowed to say "I don't know" and pay more

Every response carries inspection headers, so you can audit what happened and why:

X-Laya-Route: cheap
X-Laya-Model: glm-4.6-flash
X-Laya-Confidence: 0.91
X-Laya-Reason: classified simple, above threshold
Enter fullscreen mode Exit fullscreen mode

The benchmarks (all of them)

180 prompts, one cheap and one frontier model, blind LLM judge. I publish the wins and the caveats:

Measurement Result
Routed to cheap tier 80.6%
Estimated cost reduction 54.9%
Cheap win/tie rate vs frontier 79.3%
Routing latency p50 / p95 / p99 460 ms / 1.4 s / 2.7 s

Caveats, because they exist: one judge, one model pair, and the judge was noisy on trivial prompts. The backtest pipeline is committed (make backtest) with the dataset, the answers and the judgements, so you can re-run it on your own traffic. If your prompts are all hard, you will save nothing. Measure first.

What it deliberately does not do

  • No silent failover. If routing fails, you get a structured 503, not a surprise GPT-4 call on your bill
  • No rewriting. Your prompt arrives at the upstream model byte-identical. The proxy decides, it does not mutate
  • No vendor lock. Works with OpenAI, vLLM, Ollama, OpenRouter, Z.ai, anything OpenAI-shaped

Try it

pip install laya-router
Enter fullscreen mode Exit fullscreen mode

Define your tiers in YAML (models + prices), point your base_url at it, done. Docker image included, /metrics endpoint for Prometheus, JSONL decision log if you want to analyze your routing later.

The rest of the series

laya-router is one of four tools I'm building around the same idea: local, measured decisions in LLM pipelines.

  • laya-compactor: cut 70% of RAG context tokens with zero answer-quality loss (measured)
  • laya-triage: support ticket triage, fine-tuned from 51% to 90.5% intent accuracy
  • laya-phishield: explainable phishing detection, $0 per 1,000 emails

Each one ships with its benchmarks in the README, including the numbers that hurt.

If your LLM app ever died in production after a perfect demo, we have something to talk about :)

Top comments (0)