DEV Community

Cover image for I Compared 5 LLM Gateway Tools for Real-World Production Use

I Compared 5 LLM Gateway Tools for Real-World Production Use

Harsh Raval on September 23, 2026

When an LLM application starts handling real users, calling a model API is usually easy. The problems show up around it. One provider hits a rate l...
Collapse
 
hayrullahkar profile image
Hayrullah Kar •

The dimension I would add to this comparison is what each gateway records per attempt, not per request. Retries and provider fallbacks are the feature people buy a gateway for, and they are also what quietly doubles a bill: one answer that failed over once is two billed calls, and a log keyed by request shows it as one. If the usage view cannot break out attempts, cost per successful answer is unknowable at exactly the point it starts to matter.

Related, and the reason I ended up writing a thin layer of my own instead of adopting one: normalising the response shape across providers is the easy half. The hard half is that token accounting differs. Cached input, reasoning tokens and image tokens are not counted the same way everywhere, so a single normalised "tokens" field in a dashboard is an average of different things. Worth checking which of the five expose the provider's raw usage object alongside their own.

One operational question: for the self-hosted options, does the gateway's own failure mode get measured anywhere in your setup? It becomes a new single point in front of every model call, and its timeout budget has to be shorter than the caller's or the retries stack.

Collapse
 
devstackhub profile image
Harsh Raval Dev Stack Community •

This is a gap in the post. I listed retry frequency, fallback frequency and cost per successful request as things to measure, but I treated them as aggregate metrics and never asked whether each gateway can break them out per attempt. A log keyed by request hides exactly the case you describe, one answer that failed over once and was billed twice. If the usage view can't separate attempts, cost per successful answer can't be computed, so the metric I recommended is only usable on tools that record at that granularity. I'll add attempt-level logging as its own evaluation criterion.

The token accounting point is also fair. A single normalised tokens field mixes cached input, reasoning and image tokens, which providers count differently, so the dashboard number is an average of different things. I haven't verified which of the five expose the provider's raw usage object next to their own, and I don't want to guess from docs. That check will go in the follow-up.

Your last question is one the post skipped entirely. I compared features and operational fit, but I never covered the gateway's own failure mode. For the self-hosted options it becomes a single point in front of every model call, and if its timeout budget isn't shorter than the caller's, retries stack. That's a real production risk, and it belongs in the failure-scenario tests alongside provider outages and rate limits.

Thanks for the detailed comment. This is what I'd want the next comparison to be built around.

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

Useful spread of tools. The dimension these comparisons usually skip is streaming behaviour under concurrency β€” some gateways buffer SSE chunks, which wrecks time-to-first-token even when raw throughput looks fine. How did the five handle queuing when the upstream model server saturates? That's where a gateway differs most from pointing clients straight at the model server.

Collapse
 
devstackhub profile image
Harsh Raval Dev Stack Community •

Fair point, and I have to be upfront: the post doesn't answer this. I listed streaming performance (time to first token, interruptions, completion time) as something to measure, but I didn't run a concurrency benchmark across the five, so I can't tell you how each one behaves when the upstream saturates. I'd rather say that than hand-wave from docs.

SSE buffering is the part I'd check first, because it can hide behind healthy throughput numbers. Time to first token is the only metric that exposes a gateway holding chunks back. The gateways built on NGINX-style proxies (Kong and APISIX) are where I'd look at buffering configuration specifically, but that's a thing to verify per gateway, not a claim I can make from testing.

Your queuing question also splits into a few distinct behaviours that I think deserve separate test cases:

  • Does the gateway queue requests, shed them, or pass the upstream 429 straight through?
  • If it queues, is there a bounded queue with a timeout, or can it grow until memory or client timeouts break it?
  • How does a fallback interact with saturation? Failing over from a saturated provider can just move the overload somewhere else.
  • Are queued requests visible in the metrics, or do they only show up as inflated latency?

That last point ties back to your comparison with pointing clients straight at the model server. Direct connections at least fail loudly, while a gateway that queues quietly can turn an overload into slow, hard-to-diagnose latency.

I'm planning a follow-up around streaming and failure scenarios, and this gives me a much sharper test plan for it. Thanks for the push.

Collapse
 
kira_m_b9e617f21ee9b4fbe profile image
Jane Jones •

This is a thoughtful comparison because it moves beyond feature checklists and focuses on what actually changes when LLM infrastructure meets production traffic. I especially liked the emphasis on retries, fallbacks, streaming, and the hidden cost of reliability mechanismsβ€”availability improvements can introduce latency and additional token spend. The distinction between gateway performance in a controlled benchmark and behavior under realistic workloads is also particularly valuable. The recommendation to measure cost per successful request alongside latency and error rates gives the comparison a much more practical operational perspective. Ultimately, the article makes a strong case that choosing an LLM gateway is as much an architectural decision as it is a tooling decision.

Collapse
 
devstackhub profile image
Harsh Raval Dev Stack Community •

Thanks for the kind words, glad the cost-versus-reliability angle came through. That was the point I most wanted to make: retries and fallbacks improve availability, but they aren't free in latency or tokens, and a raw throughput number hides that completely.
Cost per successful request is the metric I'd push hardest, because it's where a gateway's reliability features show up on the bill. I'd also treat streaming as its own test (time to first token, mid-stream drops) rather than folding it into a single latency figure. Several readers have asked for a dedicated streaming and fallback-cost follow-up, and I'm planning one.
And you're right that the gateway choice is really an architectural decision. The best fit depends far more on your existing infrastructure and provider mix than on any feature list.
Thanks for reading and taking the time to comment!

Collapse
 
kartik-nvjk profile image
Kartik N V J K •

I run a gateway in front of my eval traffic and the part nobody mentions is that failover changes your results. A retry against a different provider mid-benchmark quietly shifts the distribution, so my golden set agreement drops for reasons that have nothing to do with the model. I ended up pinning eval runs to a single provider and only letting the gateway route production traffic. Did you test the five with identical traffic replays, or fresh requests each time?

Collapse
 
devstackhub profile image
Harsh Raval Dev Stack Community •

Good point, and it's a failure mode I didn't cover. Fallback quietly changes what you're measuring, because the model that answered isn't the model you thought you were scoring. Pinning eval runs to one provider and letting the gateway route only production traffic is a sound way to handle it.

To answer your question straight: the post is an architectural comparison, not a benchmark write-up, so I don't have replay-versus-fresh results to report. It says to test with representative traffic, but it doesn't say how to keep runs comparable, and identical replays are the right answer to that. Fresh requests each time add prompt variance on top of the gateway differences, so you can't tell which one moved your numbers.

Your comment also suggests a few things I'd add to the follow-up:

  • Record which provider and model actually served each response. A per-request field would let you see when a fallback fired during a run.
  • Treat a fallback during an eval run as a flagged event. Either exclude those samples or fail the run, so a silent failover can't shift your golden-set agreement.
  • Replay the same recorded traffic through each gateway, with identical failure injection, so differences come from the gateway and not from the inputs.

Thanks for raising this. It's a use case the "gateway for production traffic" framing in the post missed.

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

file:/tmp/opencode/outbound/c2.txt

Collapse
 
julianneagu profile image
Julian Neagu •

The retry layer is where this gets interesting. I’ve seen a single timeout turn into 2–3 provider calls pretty quickly. Logging only the final request makes that cost almost impossible to spot.

Collapse
 
devstackhub profile image
Harsh Raval Dev Stack Community •

Exactly, and the timeout case is the sneaky one. A timeout doesn't mean the provider didn't process the request, so you can pay for the abandoned attempt and then pay again for the retry. If the app SDK, the gateway, and a fallback policy each retry independently, one slow response can multiply into several billed calls, and the final log line shows a single "success."

The fix I'd push for is logging per attempt, not per request: attempt number, provider, model, latency, outcome, and tokens, all tied to one parent request ID. Then "retries per successful request" and "cost per successful request" become simple queries instead of guesses.

I only covered this at the metric level in the post, and I didn't compare how granularly each of the five records attempts. That would make a good follow-up, since it's a real differentiator once retries stack up. Have you found a gateway that handles this well, or did you end up adding your own attempt-level logging?

Collapse
 
michaeljohnsondz profile image
Michael Johnson •

The fallback discussion caught my attention. It's easy to think of fallbacks as a free reliability feature, but the extra latency and token usage can add up quickly.

I'd be interested to see a follow-up comparing different fallback strategies, especially primary provider failure vs rate-limit failure vs timeout.

Collapse
 
jennifer-smith profile image
Jennifer Smith •

I like that you included LiteLLM, Kong, Portkey, Bifrost, and APISIX because they approach the problem from slightly different infrastructure perspectives.

The part I'd be most interested in testing is streaming. A gateway can look great in a normal request benchmark, but streaming introduces things like time to first token, connection handling, and interruptions.

Did you consider making streaming performance a separate benchmark in a future comparison?

Collapse
 
alinashah profile image
Alina Shah •

The point about not judging an LLM gateway purely by requests per second is really important. In production, p95/p99 latency, retries, fallbacks, and provider failures can matter much more than a clean benchmark.

I'm curious how other teams decide when a gateway is actually worth adding versus keeping provider logic directly inside the application.

Collapse
 
mayur-upadhyay profile image
Mayur Upadhyay •

One thing I'd add: the cost side of fallbacks deserves its own deep dive. A fallback that "saves" the request can quietly double token spend if the secondary provider has a different pricing tier or context window, and that's easy to miss until the bill shows up. Would be curious if you've tracked cost per successful request across these five in a real workload, not just latency and error rate.

Also agree with the comments about streaming. Time to first token and how a gateway handles a dropped stream mid-response feels like it deserves a benchmark of its own, separate from the standard request/response numbers.

Collapse
 
debugtodeploy profile image
Vinay Shah •

Cost per successful request is the metric I'd want to see broken out per gateway, not just mentioned as a category. Fallbacks especially can look like a reliability win while quietly doubling spend if the secondary provider has a different pricing tier, and that gap only shows up once you're looking at the bill instead of the logs.

Streaming is the other piece worth a dedicated pass. Time to first token, connection handling, and mid-stream interruptions behave nothing like a normal request/response cycle, so lumping them into the same latency numbers probably hides more than it reveals.

Collapse
 
rafidbottler profile image
Rafid Bottler •

The fallback cost point is worth digging into further. A fallback that "saves" the request can quietly double token spend if the secondary provider has a different pricing tier or context window, and that's easy to miss until the bill shows up. Would be curious if you've tracked cost per successful request across these five in a real workload, not just latency and error rate.

Streaming feels like the bigger gap though. Time to first token and how a gateway handles a dropped stream mid-response deserves a benchmark of its own, separate from standard request/response numbers, especially since two other commenters already flagged it.