When an LLM application starts handling real users, calling a model API is usually easy. The problems show up around it. One provider hits a rate l...
For further actions, you may consider blocking this person and/or reporting abuse
The dimension I would add to this comparison is what each gateway records per attempt, not per request. Retries and provider fallbacks are the feature people buy a gateway for, and they are also what quietly doubles a bill: one answer that failed over once is two billed calls, and a log keyed by request shows it as one. If the usage view cannot break out attempts, cost per successful answer is unknowable at exactly the point it starts to matter.
Related, and the reason I ended up writing a thin layer of my own instead of adopting one: normalising the response shape across providers is the easy half. The hard half is that token accounting differs. Cached input, reasoning tokens and image tokens are not counted the same way everywhere, so a single normalised "tokens" field in a dashboard is an average of different things. Worth checking which of the five expose the provider's raw usage object alongside their own.
One operational question: for the self-hosted options, does the gateway's own failure mode get measured anywhere in your setup? It becomes a new single point in front of every model call, and its timeout budget has to be shorter than the caller's or the retries stack.
This is a gap in the post. I listed retry frequency, fallback frequency and cost per successful request as things to measure, but I treated them as aggregate metrics and never asked whether each gateway can break them out per attempt. A log keyed by request hides exactly the case you describe, one answer that failed over once and was billed twice. If the usage view can't separate attempts, cost per successful answer can't be computed, so the metric I recommended is only usable on tools that record at that granularity. I'll add attempt-level logging as its own evaluation criterion.
The token accounting point is also fair. A single normalised tokens field mixes cached input, reasoning and image tokens, which providers count differently, so the dashboard number is an average of different things. I haven't verified which of the five expose the provider's raw usage object next to their own, and I don't want to guess from docs. That check will go in the follow-up.
Your last question is one the post skipped entirely. I compared features and operational fit, but I never covered the gateway's own failure mode. For the self-hosted options it becomes a single point in front of every model call, and if its timeout budget isn't shorter than the caller's, retries stack. That's a real production risk, and it belongs in the failure-scenario tests alongside provider outages and rate limits.
Thanks for the detailed comment. This is what I'd want the next comparison to be built around.
Useful spread of tools. The dimension these comparisons usually skip is streaming behaviour under concurrency β some gateways buffer SSE chunks, which wrecks time-to-first-token even when raw throughput looks fine. How did the five handle queuing when the upstream model server saturates? That's where a gateway differs most from pointing clients straight at the model server.
Fair point, and I have to be upfront: the post doesn't answer this. I listed streaming performance (time to first token, interruptions, completion time) as something to measure, but I didn't run a concurrency benchmark across the five, so I can't tell you how each one behaves when the upstream saturates. I'd rather say that than hand-wave from docs.
SSE buffering is the part I'd check first, because it can hide behind healthy throughput numbers. Time to first token is the only metric that exposes a gateway holding chunks back. The gateways built on NGINX-style proxies (Kong and APISIX) are where I'd look at buffering configuration specifically, but that's a thing to verify per gateway, not a claim I can make from testing.
Your queuing question also splits into a few distinct behaviours that I think deserve separate test cases:
That last point ties back to your comparison with pointing clients straight at the model server. Direct connections at least fail loudly, while a gateway that queues quietly can turn an overload into slow, hard-to-diagnose latency.
I'm planning a follow-up around streaming and failure scenarios, and this gives me a much sharper test plan for it. Thanks for the push.
This is a thoughtful comparison because it moves beyond feature checklists and focuses on what actually changes when LLM infrastructure meets production traffic. I especially liked the emphasis on retries, fallbacks, streaming, and the hidden cost of reliability mechanismsβavailability improvements can introduce latency and additional token spend. The distinction between gateway performance in a controlled benchmark and behavior under realistic workloads is also particularly valuable. The recommendation to measure cost per successful request alongside latency and error rates gives the comparison a much more practical operational perspective. Ultimately, the article makes a strong case that choosing an LLM gateway is as much an architectural decision as it is a tooling decision.
Thanks for the kind words, glad the cost-versus-reliability angle came through. That was the point I most wanted to make: retries and fallbacks improve availability, but they aren't free in latency or tokens, and a raw throughput number hides that completely.
Cost per successful request is the metric I'd push hardest, because it's where a gateway's reliability features show up on the bill. I'd also treat streaming as its own test (time to first token, mid-stream drops) rather than folding it into a single latency figure. Several readers have asked for a dedicated streaming and fallback-cost follow-up, and I'm planning one.
And you're right that the gateway choice is really an architectural decision. The best fit depends far more on your existing infrastructure and provider mix than on any feature list.
Thanks for reading and taking the time to comment!
I run a gateway in front of my eval traffic and the part nobody mentions is that failover changes your results. A retry against a different provider mid-benchmark quietly shifts the distribution, so my golden set agreement drops for reasons that have nothing to do with the model. I ended up pinning eval runs to a single provider and only letting the gateway route production traffic. Did you test the five with identical traffic replays, or fresh requests each time?
Good point, and it's a failure mode I didn't cover. Fallback quietly changes what you're measuring, because the model that answered isn't the model you thought you were scoring. Pinning eval runs to one provider and letting the gateway route only production traffic is a sound way to handle it.
To answer your question straight: the post is an architectural comparison, not a benchmark write-up, so I don't have replay-versus-fresh results to report. It says to test with representative traffic, but it doesn't say how to keep runs comparable, and identical replays are the right answer to that. Fresh requests each time add prompt variance on top of the gateway differences, so you can't tell which one moved your numbers.
Your comment also suggests a few things I'd add to the follow-up:
Thanks for raising this. It's a use case the "gateway for production traffic" framing in the post missed.
file:/tmp/opencode/outbound/c2.txt
The retry layer is where this gets interesting. Iβve seen a single timeout turn into 2β3 provider calls pretty quickly. Logging only the final request makes that cost almost impossible to spot.
Exactly, and the timeout case is the sneaky one. A timeout doesn't mean the provider didn't process the request, so you can pay for the abandoned attempt and then pay again for the retry. If the app SDK, the gateway, and a fallback policy each retry independently, one slow response can multiply into several billed calls, and the final log line shows a single "success."
The fix I'd push for is logging per attempt, not per request: attempt number, provider, model, latency, outcome, and tokens, all tied to one parent request ID. Then "retries per successful request" and "cost per successful request" become simple queries instead of guesses.
I only covered this at the metric level in the post, and I didn't compare how granularly each of the five records attempts. That would make a good follow-up, since it's a real differentiator once retries stack up. Have you found a gateway that handles this well, or did you end up adding your own attempt-level logging?
The fallback discussion caught my attention. It's easy to think of fallbacks as a free reliability feature, but the extra latency and token usage can add up quickly.
I'd be interested to see a follow-up comparing different fallback strategies, especially primary provider failure vs rate-limit failure vs timeout.
I like that you included LiteLLM, Kong, Portkey, Bifrost, and APISIX because they approach the problem from slightly different infrastructure perspectives.
The part I'd be most interested in testing is streaming. A gateway can look great in a normal request benchmark, but streaming introduces things like time to first token, connection handling, and interruptions.
Did you consider making streaming performance a separate benchmark in a future comparison?
The point about not judging an LLM gateway purely by requests per second is really important. In production, p95/p99 latency, retries, fallbacks, and provider failures can matter much more than a clean benchmark.
I'm curious how other teams decide when a gateway is actually worth adding versus keeping provider logic directly inside the application.
One thing I'd add: the cost side of fallbacks deserves its own deep dive. A fallback that "saves" the request can quietly double token spend if the secondary provider has a different pricing tier or context window, and that's easy to miss until the bill shows up. Would be curious if you've tracked cost per successful request across these five in a real workload, not just latency and error rate.
Also agree with the comments about streaming. Time to first token and how a gateway handles a dropped stream mid-response feels like it deserves a benchmark of its own, separate from the standard request/response numbers.
Cost per successful request is the metric I'd want to see broken out per gateway, not just mentioned as a category. Fallbacks especially can look like a reliability win while quietly doubling spend if the secondary provider has a different pricing tier, and that gap only shows up once you're looking at the bill instead of the logs.
Streaming is the other piece worth a dedicated pass. Time to first token, connection handling, and mid-stream interruptions behave nothing like a normal request/response cycle, so lumping them into the same latency numbers probably hides more than it reveals.
The fallback cost point is worth digging into further. A fallback that "saves" the request can quietly double token spend if the secondary provider has a different pricing tier or context window, and that's easy to miss until the bill shows up. Would be curious if you've tracked cost per successful request across these five in a real workload, not just latency and error rate.
Streaming feels like the bigger gap though. Time to first token and how a gateway handles a dropped stream mid-response deserves a benchmark of its own, separate from standard request/response numbers, especially since two other commenters already flagged it.