If you look at vendor landing pages or benchmarks on social media, every inference provider claims to be “the fastest engine on Earth.” You see sleek bar charts showing thousands of tokens per second, single-digit Time-To-First-Token, and promises of dramatic cost savings.
Then you deploy a 27-billion parameter reasoning model like Qwen 3.8 27B into production, route real multi-turn traffic through those endpoints, and practical systems trade-offs immediately emerge. Single-stream latency behaves very differently from high-concurrency batch throughput. Hardware interconnects dictate whether Tensor Parallelism flies or grinds to a halt. And subtle gateway interpretations of reasoning tokens can quietly balloon your generation budgets.
Over the past several weeks, we ran an exhaustive series of empirical benchmarks across our research platform (gft-studio) to evaluate Qwen 3.8 27B across dedicated infrastructure and leading inference providers: Together AI, Fireworks AI, Nebius, Doubleword, and g factor.
Below, we share the verified engineering telemetry: how parallel topology (Tensor Parallelism vs. Data Parallelism) shapes decode latency, how next-generation B200 hardware scales over H100 baselines, the real-world sweet spot of multi-token speculative decoding (MTP4 vs. MTP8), how prefix caching and prompt speculation behave on structured workflows, and what happens to latency when concurrency pushes to 64 parallel streams.
Start here
Explore complete concurrency 1–64 benchmarks across Together, Fireworks FP8, Doubleword, vanilla vLLM, and g factor on Qwen 3.8 27B. Discover why single-node Tensor Parallelism hits 190 tok/s at c1 while cross-node TP stalls, how MTP4 pushes dual-H100 throughput to 769 tok/s, and how prefix caching unlocks 1,140+ tok/s on structured prompts.
- TTFT: Time-To-First-Token: the elapsed latency from HTTP request submission until the first stream chunk arrives.
- ITL: Inter-Token Latency: the time required to generate and emit each subsequent token during the decode phase.
- TP2 (Tensor Parallelism): Splitting individual weight matrices across 2 GPUs over high-speed NVLink so both GPUs collaborate on every token.
- DP2 (Data Parallelism): Running two independent replicas of the model on separate GPUs, routing requests via a load balancer.
- MTP (Multi-Token Prediction): A speculative decoding technique where additional lightweight heads predict future tokens in a single forward pass.
1. The Experimental Setup: Isolating Real Performance
To make an inference benchmark meaningful, you have to eliminate confounding variables. Comparing a 7B model on FP8 with a 70B model on BF16 tells you nothing. Comparing an API called from a laptop in London with a server hosted in Oregon tells you about transit latency, not engine throughput.
Here is how we standardized our test harness:
-
The Model:
Qwen/Qwen3.8-27B(and its official FP8 quantized variant), pinned to tokenizer revision1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. - The Workload: Standardized AIPerf 0.12.0 test suites streaming over public HTTPS. Prompts averaged ~564 input tokens, requesting exactly 128 generated output tokens at temperature 0 and random seed 42.
- Concurrencies Tested: Standard evaluation across concurrency levels 1, 2, 4, 8, 16, 32, and 64. Low-concurrency cells ran 16 warmup requests followed by 60 strictly measured requests across three separate repetitions (180 measured requests per cell). High-concurrency cells (c16, c32, c64) ran 64 warmup requests and 256 measured requests per repeat.
- The Hardware Baseline: A fixed hardware budget of 2× NVIDIA H100 80GB SXM GPUs per target (except where explicitly comparing next-generation B200 accelerators).
2. The Full Concurrency 1–64 Comparison
In real-world serving, systems rarely operate at a single concurrency point. Early morning traffic might see solitary interactive queries, while daytime peaks bombard your cluster with dozens of simultaneous streams.
Below is the full empirical comparison across Together AI, Fireworks AI (FP8), Doubleword, vanilla standalone vLLM, and g factor with MTP4 speculative decoding:

Output throughput from concurrency 1 to 64 on 2x H100 (or the provider's equivalent). Single-node TP2 (Together) leads at low concurrency; at c64 four of the five engines land within about 10% of each other.
| Concurrency | g factor MTP4 | Together FP8 TP2 | Fireworks FP8 DP2 | Doubleword Public FP8 | Vanilla vLLM FP8 DP2 |
|---|---|---|---|---|---|
| c = 1 | 133.00 | 189.61 | 114.49 | 65.86 | 69.29 |
| c = 2 | 250.65 | 350.57 | 227.15 | 138.48 | 136.81 |
| c = 4 | 437.41 | 612.35 | 401.46 | 291.29 | 253.51 |
| c = 8 | 769.13 | 1033.44 | 640.02 | 534.71 | 459.16 |
| c = 16 | 1252.34 | 1590.84 | 1056.65 | 935.52 | 836.18 |
| c = 32 | 1904.97 | 2579.22 | 1146.75 | 1716.53 | 1601.44 |
| c = 64 | 2611.88 | 2701.47 | 1145.08 | 2649.71 | 2462.44 |
Looking across this spectrum reveals three distinct operational regimes:
- Low Concurrency (c = 1 to 2): Together AI leads single-stream speed (189.61 tok/s at c1, 350.57 tok/s at c2). Because Together serves the model via single-node Tensor Parallelism (TP2), both H100s collaborate on every token decode step. g factor with MTP4 delivers 133.00 tok/s at c1 and 250.65 tok/s at c2, comfortably outperforming Fireworks (114.49 / 227.15 tok/s) and Vanilla vLLM (69.29 / 136.81 tok/s).
- Medium Concurrency (c = 4 to 8): As requests multiply, multi-replica data parallelism (DP2) hits its stride. g factor reaches 437.41 tok/s at c4 and 769.13 tok/s at c8, pulling ahead of Fireworks DP2 (640.02 tok/s) and Vanilla vLLM (459.16 tok/s).
- High Concurrency Saturation (c = 16 to 64): At c32 and c64, all modern engines push past 1,500 to 2,600 tokens per second. Together reaches 2,701 tok/s, Doubleword reaches 2,650 tok/s, and g factor reaches 2,612 tok/s. But raw throughput at high concurrency is only half the story—you also have to examine the latency bill.
The Latency Bill Under Load
It is easy to generate thousands of tokens per second if you let requests sit in a queue. What matters for interactive user experience is Time-To-First-Token (TTFT): how long the user stares at a blank screen before text begins to stream.
| Concurrency | Together FP8 | Fireworks FP8 | Doubleword FP8 | Vanilla vLLM FP8 |
|---|---|---|---|---|
| c = 1 | 141.6 ms | 1088.4 ms | 1093.1 ms | 278.0 ms |
| c = 2 | 174.3 ms | 390.4 ms | 1059.8 ms | 468.2 ms |
| c = 4 | 217.1 ms | 435.3 ms | 826.8 ms | 538.0 ms |
| c = 8 | 221.8 ms | 1323.3 ms | 786.7 ms | 656.9 ms |
| c = 16 | 269.1 ms | 1034.9 ms | 919.5 ms | 705.6 ms |
| c = 32 | 245.8 ms | 2687.1 ms | 975.5 ms | 736.0 ms |
| c = 64 | 1701.7 ms | 7714.6 ms | 1598.6 ms | 1195.3 ms |
Notice what happens between c32 and c64. On Fireworks, throughput stays essentially flat (1,146 tok/s → 1,145 tok/s), while p95 TTFT surges from 2.69 seconds to 7.71 seconds. The GPUs are fully saturated; adding more requests in flight simply queues them up at the door without producing more tokens per second.

p95 Time-To-First-Token on a log scale. Fireworks climbs to 7.7 s at c64 while its throughput stays flat: extra requests only wait in the queue.
3. Architectural Topology: Tensor Parallelism vs. Independent Replicas
The most important architectural lesson from our benchmark is that two GPUs do not make a system; how those two GPUs are connected makes the system.
Why Together Won Single-Stream Latency: High-Speed NVLink
Look at the single-stream results: Together achieved 189.6 tok/s and an Inter-Token Latency (ITL) of just 3.9 milliseconds, compared to 69–133 tok/s on independent single-GPU replicas.
Why? Because Together deployed Tensor Parallelism (TP2) inside a single physical server. In TP2, every linear layer in Qwen 3.8 27B is sliced across both GPUs. For every single token decode step, GPU 0 and GPU 1 compute their respective matrix slices and exchange intermediate activations via an all-reduce collective. At concurrency 1, both H100s collaborate on that single user’s request simultaneously.
In Data Parallelism (DP2), by contrast, GPU 0 handles User A while GPU 1 handles User B. At concurrency 1, User A only uses one GPU, with decode speed physically bounded by that single chip’s memory bandwidth (3.35 TB/s on H100 SXM).
Evaluating Cross-Node Tensor Parallelism
Seeing Together’s single-stream TP2 speed, we tested an experiment: what happens if you run TP2 across two separate 1× H100 cloud instances connected via standard datacenter virtual networking (VPC) rather than intra-node NVLink?
The result demonstrated the severe cost of network barrier latency:
| Concurrency | DP2 (2 Separate Nodes) | Cross-Node TP2 (Over Network) | Single-Node TP2 (NVLink) | Network Interconnect Penalty |
|---|---|---|---|---|
| c = 1 | 73.71 tok/s | 24.98 tok/s | 189.61 tok/s | 2.95x slower than DP2 |
| c = 4 | 274.64 tok/s | 56.47 tok/s | 612.35 tok/s | 4.86x slower than DP2 |
| c = 8 | 483.19 tok/s | 75.75 tok/s | 1,033.44 tok/s | 6.38x slower than DP2 |
At concurrency 8, cross-node TP2 throughput collapsed from 483 tok/s down to 75.75 tok/s. The physical explanation is straightforward:

The same two H100s at concurrency 8: NVLink inside one server, two independent replicas, or tensor parallelism stretched over the cloud network.
During single-token decoding, GEMV computations for a 27B model take only 5 to 10 microseconds. Within a single chassis over NVLink, exchanging activations takes ~2 microseconds over 900 GB/s channels. Across separate physical servers over standard cloud networking, that same transfer requires 800 to 2,000 microseconds. The GPUs spent over 95% of their execution time stalled at network barriers waiting for TCP packets.
The Golden Rule of Topology: Never run Tensor Parallelism across physical machines unless you have dedicated multi-rail InfiniBand. If your GPUs reside on separate nodes, always deploy independent replicas with data parallelism (DP).
4. The Hardware Leap: What Happens on NVIDIA B200?
While H100 remains the workhorse of enterprise inference, next-generation NVIDIA Blackwell (B200) accelerators are entering production. We benchmarked Fireworks AI running Qwen 3.8 27B on a dedicated 2× B200 deployment:
| Concurrency | 2× H100 BF16 (Fireworks) | 2× B200 BF16 (Fireworks) | Observed Hardware Speedup |
|---|---|---|---|
| c = 1 | 97.09 tok/s | 153.38 – 161.32 tok/s | 1.58x – 1.66x |
| c = 4 | 347.12 tok/s | 571.82 – 597.11 tok/s | 1.65x – 1.72x |
| c = 8 | 622.18 tok/s | 979.49 – 1,001.32 tok/s | 1.57x – 1.61x |
Moving from H100 to B200 delivered a clean 1.6x to 1.7x throughput increase with zero code changes. This speedup is directly explained by memory hardware specifications:
- NVIDIA H100 SXM features 3.35 TB/s of HBM3 memory bandwidth.
- NVIDIA B200 features 8.00 TB/s of ultra-dense HBM3e bandwidth (a 2.38x hardware leap).
Because memory-bound autoregressive decoding scales near-linearly with memory bandwidth, B200 allows a 2-GPU cluster to cross the 1,000 tok/s barrier even in unquantized 16-bit precision.
5. Case Study: Reasoning Token Budgets Across API Gateways
During our evaluation of Qwen3.8-27B-FP8 on Doubleword’s public API, we encountered an instructive systems interaction that illustrates how reasoning-capable models interface with standard throughput benchmarks.
We configured our AIPerf test harness with standard parameters to measure decode throughput at a fixed length:
{"max_completion_tokens": 128, "temperature": 0}
While standard non-reasoning requests complete in 1 to 2 seconds for 128 tokens, these initial requests ran for approximately 247 seconds. An inspection of the returned payload counters explained the extended duration:

The 128-token cap applied only to the visible answer. The model first generated 15,774 hidden reasoning tokens, so each request ran for about 247 seconds.
| Metric | Configured Parameter | Returned API Telemetry |
|---|---|---|
| Reasoning Tokens (<think>) | Default reasoning; cap requested | 15,774 tokens |
| Answer Tokens | 128 target | 130 tokens |
| Total Generated Tokens | 128 completion cap | 15,904 tokens (full reasoning trajectory) |
Gateway Semantics and Parameter Interpretation
Tracing the HTTP response headers revealed routing through an upstream gateway adapter. In many recent reasoning model architectures, API proxies interpret max_tokens or max_completion_tokens as applying strictly to the final visible answer, leaving internal reasoning tokens (within <think>...</think> tags) unconstrained.
When no explicit reasoning limit is enforced, the model executes its full natural thought trajectory (~15,774 reasoning tokens) prior to producing the concise final response.
Standardizing the Benchmark: To evaluate pure decode throughput at a controlled output length, Doubleword documents explicitly setting reasoning_effort: "none":
{"max_completion_tokens": 128, "reasoning_effort": "none"}
With explicit non-thinking parameters configured, Doubleword’s public API completed all 540 measured requests cleanly within the target token bounds:
- Concurrency 1: 65.86 tok/s (TTFT p50: 708 ms, ITL: 12.4 ms)
- Concurrency 4: 291.29 tok/s (TTFT p50: 725 ms, ITL: 12.5 ms)
- Concurrency 8: 534.71 tok/s (TTFT p50: 708 ms, ITL: 12.7 ms)
Benchmarking Note on Token Accounting: When measuring reasoning models, it is essential to audit both internal thought tokens and final visible tokens. Standardizing reasoning parameters is necessary to ensure fair, reproducible throughput comparisons across providers.
6. Speculative Decoding: Tuning MTP4 vs. MTP8
Multi-Token Prediction (MTP) is one of the most effective ways to accelerate autoregressive decoding without loading a separate draft model. Instead of predicting a single token per forward pass, lightweight auxiliary heads speculatively propose multiple future tokens, which the main model verifies in parallel.

How multi-token prediction works: cheap heads guess several tokens ahead, the main model checks all guesses in one pass, keeps the matching prefix and adds one token of its own.
Our initial baseline used MTP1 (predicting 1 draft token ahead), yielding 613.54 tok/s at concurrency 8. We then ran an ablation study comparing MTP4 (4 draft tokens) against MTP8 (8 draft tokens) on identical pairs of H100 GPUs with FP8 weights and attention KV:
| Profile | Concurrency | Mean Throughput (tok/s) | TTFT p50 (ms) | ITL p50 (ms) | End-to-End p95 (ms) | Draft Token Acceptance |
|---|---|---|---|---|---|---|
| g factor MTP4 | 1 | 133.00 [132.15–133.54] | 253.95 ms | 5.22 ms | 1,221.62 ms | 59.36% |
| g factor MTP4 | 4 | 437.41 [434.23–440.52] | 290.94 ms | 6.41 ms | 1,407.07 ms | 59.36% |
| g factor MTP4 | 8 | 769.13 [761.22–774.06] | 303.63 ms | 6.93 ms | 1,666.06 ms | 59.36% |
| g factor MTP8 | 1 | 129.81 [129.36–130.14] | 259.32 ms | 5.13 ms | 1,423.34 ms | 39.02% |
| g factor MTP8 | 4 | 417.11 [415.15–420.33] | 302.78 ms | 6.60 ms | 1,601.21 ms | 39.02% |
| g factor MTP8 | 8 | 767.94 [744.30–792.23] | 315.16 ms | 6.77 ms | 1,883.54 ms | 39.02% |
Why did eight draft tokens fail to beat four?
Speculative decoding is a fundamental trade-off between verification speed and draft accuracy:
- With MTP4, the model accepted 59.36% of all draft tokens. Every step generated an average of 3.37 accepted tokens with minimal verification overhead.
- With MTP8, predicting 8 tokens ahead into the future proved significantly harder: acceptance plunged to 39.02%. While the average step size increased slightly to 4.12 tokens, the engine wasted GPU cycles generating and verifying rejected branches.
- As a result, MTP8 yielded 2% to 4% lower throughput and increased p95 end-to-end latency by 13% to 17%.

MTP4 vs MTP8 at concurrency 4: deeper drafts are accepted less often, so the longer step does not pay for the extra verification work.
Conclusion: More speculative depth is not automatically better. For Qwen 3.8 27B, MTP4 is the empirical sweet spot, achieving 769 tok/s on 2× H100 without latency penalties.
7. Concurrency Scaling: Throughput Has a Latency Bill
Concurrency 8 is a standard comparison point, but it does not represent the saturation capacity of two H100s. We tested the scaling behavior of g factor under heavy loads of 16, 32, and 64 concurrent streams:
| Concurrency | MTP1 Throughput | MTP4 Tuned Throughput | MTP4 TTFT p95 (ms) | MTP4 End-to-End p95 (ms) | Interactive Serving SLA |
|---|---|---|---|---|---|
| c = 16 | 1,013.98 tok/s | 1,252.34 tok/s | 530.00 ms | 2,065.24 ms | Optimal interactive SLA |
| c = 32 | 1,604.98 tok/s | 1,904.97 tok/s | 794.38 ms | 2,780.32 ms | High-throughput sweet spot |
| c = 64 | 1,645.97 tok/s | 2,611.88 tok/s | 1,433.17 ms | 4,111.71 ms | Queueing delay boundary |
2,612 tok/s on two H100s looks incredible on a slide, but you have to check the latency bill:
- At c = 32, the cluster operates in its optimal production window: generating 1,905 tok/s while keeping p95 Time-To-First-Token well under 800 ms.
- At c = 64, throughput expands to 2,612 tok/s, but prefill queueing delays cause p95 TTFT to climb to 1.43 seconds.
In production inference, peak throughput numbers are meaningless without latency percentiles. Always size your concurrency targets to your application’s TTFT service level objectives.
8. Prefix Caching & Prompt Speculation: Give Them Something to Reuse
In short-prompt benchmarks (~500 input tokens), prefix caching typically reports near-zero cache hits because each request presents completely unique text.
To evaluate how caching and speculation perform in real agentic workflows (where prompts share long system prompts, tool definitions, and conversation histories), we tested a workload with a 2,820-token shared prefix and a unique prompt tail (total input: 2,870 tokens):

Why the shared prefix helps twice: the cached prefix skips most of prefill, and suffix speculation copies tool-call text that already appears in the prompt.
| Optimization Strategy | c=1 tok/s | c=4 tok/s | c=8 tok/s | Prefix Cache Hit Ratio | Draft Token Acceptance |
|---|---|---|---|---|---|
| Baseline (Cache Disabled) | 57.92 | 209.92 | 349.84 | Disabled | — |
| Automatic Prefix Caching | 60.70 | 230.88 | 418.49 | 80.47% | — |
| Prefix Cache + N-gram (k=4) | 172.62 | 558.05 | 1,046.40 | 83.76% | 90.18% |
| Prefix Cache + Suffix Speculation | 180.36 | 655.77 | 1,140.26 | 83.76% | 99.12% |
The empirical takeaways for long-context and agentic workflows are substantial:
- Prefix Caching Eliminates Prefill Latency: Enabling prefix caching achieved an 80.47% cache hit rate, reducing concurrency-8 TTFT p50 from 855 ms down to 323 ms and delivering an immediate +19.6% throughput boost.
- Context-Aware Speculation (N-gram & Suffix): When repetitive or structured patterns exist in the prompt, n-gram and suffix speculation propose continuations directly from the cached context without requiring an auxiliary neural net. This accelerated decode throughput by 2.7x to 2.9x, pushing throughput to 1,140 tok/s with a 99% draft acceptance rate.

Concurrency-8 throughput on a 2,870-token prompt with a 2,820-token shared prefix, as each optimization is switched on.
9. Architectural Rules for Honest Benchmarks
When evaluating published inference benchmarks or sizing dedicated clusters, five architectural dimensions determine real-world performance:
- What was the GPU interconnect? Is that 1,000 tok/s running on a single server with 900 GB/s NVLink (TP2), or across independent nodes (DP2)?
- What precision was actually served? Was it BF16, calibrated FP8, or aggressive 4-bit quantization? Has domain accuracy been qualified on target tasks?
- Where was the client running? Were requests sent across the public internet or through internal cluster networking?
- How were reasoning tokens budgeted? Does the token limit include the entire completion, and were internal thought chains isolated during speed testing?
- What is the latency trade-off? High batch throughput is easy to achieve at high concurrency, but does the Time-To-First-Token still meet your interactive user requirements?
The Bottom Line: High-throughput inference is not a black-box commodity. It is a systems discipline that balances memory bandwidth, network topology, speculative decoding depth, and cache management. When engineered properly, a dedicated 2-GPU cluster can deliver world-class latency, break past 1,000 tokens per second, and provide complete data sovereignty at predictable costs. That is what we set up in a private LLM deployment.
Originally published at g-ftech.com.
Top comments (0)