Part 1 ran a fresh-handshake load test on a Graviton3 and measured what the ML-KEM half of X25519MLKEM768 adds to TLS 1.3 handshake latency. The headline was that hybrid PQ costs about an order of magnitude more at p95 under a stream of brand-new connections.
That's a correct number for a benchmark almost nobody runs in production.
Real HTTPS traffic doesn't handshake once per request. Browsers pool connections. Load balancers hold them open. TLS session tickets let returning clients skip the key exchange entirely. So the question Part 1 left open is: once we measure in a shape that looks like production, what does the PQ cost actually become?
What changed from Phase 1
Same infra, two toggles.
Server (terraform/target/user-data-fast.sh.tftpl):
ssl_session_cache shared:SSL:10m;
ssl_session_timeout 1h;
ssl_session_tickets on;
Client (Gatling):
http.baseUrl(target)
.shareConnections() // ← the key change
.disableCaching()
...
With shareConnections(), each virtual user does N requests over one TLS session instead of opening a new connection per request. That's how production clients behave.
Everything else is identical to Phase 1: c7g.large Graviton3 target running nginx-pq against the same c7g.large loadgen, OpenSSL 3.5, TLS 1.3, hybrid X25519MLKEM768 KEM, ML-DSA-65 server cert for PQ-capable clients and an ECDSA P-256 cert for BoringSSL-based ones. Full Phase 2 code is on GitHub.
Two scenarios on top of that.
Scenario 02 — session resumption
Four arms: classical-fresh, classical-resumed, pq-fresh, pq-resumed. 300 user-connections per second, 10 requests per user, 5-minute measurement, 3 trials per arm, 12 trials total.
The p95 for each arm, averaged across trials:
| arm | p95 (ms) |
|---|---|
| classical-fresh | 4 |
| classical-resumed | 1 |
| pq-fresh | 5 |
| pq-resumed | 1 |
Fresh PQ adds 1 ms at p95. That's smaller than Part 1's number, mostly because Phase 2 uses hybrid session-cache-enabled nginx-pq with a lighter config path for the sub-ms responses. The relative cost (+25% vs classical fresh) is still real.
Resumed PQ adds nothing. The resumed numbers for classical and PQ are byte-for-byte identical: 1 ms at p95, 1 ms at p99.
Scenario 02 — response time by percentile, PQ overhead fresh vs resumed

This makes sense once we remember what resumption does. A resumed TLS 1.3 session skips the key-exchange exchange entirely: the client proves it has the master secret via a PSK and the connection picks up where it left off. ML-KEM, X25519, ECDH — all irrelevant after the first handshake. From there every record is wrapped in AES-256-GCM, which doesn't care which KEM negotiated the keys that derived the GCM key.
The resumption gain itself (3-4 ms off the p95) is bigger than the PQ cost (1 ms added). Any production fleet already running session resumption gets PQ migration essentially for free at this percentile.
Scenario 03 — payload sweep
Six arms: classical and pq × response sizes 100B, 10K, 100K. All connections resumed (shareConnections on, 10 requests per user). 200 users per second, 5-minute measurement, 3 trials per arm.
| arm | p95 (ms) | max (ms) |
|---|---|---|
| classical-100B | 1 | 7.00 |
| classical-10K | 1 | 5.33 |
| classical-100K | 1 | 7.33 |
| pq-100B | 1 | 6.33 |
| pq-10K | 1 | 5.67 |
| pq-100K | 1 | 6.00 |
Identical at every percentile. Max response time wobbles 5.33 → 7.33 ms but the spread is network jitter, not a classical-vs-PQ signal.
Scenario 03 — p95 across payload sizes, PQ overhead flat at zero

Same reason as before. Once the handshake amortizes over 10 requests, what is left is pure record-layer throughput, and the record layer is identical for each group. The AES-GCM path through c7g.large is the same whether the master secret came from X25519 or X25519MLKEM768.
The 1M arm I wanted to add didn't produce usable data: 200 users/sec × 10 requests × 1 MB = ~16 Gbps steady aggregate throughput, and c7g.large has a 0.75 Gbps NIC baseline (12.5 Gbps burst, but nothing like 16 Gbps sustained). The target hit NIC saturation and the trials filled the 60s-timeout bucket. That measurement was bandwidth-bound, not crypto-bound. 1M belongs in its own low-rate scenario - targeting for next phase.
What this changes
Three places PQ overhead can show up in a TLS deployment: handshake, cert chain, record layer. In Phase 2, we have measured two of them.
- Handshake: 1 ms added at p95 for fresh connections. Zero for resumed.
- Record layer: no measurable difference at any payload size up to 100 KB.
Which leaves cert-chain size — ML-DSA-65 signatures (size around ~3.3 kB ) vs ECDSA (~70 B) — as the remaining lever, and that one lands on cold or infrequently-resumed clients where the bigger ServerHello matters. That could be our next part (phase 3) area.
If we believe our production traffic looks more like scenario 02 and scenario 03 than like Phase 1 — and for most HTTPS services behind a connection pool, it does — then the practical cost of migrating to X25519MLKEM768 on current ARM server hardware is on the order of 1 ms at p95, paid once per client per session cache lifetime.
A few honest caveats
Gatling 3.15 reports response times in integer milliseconds, so sub-millisecond latencies round down to zero. The mean, p50, and p75 columns in both scenarios read as 0 ms for most arms — that is a measurement limitation, not real floor behavior. At p95 and above the numbers are above 1 ms and real, which is what the comparison relies on.
Three trials per arm is enough for the clean comparisons made here but tight for tail percentiles. The p99 column has noise comparable to the signal; I'd want 10 trials before saying anything confidently about p99 behavior.
NIC saturation at 1M payloads means Phase 2 is silent on how PQ scales under high-bandwidth sustained workloads. That is a different question to answer with different setup.
What's next
Phase 3 candidates, roughly ordered:
- Cert-chain size impact: how much does the ~3 kB ML-DSA-65 signature cost at the first-byte percentile across varying RTT? ML-DSA-65 vs ECDSA, same target.
- 1M payload at a sane rate (~20 rps) — bandwidth characterization, not primarily crypto.
- Realistic client mix: 80% resumed + 20% fresh, measuring aggregate p95 as a weighted average the way CDN telemetry would see it.
Code and raw results for Phase 2 are on GitHub. To reproduce: ./scenarios/02-resumption/run.sh 300 3 10 and ./scenarios/03-payload-sweep/run.sh 200 3 10. Total cost on on-demand c7g.large is about $0.50. Based on my experience it can and will interrupt a multi-hour run.
Top comments (0)