There's a lot of discussion going on about post-quantum TLS. Both AWS and Cloudflare claim the overhead is minimal — but both are vendor-published, on infrastructure most people don't have. This is the starting post of a different attempt: an independent, reproducible benchmark of X25519MLKEM768 (the NIST-standardized hybrid PQ key exchange, FIPS 203) versus classical X25519. I ran it on two c7g.large EC2 instances in the same AWS zone. The full run cost less than a dollar.
Complete setup — Terraform, Gatling simulation, driver script, raw Gatling result files — is in the repo. This can be reproduced in around ~50 minutes for under $0.30 (based on my experience, accounting for the occasional disruption during a test).
Setup
| Piece | Choice |
|---|---|
| Target | c7g.large spot, Ubuntu 24.04, nginx built from source against OpenSSL 3.5.0 |
| Load generator | c7g.large spot, JDK 21, Gatling 3.15.1 (Java DSL) |
| Placement | Same subnet, same AZ (us-east-1) — cross-AZ latency would swamp the effect |
| Protocol | TLS 1.3 only |
| Cipher suites | TLS_AES_256_GCM_SHA384, TLS_CHACHA20_POLY1305_SHA256 |
| Server cert | Self-signed ECDSA P-256 |
| KEM (A/B knob) | X25519 vs X25519MLKEM768 |
| Response body | 3 bytes ("OK\n") — the point is handshake, not transfer |
| Session resumption | Disabled (fresh handshakes only) |
| Trials | 3 per arm, 5 min each, 30 s warmup dropped |
Note on the self-signed cert
The server serves an ECDSA cert, not an ML-DSA-65 one.
The reason: Gatling uses Netty's BoringSSL under the hood, and BoringSSL doesn't advertise ML-DSA-65 in its supported signature algorithms. If nginx only had an ML-DSA-65 cert to serve, BoringSSL wouldn't be able to complete the handshake at all.
This actually mirrors real-world PQ TLS in 2026: no public CA issues ML-DSA-65 certs. Everyone shipping "PQ TLS" today serves a classical cert with a hybrid KEM. This benchmark measures the KEM overhead under exactly that pattern.
First attempt: what didn't work
My first serious run was at 1000 req/s. The numbers were absurd:
| Metric | Classical | PQ |
|---|---|---|
| min | 1 ms | 1 ms |
| mean | 1,318 ms | 1,230 ms |
| p99 | 6,440 ms | 6,091 ms |
| max | 7,648 ms | 7,548 ms |
Not "PQ is faster than classical" — that's not physically possible when the PQ arm does the same work as classical plus more crypto. The shape (min=1ms, mean=1300ms) is the fingerprint of client-side CPU saturation. 1000 fresh handshakes/sec on a 2-vCPU loadgen means each vCPU is doing 500 handshakes/sec of X25519 or ML-KEM math — beyond a c7g.large's capacity. Requests queue for CPU, means shoot into the seconds range, and any actual PQ vs classical delta gets swallowed by queuing noise.
Cross-check: PQ p99 across the 3 trials ranged 5892 → 6091 → 7623 ms. That trial-to-trial variance is bigger than any real signal I could measure — a red flag on its own.
Lesson: benchmark your load generator before you trust its numbers. Client-side CPU cost is easy to forget when the discussion is all about server overhead.
The real numbers, at 300 req/s
Dropping to 300 req/s puts the loadgen well inside its capacity envelope. Now I'm measuring the handshake itself.
Median of 3 trials, latency in milliseconds:
| Metric | Classical (X25519) | PQ (X25519MLKEM768) | Overhead |
|---|---|---|---|
| min | 1 | 1 | — |
| mean | 2 | 2 | — |
| p50 | 1 | 2 | +1 ms |
| p95 | 4 | 4 | — |
| p99 | 5 | 8 | +3 ms (+60%) |
| max | 14 | 41 | +27 ms |
Trial-by-trial p99 (to show the variance honestly):
| Trial | Classical p99 | PQ p99 |
|---|---|---|
| 1 | 4 | 5 |
| 2 | 5 | 8 |
| 3 | 5 | 39 |
Trial 3's PQ p99 of 39 ms is a real outlier — likely a spot instance CPU steal event or a JVM GC pause, though I haven't isolated the cause yet. Something to nail down in Phase 2 with more trials.
What this tells me
Three findings worth stating:
1. In the common case, PQ overhead is essentially invisible. At 300 req/s on Graviton3, X25519MLKEM768 costs the same 2 ms mean handshake as classical X25519. p50 and p95 are the same at ms-resolution. If your SLO is "handshake under 50 ms", PQ vs classical is not what you should be worrying about.
2. The cost shows up in the tail. Median p99 grew from 5 ms to 8 ms. In one of three trials the PQ p99 was 39 ms — 8x what the median trial saw. This matches what I expected theoretically: KEM operations have more variance than raw X25519, and worst-case scheduling amplifies the difference.
3. This is a floor, not a ceiling. At sub-10 ms latencies, Gatling's integer-millisecond quantization hides sub-ms differences. The "no difference" reading at mean/p50/p95 may be masking real ~0.5 ms deltas. Higher-precision timing and higher throughput on beefier instances would tell more.
What I can't say from this run
- How PQ overhead scales with concurrent connection count. All measurement here is at a fixed rate with the loadgen sized carefully to avoid saturation.
- What happens with session resumption. Session resumption skips the KEM entirely. Real traffic is a mix — measuring just the fresh case is the worst case.
- Whether Graviton3's specific silicon helps or hurts PQ. Would need x86 (c7i.large) or AMD (c6a.large) runs for comparison.
- Anything about payload size scaling. Every response was 3 bytes.
All of these are Phase 2 targets.
Reproduce it
git clone https://github.com/asifmahbubnaeem/pqc-bench
cd pqc-bench
# Configure AWS profile + AMI IDs in terraform/bench/terraform.tfvars
make bench-up
make wait-ready
./scenarios/01-fresh-handshake-latency/run.sh 300 3
make bench-down
Wall time: ~50 min. AWS cost: ~$0.15. The raw Gatling result trees, including per-trial HTML reports and simulation.log files, can be found in results/2026-09-27-scenario-01-rate300/.
What's next
Phase 2 questions:
- Payload sweep: 100 B / 1 KB / 100 KB / 1 MB response. Does handshake overhead disappear into transfer time?
- Session resumption: Full mix, 0-RTT, 1-RTT. What's the effective overhead when 90% of connections resume?
- Cross-architecture: Same benchmark on c7g (Graviton3, ARM), c7i (Intel Sapphire Rapids), c6a (AMD EPYC). Does the KEM math favor any one microarchitecture?
- Concurrency sweep: 100 / 500 / 2000 concurrent connections at fixed rate — where does saturation actually hit, and does the answer differ for PQ vs classical?
I'll post those over the coming weekends. If there's a specific angle you'd want measured, contact me.
Data, methodology and all scripts live in the repo. Corrections and independent reruns very welcome — that's the whole point.
Draft assistance from Claude; all measurements, interpretations, and the setup work are mine and independently verified against the raw data in the repo.
Top comments (1)