DEV Community

asif naeem
asif naeem

Posted on

PQC-Bench Part 1 — measuring X25519MLKEM768 costs on a $0.04/hour Graviton3

There's a lot of discussion going on about post-quantum TLS. Both AWS and Cloudflare claim the overhead is minimal — but both are vendor-published, on infrastructure most people don't have. This is the starting post of a different attempt: an independent, reproducible benchmark of X25519MLKEM768 (the NIST-standardized hybrid PQ key exchange, FIPS 203) versus classical X25519. I ran it on two c7g.large EC2 instances in the same AWS zone. The full run cost less than a dollar.

Complete setup — Terraform, Gatling simulation, driver script, raw Gatling result files — is in the repo. This can be reproduced in around ~50 minutes for under $0.30 (based on my experience, accounting for the occasional disruption during a test).

Setup

Piece Choice
Target c7g.large spot, Ubuntu 24.04, nginx built from source against OpenSSL 3.5.0
Load generator c7g.large spot, JDK 21, Gatling 3.15.1 (Java DSL)
Placement Same subnet, same AZ (us-east-1) — cross-AZ latency would swamp the effect
Protocol TLS 1.3 only
Cipher suites TLS_AES_256_GCM_SHA384, TLS_CHACHA20_POLY1305_SHA256
Server cert Self-signed ECDSA P-256
KEM (A/B knob) X25519 vs X25519MLKEM768
Response body 3 bytes ("OK\n") — the point is handshake, not transfer
Session resumption Disabled (fresh handshakes only)
Trials 3 per arm, 5 min each, 30 s warmup dropped

Note on the self-signed cert

The server serves an ECDSA cert, not an ML-DSA-65 one.

The reason: Gatling uses Netty's BoringSSL under the hood, and BoringSSL doesn't advertise ML-DSA-65 in its supported signature algorithms. If nginx only had an ML-DSA-65 cert to serve, BoringSSL wouldn't be able to complete the handshake at all.

This actually mirrors real-world PQ TLS in 2026: no public CA issues ML-DSA-65 certs. Everyone shipping "PQ TLS" today serves a classical cert with a hybrid KEM. This benchmark measures the KEM overhead under exactly that pattern.

First attempt: what didn't work

My first serious run was at 1000 req/s. The numbers were absurd:

Metric Classical PQ
min 1 ms 1 ms
mean 1,318 ms 1,230 ms
p99 6,440 ms 6,091 ms
max 7,648 ms 7,548 ms

Not "PQ is faster than classical" — that's not physically possible when the PQ arm does the same work as classical plus more crypto. The shape (min=1ms, mean=1300ms) is the fingerprint of client-side CPU saturation. 1000 fresh handshakes/sec on a 2-vCPU loadgen means each vCPU is doing 500 handshakes/sec of X25519 or ML-KEM math — beyond a c7g.large's capacity. Requests queue for CPU, means shoot into the seconds range, and any actual PQ vs classical delta gets swallowed by queuing noise.

Cross-check: PQ p99 across the 3 trials ranged 5892 → 6091 → 7623 ms. That trial-to-trial variance is bigger than any real signal I could measure — a red flag on its own.

Lesson: benchmark your load generator before you trust its numbers. Client-side CPU cost is easy to forget when the discussion is all about server overhead.

The real numbers, at 300 req/s

Dropping to 300 req/s puts the loadgen well inside its capacity envelope. Now I'm measuring the handshake itself.

Median of 3 trials, latency in milliseconds:

Metric Classical (X25519) PQ (X25519MLKEM768) Overhead
min 1 1 —
mean 2 2 —
p50 1 2 +1 ms
p95 4 4 —
p99 5 8 +3 ms (+60%)
max 14 41 +27 ms

Trial-by-trial p99 (to show the variance honestly):

Trial Classical p99 PQ p99
1 4 5
2 5 8
3 5 39

Trial 3's PQ p99 of 39 ms is a real outlier — likely a spot instance CPU steal event or a JVM GC pause, though I haven't isolated the cause yet. Something to nail down in Phase 2 with more trials.

What this tells me

Three findings worth stating:

1. In the common case, PQ overhead is essentially invisible. At 300 req/s on Graviton3, X25519MLKEM768 costs the same 2 ms mean handshake as classical X25519. p50 and p95 are the same at ms-resolution. If your SLO is "handshake under 50 ms", PQ vs classical is not what you should be worrying about.

2. The cost shows up in the tail. Median p99 grew from 5 ms to 8 ms. In one of three trials the PQ p99 was 39 ms — 8x what the median trial saw. This matches what I expected theoretically: KEM operations have more variance than raw X25519, and worst-case scheduling amplifies the difference.

3. This is a floor, not a ceiling. At sub-10 ms latencies, Gatling's integer-millisecond quantization hides sub-ms differences. The "no difference" reading at mean/p50/p95 may be masking real ~0.5 ms deltas. Higher-precision timing and higher throughput on beefier instances would tell more.

What I can't say from this run

  • How PQ overhead scales with concurrent connection count. All measurement here is at a fixed rate with the loadgen sized carefully to avoid saturation.
  • What happens with session resumption. Session resumption skips the KEM entirely. Real traffic is a mix — measuring just the fresh case is the worst case.
  • Whether Graviton3's specific silicon helps or hurts PQ. Would need x86 (c7i.large) or AMD (c6a.large) runs for comparison.
  • Anything about payload size scaling. Every response was 3 bytes.

All of these are Phase 2 targets.

Reproduce it

git clone https://github.com/asifmahbubnaeem/pqc-bench
cd pqc-bench
# Configure AWS profile + AMI IDs in terraform/bench/terraform.tfvars
make bench-up
make wait-ready
./scenarios/01-fresh-handshake-latency/run.sh 300 3
make bench-down
Enter fullscreen mode Exit fullscreen mode

Wall time: ~50 min. AWS cost: ~$0.15. The raw Gatling result trees, including per-trial HTML reports and simulation.log files, can be found in results/2026-09-27-scenario-01-rate300/.

What's next

Phase 2 questions:

  • Payload sweep: 100 B / 1 KB / 100 KB / 1 MB response. Does handshake overhead disappear into transfer time?
  • Session resumption: Full mix, 0-RTT, 1-RTT. What's the effective overhead when 90% of connections resume?
  • Cross-architecture: Same benchmark on c7g (Graviton3, ARM), c7i (Intel Sapphire Rapids), c6a (AMD EPYC). Does the KEM math favor any one microarchitecture?
  • Concurrency sweep: 100 / 500 / 2000 concurrent connections at fixed rate — where does saturation actually hit, and does the answer differ for PQ vs classical?

I'll post those over the coming weekends. If there's a specific angle you'd want measured, contact me.


Data, methodology and all scripts live in the repo. Corrections and independent reruns very welcome — that's the whole point.

Draft assistance from Claude; all measurements, interpretations, and the setup work are mine and independently verified against the raw data in the repo.

Top comments (1)