Part 1 measured post-quantum TLS handshake cost at same-AZ latency and found about 1 ms added per handshake. Part 2 showed that session resumption makes that cost disappear entirely for ongoing connections. That left one PQ overhead lever unmeasured: cert chain size. ML-DSA-65 signatures are around 3.3 kB and the public key is another ~1.9 kB. Combined, the whole cert file is about 10× the size of an ECDSA P-256 cert. Common intuition says that would hurt the handshake at high RTT, because every extra TCP segment to deliver the cert costs an extra round-trip.
The data says the opposite. At realistic WAN RTT, cert size is statistically invisible.
Setup
Same c7g.large Graviton3 target in us-east-1 as Phase 1 and 2. New dimension: a single loadgen deployed in three different AWS regions per night, hitting the same target through its public IP. That gave three natural RTT points:
| Loadgen region | Measured RTT to target | Role |
|---|---|---|
| us-east-1 (same-AZ) | ~1 ms | baseline, matches Part 1/2 conditions |
| us-west-2 (Oregon) | ~70 ms | typical North-America east↔west |
| ap-northeast-1 (Tokyo) | ~148 ms | trans-Pacific |
Target serves one of three cert types per run, picked by a cert_type terraform variable (ECDSA P-256 / RSA-2048 / ML-DSA-65). Each cert type is tested with 3 trials × 60 s per trial. 27 trials total.
Cert file sizes on disk (the input to this experiment):
| Cert type | File size | Signature size | Public key |
|---|---|---|---|
| ECDSA P-256 | 741 B | ~70 B | ~65 B |
| RSA-2048 | 1,277 B | 256 B | ~270 B |
| ML-DSA-65 | 7,680 B | 3,309 B | 1,952 B |
Ratio 1 : 1.7 : 10.4. ML-DSA-65 is unambiguously the big one. The question Phase 3 asks: does that size translate into measurable handshake time?
Methodology detour: Gatling can't parse ML-DSA certs
The first attempt used Gatling for consistency with Phase 1/2. It failed immediately on the ML-DSA arm: 100% handshake failures. JDK 21's X509 parser rejects the ML-DSA-65 public key OID (2.16.840.1.101.3.4.3.17) during certificate deserialization, which happens before any trust-all client config can take effect. Even with -Dgatling.http.ssl.trustAll=true, the handshake never gets past the Certificate message.
OpenSSL 3.5 handles all three sig algorithms natively, so Phase 3 switched the measurement tool from Gatling to openssl s_time. Trade-off: no per-handshake percentile data, just aggregate mean (count × wall time → average). For a cert-size experiment that's acceptable — the signal shows up as a mean delta, not a tail distribution. Switching tools mid-phase was ugly but is itself a real-world finding worth noting: the client-side tooling that will dominate real-world PQ TLS adoption is still catching up to the server side.
Results
Mean handshake time per trial, across all 27 trials:
| Region | ecdsa | rsa-2048 | ml-dsa-65 | Max spread |
|---|---|---|---|---|
| us-east-1 (~1 ms) | 1.67 ms | 2.86 ms | 3.16 ms | 1.49 ms |
| us-west-2 (~70 ms) | 209.9 ms | 211.6 ms | 211.3 ms | 1.7 ms |
| ap-northeast-1 (~148 ms) | 445.3 ms | 444.2 ms | 445.3 ms | 1.1 ms |
At same-AZ, there is a measurable delta: ML-DSA-65 adds about 1.5 ms over ECDSA (+89%). That matches intuition — the extra cert bytes and the extra signature parse/verify CPU time cost about 1.5 ms when everything is sub-millisecond.
At 70 ms RTT, the delta collapses to 0.7%. At 148 ms RTT, it is statistically zero — ECDSA and ML-DSA-65 numbers are identical to one decimal place.
Trial-to-trial standard deviation was <3 ms at every region, so these aren't noisy numbers hiding a signal. The signal just isn't there.
Why cert size doesn't matter at WAN RTT
TCP's initial congestion window (initcwnd) on modern Linux kernels is 10 segments ≈ 14.6 kB (10 × ~1460-byte MSS). That's the amount of data a server can send in the very first flight after the handshake completes without waiting for the client to ACK anything.
A full TLS 1.3 ServerHello carrying the entire ML-DSA-65 cert chain is around 7.7 kB. Smaller than initcwnd. Which means the server delivers the whole thing in a single flight, in a single round-trip, regardless of whether it's a 700 B ECDSA cert or a 7.7 kB ML-DSA cert. No extra round-trip to pay for.
What would break this: a cert bigger than initcwnd. SLH-DSA signatures are ~17 kB — comfortably past the limit. That's the one PQ sig algorithm where cert size should measurably hurt at WAN RTT. That’s on the Phase 4 agenda.
Closing the Phase 2 1M-payload note
Part 2 flagged that the 1 MB payload arm failed because of NIC (network interface card) saturation — 200 users/sec × 10 reqs × 1 MB ≈ 16 Gbps is well above the c7g.large’s sustained throughput. The measurement was bandwidth-bound, not crypto-bound. As promised, here’s the rerun at a low rate.
Rerun done at 20 users/sec × 10 reqs × 1 MB ≈ 1.6 Gbps (within instance burst envelope):
| Arm | req_total | p95 | p99 | max | % success |
|---|---|---|---|---|---|
| classical-1M | 60,000 | 2 ms | 2.33 ms | 5.33 ms | 100% |
| pq-1M | 60,000 | 2 ms | 2.00 ms | 5.33 ms | 100% |
Classical and PQ are statistically identical at every percentile, confirming the Phase 2 failure was 100% NIC saturation. Once the handshake completes, every byte of application data is encrypted with AES-256-GCM regardless of which KEM negotiated the key — so the KEM choice stops mattering the moment the handshake amortizes.
What this changes
Three places PQ TLS overhead could show up:
- Handshake KEM: ~1 ms at same-AZ (Part 1), 0 ms when resumed (Part 2)
- Cert chain size: 1.5 ms at same-AZ, 0 ms at any WAN RTT (Part 3)
- Record layer: identical across KEMs regardless of payload size (Part 2 + 3)
Across all three levers in all three parts, the only consistently measurable PQ cost is Part 1's +1 ms fresh handshake at low RTT. Everywhere else, PQ hides inside the regular protocol overhead.
If you’re planning the migration to X25519MLKEM768 + ML-DSA-65 and bracing for a user-perceived latency hit at WAN RTT because of certificate size, the data suggests you can stop worrying about that specific lever.
A few honest caveats
- This finding breaks if the cert gets bigger than ~15 kB. TCP ships the first ~15 kB of data in one round-trip; ML-DSA’s 7.7 kB fits comfortably. SLH-DSA signatures are around 17 kB — big enough to need a second round-trip, and that would show up at WAN RTT. Different signature, different story.
- Signing takes CPU time too, not just bandwidth. ML-DSA signing on c7g.large is a few hundred microseconds — too small to notice next to a 70 ms round-trip. On smaller machines (think Raspberry Pi–class hardware or t4g.small), the CPU cost would be a bigger slice of the total and might become visible.
- Three trials per arm is enough for a solid mean, but tight for tail behavior. For p99-level claims I’d want 10+ trials.
- This measures the TLS handshake only. The rest of the PQ cert story — longer chains when an intermediate CA is also signed with ML-DSA, bigger OCSP/CRL fetches to check revocation — is a different experiment.
What's next
Phase 4 candidates:
-
SLH-DSA cert (~17 kB) at WAN RTT — the one PQ sig that should break the
initcwndboundary and finally show the cert-size effect people expect. - Cross-architecture: c7g (Graviton3) vs c7i (Sapphire Rapids) vs c6a (AMD). Does the ML-DSA signing-side CPU cost change the picture on x86?
- Realistic client mix: 80% resumed + 20% fresh across varying RTT, weighted average.
Code and raw results on GitHub. To reproduce the cert-size scenario: see scenarios/04-cert-chain-size/ and docs/03-phase-3-walkthrough-night1.md. Cost for all 27 trials across 3 regions: about $0.75 on on-demand c7g.large.
Full series:
- Part 1 — Fresh-handshake latency
- Part 2 — Session resumption + payload sweep
- Part 3 — this post
Top comments (0)