DEV Community

asif naeem
asif naeem

Posted on

pqc-bench Part 3: ML-DSA-65 certs are 10 bigger than ECDSA — and don't cost a thing on real networks

Part 1 measured post-quantum TLS handshake cost at same-AZ latency and found about 1 ms added per handshake. Part 2 showed that session resumption makes that cost disappear entirely for ongoing connections. That left one PQ overhead lever unmeasured: cert chain size. ML-DSA-65 signatures are around 3.3 kB and the public key is another ~1.9 kB. Combined, the whole cert file is about 10× the size of an ECDSA P-256 cert. Common intuition says that would hurt the handshake at high RTT, because every extra TCP segment to deliver the cert costs an extra round-trip.

The data says the opposite. At realistic WAN RTT, cert size is statistically invisible.

Setup

Same c7g.large Graviton3 target in us-east-1 as Phase 1 and 2. New dimension: a single loadgen deployed in three different AWS regions per night, hitting the same target through its public IP. That gave three natural RTT points:

Loadgen region Measured RTT to target Role
us-east-1 (same-AZ) ~1 ms baseline, matches Part 1/2 conditions
us-west-2 (Oregon) ~70 ms typical North-America east↔west
ap-northeast-1 (Tokyo) ~148 ms trans-Pacific

Target serves one of three cert types per run, picked by a cert_type terraform variable (ECDSA P-256 / RSA-2048 / ML-DSA-65). Each cert type is tested with 3 trials × 60 s per trial. 27 trials total.

Cert file sizes on disk (the input to this experiment):

Cert type File size Signature size Public key
ECDSA P-256 741 B ~70 B ~65 B
RSA-2048 1,277 B 256 B ~270 B
ML-DSA-65 7,680 B 3,309 B 1,952 B

Ratio 1 : 1.7 : 10.4. ML-DSA-65 is unambiguously the big one. The question Phase 3 asks: does that size translate into measurable handshake time?

Methodology detour: Gatling can't parse ML-DSA certs

The first attempt used Gatling for consistency with Phase 1/2. It failed immediately on the ML-DSA arm: 100% handshake failures. JDK 21's X509 parser rejects the ML-DSA-65 public key OID (2.16.840.1.101.3.4.3.17) during certificate deserialization, which happens before any trust-all client config can take effect. Even with -Dgatling.http.ssl.trustAll=true, the handshake never gets past the Certificate message.

OpenSSL 3.5 handles all three sig algorithms natively, so Phase 3 switched the measurement tool from Gatling to openssl s_time. Trade-off: no per-handshake percentile data, just aggregate mean (count × wall time → average). For a cert-size experiment that's acceptable — the signal shows up as a mean delta, not a tail distribution. Switching tools mid-phase was ugly but is itself a real-world finding worth noting: the client-side tooling that will dominate real-world PQ TLS adoption is still catching up to the server side.

Results

Mean handshake time per trial, across all 27 trials:

Region ecdsa rsa-2048 ml-dsa-65 Max spread
us-east-1 (~1 ms) 1.67 ms 2.86 ms 3.16 ms 1.49 ms
us-west-2 (~70 ms) 209.9 ms 211.6 ms 211.3 ms 1.7 ms
ap-northeast-1 (~148 ms) 445.3 ms 444.2 ms 445.3 ms 1.1 ms

At same-AZ, there is a measurable delta: ML-DSA-65 adds about 1.5 ms over ECDSA (+89%). That matches intuition — the extra cert bytes and the extra signature parse/verify CPU time cost about 1.5 ms when everything is sub-millisecond.

At 70 ms RTT, the delta collapses to 0.7%. At 148 ms RTT, it is statistically zero — ECDSA and ML-DSA-65 numbers are identical to one decimal place.

Trial-to-trial standard deviation was <3 ms at every region, so these aren't noisy numbers hiding a signal. The signal just isn't there.

Why cert size doesn't matter at WAN RTT

TCP's initial congestion window (initcwnd) on modern Linux kernels is 10 segments ≈ 14.6 kB (10 × ~1460-byte MSS). That's the amount of data a server can send in the very first flight after the handshake completes without waiting for the client to ACK anything.

A full TLS 1.3 ServerHello carrying the entire ML-DSA-65 cert chain is around 7.7 kB. Smaller than initcwnd. Which means the server delivers the whole thing in a single flight, in a single round-trip, regardless of whether it's a 700 B ECDSA cert or a 7.7 kB ML-DSA cert. No extra round-trip to pay for.

What would break this: a cert bigger than initcwnd. SLH-DSA signatures are ~17 kB — comfortably past the limit. That's the one PQ sig algorithm where cert size should measurably hurt at WAN RTT. That’s on the Phase 4 agenda.

Closing the Phase 2 1M-payload note

Part 2 flagged that the 1 MB payload arm failed because of NIC (network interface card) saturation — 200 users/sec × 10 reqs × 1 MB ≈ 16 Gbps is well above the c7g.large’s sustained throughput. The measurement was bandwidth-bound, not crypto-bound. As promised, here’s the rerun at a low rate.

Rerun done at 20 users/sec × 10 reqs × 1 MB ≈ 1.6 Gbps (within instance burst envelope):

Arm req_total p95 p99 max % success
classical-1M 60,000 2 ms 2.33 ms 5.33 ms 100%
pq-1M 60,000 2 ms 2.00 ms 5.33 ms 100%

Classical and PQ are statistically identical at every percentile, confirming the Phase 2 failure was 100% NIC saturation. Once the handshake completes, every byte of application data is encrypted with AES-256-GCM regardless of which KEM negotiated the key — so the KEM choice stops mattering the moment the handshake amortizes.

What this changes

Three places PQ TLS overhead could show up:

  1. Handshake KEM: ~1 ms at same-AZ (Part 1), 0 ms when resumed (Part 2)
  2. Cert chain size: 1.5 ms at same-AZ, 0 ms at any WAN RTT (Part 3)
  3. Record layer: identical across KEMs regardless of payload size (Part 2 + 3)

Across all three levers in all three parts, the only consistently measurable PQ cost is Part 1's +1 ms fresh handshake at low RTT. Everywhere else, PQ hides inside the regular protocol overhead.

If you’re planning the migration to X25519MLKEM768 + ML-DSA-65 and bracing for a user-perceived latency hit at WAN RTT because of certificate size, the data suggests you can stop worrying about that specific lever.

A few honest caveats

  • This finding breaks if the cert gets bigger than ~15 kB. TCP ships the first ~15 kB of data in one round-trip; ML-DSA’s 7.7 kB fits comfortably. SLH-DSA signatures are around 17 kB — big enough to need a second round-trip, and that would show up at WAN RTT. Different signature, different story.
  • Signing takes CPU time too, not just bandwidth. ML-DSA signing on c7g.large is a few hundred microseconds — too small to notice next to a 70 ms round-trip. On smaller machines (think Raspberry Pi–class hardware or t4g.small), the CPU cost would be a bigger slice of the total and might become visible.
  • Three trials per arm is enough for a solid mean, but tight for tail behavior. For p99-level claims I’d want 10+ trials.
  • This measures the TLS handshake only. The rest of the PQ cert story — longer chains when an intermediate CA is also signed with ML-DSA, bigger OCSP/CRL fetches to check revocation — is a different experiment.

What's next

Phase 4 candidates:

  1. SLH-DSA cert (~17 kB) at WAN RTT — the one PQ sig that should break the initcwnd boundary and finally show the cert-size effect people expect.
  2. Cross-architecture: c7g (Graviton3) vs c7i (Sapphire Rapids) vs c6a (AMD). Does the ML-DSA signing-side CPU cost change the picture on x86?
  3. Realistic client mix: 80% resumed + 20% fresh across varying RTT, weighted average.

Code and raw results on GitHub. To reproduce the cert-size scenario: see scenarios/04-cert-chain-size/ and docs/03-phase-3-walkthrough-night1.md. Cost for all 27 trials across 3 regions: about $0.75 on on-demand c7g.large.

Full series:

Top comments (0)