<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: asif naeem</title>
    <description>The latest articles on DEV Community by asif naeem (@asif_naeem_bd5842be56bfcd).</description>
    <link>https://dev.to/asif_naeem_bd5842be56bfcd</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4145821%2Ff4f7c43a-a0dc-459b-9bb2-f05dd6d6a755.jpg</url>
      <title>DEV Community: asif naeem</title>
      <link>https://dev.to/asif_naeem_bd5842be56bfcd</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/asif_naeem_bd5842be56bfcd"/>
    <language>en</language>
    <item>
      <title>PQC-Bench Part 1 — measuring X25519MLKEM768 costs on a $0.04/hour Graviton3</title>
      <dc:creator>asif naeem</dc:creator>
      <pubDate>Mon, 28 Sep 2026 19:46:30 +0000</pubDate>
      <link>https://dev.to/asif_naeem_bd5842be56bfcd/pqc-bench-part-1-measuring-x25519mlkem768-costs-on-a-004hour-graviton3-2b2d</link>
      <guid>https://dev.to/asif_naeem_bd5842be56bfcd/pqc-bench-part-1-measuring-x25519mlkem768-costs-on-a-004hour-graviton3-2b2d</guid>
      <description>&lt;p&gt;There's a lot of discussion going on about post-quantum TLS. Both AWS and Cloudflare claim the overhead is minimal — but both are vendor-published, on infrastructure most people don't have. This is the starting post of a different attempt: an independent, reproducible benchmark of &lt;code&gt;X25519MLKEM768&lt;/code&gt; (the NIST-standardized hybrid PQ key exchange, FIPS 203) versus classical &lt;code&gt;X25519&lt;/code&gt;. I ran it on two &lt;code&gt;c7g.large&lt;/code&gt; EC2 instances in the same AWS zone. The full run cost less than a dollar.&lt;/p&gt;

&lt;p&gt;Complete setup — Terraform, Gatling simulation, driver script, raw Gatling result files — is &lt;a href="https://github.com/asifmahbubnaeem/pqc-bench/tree/main" rel="noopener noreferrer"&gt;in the repo&lt;/a&gt;. This can be reproduced in around ~50 minutes for under $0.30 (based on my experience, accounting for the occasional disruption during a test).&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Piece&lt;/th&gt;
&lt;th&gt;Choice&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Target&lt;/td&gt;
&lt;td&gt;c7g.large spot, Ubuntu 24.04, nginx built from source against OpenSSL 3.5.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load generator&lt;/td&gt;
&lt;td&gt;c7g.large spot, JDK 21, Gatling 3.15.1 (Java DSL)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Placement&lt;/td&gt;
&lt;td&gt;Same subnet, same AZ (us-east-1) — cross-AZ latency would swamp the effect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Protocol&lt;/td&gt;
&lt;td&gt;TLS 1.3 only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cipher suites&lt;/td&gt;
&lt;td&gt;TLS_AES_256_GCM_SHA384, TLS_CHACHA20_POLY1305_SHA256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Server cert&lt;/td&gt;
&lt;td&gt;Self-signed ECDSA P-256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;KEM (A/B knob)&lt;/td&gt;
&lt;td&gt;X25519 vs X25519MLKEM768&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Response body&lt;/td&gt;
&lt;td&gt;3 bytes ("OK\n") — the point is handshake, not transfer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session resumption&lt;/td&gt;
&lt;td&gt;Disabled (fresh handshakes only)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trials&lt;/td&gt;
&lt;td&gt;3 per arm, 5 min each, 30 s warmup dropped&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Note on the self-signed cert
&lt;/h3&gt;

&lt;p&gt;The server serves an &lt;strong&gt;ECDSA cert&lt;/strong&gt;, not an ML-DSA-65 one.&lt;/p&gt;

&lt;p&gt;The reason: Gatling uses Netty's BoringSSL under the hood, and BoringSSL doesn't advertise ML-DSA-65 in its supported signature algorithms. If nginx only had an ML-DSA-65 cert to serve, BoringSSL wouldn't be able to complete the handshake at all.&lt;/p&gt;

&lt;p&gt;This actually mirrors real-world PQ TLS in 2026: no public CA issues ML-DSA-65 certs. Everyone shipping "PQ TLS" today serves a classical cert with a hybrid KEM. This benchmark measures the KEM overhead under exactly that pattern.&lt;/p&gt;

&lt;h2&gt;
  
  
  First attempt: what didn't work
&lt;/h2&gt;

&lt;p&gt;My first serious run was at 1000 req/s. The numbers were absurd:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Classical&lt;/th&gt;
&lt;th&gt;PQ&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;min&lt;/td&gt;
&lt;td&gt;1 ms&lt;/td&gt;
&lt;td&gt;1 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;mean&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,318 ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,230 ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p99&lt;/td&gt;
&lt;td&gt;6,440 ms&lt;/td&gt;
&lt;td&gt;6,091 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;7,648 ms&lt;/td&gt;
&lt;td&gt;7,548 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Not "PQ is faster than classical" — that's not physically possible when the PQ arm does the same work as classical plus more crypto. The shape (min=1ms, mean=1300ms) is the fingerprint of &lt;strong&gt;client-side CPU saturation&lt;/strong&gt;. 1000 fresh handshakes/sec on a 2-vCPU loadgen means each vCPU is doing 500 handshakes/sec of X25519 or ML-KEM math — beyond a c7g.large's capacity. Requests queue for CPU, means shoot into the seconds range, and any actual PQ vs classical delta gets swallowed by queuing noise.&lt;/p&gt;

&lt;p&gt;Cross-check: PQ p99 across the 3 trials ranged 5892 → 6091 → 7623 ms. That trial-to-trial variance is bigger than any real signal I could measure — a red flag on its own.&lt;/p&gt;

&lt;p&gt;Lesson: &lt;strong&gt;benchmark your load generator before you trust its numbers.&lt;/strong&gt; Client-side CPU cost is easy to forget when the discussion is all about server overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real numbers, at 300 req/s
&lt;/h2&gt;

&lt;p&gt;Dropping to 300 req/s puts the loadgen well inside its capacity envelope. Now I'm measuring the handshake itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Median of 3 trials, latency in milliseconds:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Classical (X25519)&lt;/th&gt;
&lt;th&gt;PQ (X25519MLKEM768)&lt;/th&gt;
&lt;th&gt;Overhead&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;min&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mean&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p50&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;+1 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p95&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;p99&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+3 ms (+60%)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;41&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;+27 ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Trial-by-trial p99&lt;/strong&gt; (to show the variance honestly):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Trial&lt;/th&gt;
&lt;th&gt;Classical p99&lt;/th&gt;
&lt;th&gt;PQ p99&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;39&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Trial 3's PQ p99 of 39 ms is a real outlier — likely a spot instance CPU steal event or a JVM GC pause, though I haven't isolated the cause yet. Something to nail down in Phase 2 with more trials.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this tells me
&lt;/h2&gt;

&lt;p&gt;Three findings worth stating:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. In the common case, PQ overhead is essentially invisible.&lt;/strong&gt; At 300 req/s on Graviton3, &lt;code&gt;X25519MLKEM768&lt;/code&gt; costs the same 2 ms mean handshake as classical &lt;code&gt;X25519&lt;/code&gt;. p50 and p95 are the same at ms-resolution. If your SLO is "handshake under 50 ms", PQ vs classical is not what you should be worrying about.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. The cost shows up in the tail.&lt;/strong&gt; Median p99 grew from 5 ms to 8 ms. In one of three trials the PQ p99 was 39 ms — 8x what the median trial saw. This matches what I expected theoretically: KEM operations have more variance than raw X25519, and worst-case scheduling amplifies the difference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. This is a floor, not a ceiling.&lt;/strong&gt; At sub-10 ms latencies, Gatling's integer-millisecond quantization hides sub-ms differences. The "no difference" reading at mean/p50/p95 may be masking real ~0.5 ms deltas. Higher-precision timing and higher throughput on beefier instances would tell more.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I can't say from this run
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;How PQ overhead scales with concurrent connection count.&lt;/strong&gt; All measurement here is at a fixed rate with the loadgen sized carefully to avoid saturation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happens with session resumption.&lt;/strong&gt; Session resumption skips the KEM entirely. Real traffic is a mix — measuring just the fresh case is the worst case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Whether Graviton3's specific silicon helps or hurts PQ.&lt;/strong&gt; Would need x86 (c7i.large) or AMD (c6a.large) runs for comparison.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anything about payload size scaling.&lt;/strong&gt; Every response was 3 bytes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of these are Phase 2 targets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/asifmahbubnaeem/pqc-bench
&lt;span class="nb"&gt;cd &lt;/span&gt;pqc-bench
&lt;span class="c"&gt;# Configure AWS profile + AMI IDs in terraform/bench/terraform.tfvars&lt;/span&gt;
make bench-up
make wait-ready
./scenarios/01-fresh-handshake-latency/run.sh 300 3
make bench-down
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wall time: ~50 min. AWS cost: ~$0.15. The raw Gatling result trees, including per-trial HTML reports and simulation.log files, can be found in &lt;code&gt;results/2026-09-27-scenario-01-rate300/&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;Phase 2 questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Payload sweep&lt;/strong&gt;: 100 B / 1 KB / 100 KB / 1 MB response. Does handshake overhead disappear into transfer time?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session resumption&lt;/strong&gt;: Full mix, 0-RTT, 1-RTT. What's the effective overhead when 90% of connections resume?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-architecture&lt;/strong&gt;: Same benchmark on c7g (Graviton3, ARM), c7i (Intel Sapphire Rapids), c6a (AMD EPYC). Does the KEM math favor any one microarchitecture?&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concurrency sweep&lt;/strong&gt;: 100 / 500 / 2000 concurrent connections at fixed rate — where does saturation actually hit, and does the answer differ for PQ vs classical?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'll post those over the coming weekends. If there's a specific angle you'd want measured, &lt;a href="mailto:asifnaim0123@gmail.com"&gt;contact me&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Data, methodology and all scripts live in the repo. Corrections and independent reruns very welcome — that's the whole point.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Draft assistance from Claude; all measurements, interpretations, and the setup work are mine and independently verified against the raw data in the repo.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>tls</category>
      <category>security</category>
      <category>aws</category>
      <category>benchmark</category>
    </item>
  </channel>
</rss>
