<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Scale Vector</title>
    <description>The latest articles on DEV Community by Scale Vector (@scalevector).</description>
    <link>https://dev.to/scalevector</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147968%2Fff207aab-11e2-4ab5-ae30-12c9632a2498.png</url>
      <title>DEV Community: Scale Vector</title>
      <link>https://dev.to/scalevector</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/scalevector"/>
    <language>en</language>
    <item>
      <title>What Actually Happens When 100,000 Users Click 'Buy' at the Exact Same Millisecond</title>
      <dc:creator>Scale Vector</dc:creator>
      <pubDate>Mon, 28 Sep 2026 21:34:09 +0000</pubDate>
      <link>https://dev.to/scalevector/what-actually-happens-when-100000-users-click-buy-at-the-exact-same-millisecond-1je8</link>
      <guid>https://dev.to/scalevector/what-actually-happens-when-100000-users-click-buy-at-the-exact-same-millisecond-1je8</guid>
      <description>&lt;p&gt;During a high-concurrency flash sale—such as ten limited-inventory items going live at 08:00 AM—traffic does not arrive in a smooth, manageable curve. It arrives as an instantaneous impulse function: 100,000 requests hitting the ingress layer within the same millisecond window.&lt;/p&gt;

&lt;p&gt;In a conventional architecture where an API gateway proxies requests directly to an application cluster backed by a relational database, this volume will not merely oversell inventory. It causes a cascading system collapse across network, gateway, and storage layers.&lt;/p&gt;




&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Hj0tPk9ZG3o?start=1" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;




&lt;h2&gt;
  
  
  Architecture at a Glance
&lt;/h2&gt;

&lt;p&gt;Why unbuffered architectures fail during sudden demand spikes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Network Layer (Linux Kernel):&lt;/strong&gt; Incoming SYN packets saturate kernel socket queues, causing silent connection drops before the application runtime receives execution time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gateway Layer (Reverse Proxy):&lt;/strong&gt; High connection churn exhausts ephemeral ports into &lt;code&gt;TIME_WAIT&lt;/code&gt; state, surfacing cascading &lt;code&gt;HTTP 504 Gateway Timeout&lt;/code&gt; errors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage Layer (Database):&lt;/strong&gt; Pessimistic locking (&lt;code&gt;SELECT ... FOR UPDATE&lt;/code&gt;) causes queueing across connection pools, stalling unrelated services across the platform.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A resilient architecture shifts concurrency management out of the relational engine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ 100,000 Users ]
       │
       ▼ (1. Rate limit &amp;amp; In-Memory Check)
[ Edge / Ingress ] ──(Lua Atomic Decr)──► [ Redis Cluster ] 
       │                                         │
       │ (Only 10 requests pass)                 │ (99,990 rejected instantly)
       ▼                                         ▼
[ Kafka Queue ] ────────────────────────► [ Background Workers ]
                                                 │
                                                 ▼ (Optimistic Lock)
                                          [ PostgreSQL / SQL Server ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Failure Modes Across the Stack
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Kernel Backlog Saturation (Network Layer)
&lt;/h3&gt;

&lt;p&gt;Before an HTTP payload reaches user space, the Linux kernel manages connection handshakes across two primary queues:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client ──► [ SYN Queue ] ──► [ Accept Queue ] ──► Application accept()
                 ▲
                 └── Saturated queue: Packets dropped silently
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Failure Mechanism:&lt;/strong&gt; Default kernel configurations often set &lt;code&gt;net.core.somaxconn&lt;/code&gt; to conservative values (such as &lt;code&gt;128&lt;/code&gt; or &lt;code&gt;4096&lt;/code&gt;). Under an instantaneous 100,000-connection burst, the SYN backlog overflows immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observed Behavior:&lt;/strong&gt; Clients experience stalled TLS handshakes and connection timeouts (&lt;code&gt;ETIMEDOUT&lt;/code&gt;). The application logs zero errors because packets are dropped at the kernel boundary.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Ephemeral Port Starvation (Gateway Layer)
&lt;/h3&gt;

&lt;p&gt;Reverse proxies and load balancers (such as Envoy, Nginx, or an AWS ALB) maintain outbound connections toward backend instances.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Outbound connections allocate an ephemeral port from the local range (typically &lt;code&gt;32768&lt;/code&gt;–&lt;code&gt;60999&lt;/code&gt;, providing roughly 28,232 ports).&lt;/li&gt;
&lt;li&gt;When connections close rapidly without connection reuse, sockets transition to the &lt;code&gt;TIME_WAIT&lt;/code&gt; state for the duration of &lt;code&gt;2 * MSL&lt;/code&gt; (typically 60 seconds).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Observed Behavior:&lt;/strong&gt; The proxy exhausts its ephemeral port pool, unable to open new sockets to upstream backends. Callers receive &lt;code&gt;HTTP 504 Gateway Timeout&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Database Connection Pool and Lock Exhaustion (Storage Layer)
&lt;/h3&gt;

&lt;p&gt;For requests that reach the relational database, naive inventory deduction relies on row-level pessimistic locking:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Anti-pattern: Naive pessimistic row lock&lt;/span&gt;
&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;inventory&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt; &lt;span class="k"&gt;FOR&lt;/span&gt; &lt;span class="k"&gt;UPDATE&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;inventory&lt;/span&gt; &lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;Warning: When row 42 is locked exclusively, the first transaction acquires the lock while subsequent concurrent transactions queue behind it. A standard application pool of 100 connections is fully saturated in milliseconds. Once the pool is depleted, every other endpoint sharing the pool—including authentication, user profiles, and product catalogs—stalls entirely.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Core Architectural Patterns for Flash Traffic
&lt;/h2&gt;

&lt;p&gt;To protect the database from lock contention, the architecture must filter non-viable requests in memory at the edge.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Atomic In-Memory Reservation (Redis and Lua)
&lt;/h3&gt;

&lt;p&gt;Stock validation and reservation are executed atomically in memory using an embedded Lua script on Redis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight lua"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Atomic stock reservation in Redis&lt;/span&gt;
&lt;span class="kd"&gt;local&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'get'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;stock&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="nb"&gt;tonumber&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stock&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;then&lt;/span&gt;
    &lt;span class="n"&gt;redis&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;'decr'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;KEYS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="c1"&gt;-- Reservation acquired&lt;/span&gt;
&lt;span class="k"&gt;else&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="c1"&gt;-- Insufficient stock; reject&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Latency:&lt;/strong&gt; Sub-millisecond execution (&amp;lt; 0.5 ms).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Throughput:&lt;/strong&gt; Only the initial 10 requests obtain a successful return code. The remaining 99,990 requests are rejected immediately at the ingress boundary without touching persistent storage.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Asynchronous Buffering (Kafka / Event Queue)
&lt;/h3&gt;

&lt;p&gt;Successful reservations are not persisted via synchronous writes. Instead, they are published to a durable append-only log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ Checkout Event ] ──► [ Kafka Topic: orders ] ──► [ Consumer Pool ]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This decouples request ingress from disk I/O, allowing persistent database writes to execute at a steady, controlled rate irrespective of the traffic burst.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Optimistic Concurrency Control on Settlement
&lt;/h3&gt;

&lt;p&gt;When background consumers persist the order into the primary relational store, they use Optimistic Concurrency Control (OCC) rather than exclusive locks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Recommended: Optimistic concurrency control&lt;/span&gt;
&lt;span class="k"&gt;UPDATE&lt;/span&gt; &lt;span class="n"&gt;inventory&lt;/span&gt; 
&lt;span class="k"&gt;SET&lt;/span&gt; &lt;span class="n"&gt;allocated&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;allocated&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;product_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;42&lt;/span&gt; &lt;span class="k"&gt;AND&lt;/span&gt; &lt;span class="k"&gt;version&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the version check fails due to a concurrent write, zero rows are updated. The worker backs off with jitter and retries. This eliminates database-level row lock waits and prevents thread pool starvation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Production Tuning Reference
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Failure Mode&lt;/th&gt;
&lt;th&gt;Mitigation Strategy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Linux Kernel&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;SYN backlog queue saturation&lt;/td&gt;
&lt;td&gt;Tune &lt;code&gt;sysctl -w net.core.somaxconn=65535&lt;/code&gt; and &lt;code&gt;net.ipv4.tcp_max_syn_backlog=65535&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Reverse Proxy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Ephemeral port exhaustion (&lt;code&gt;TIME_WAIT&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Enable HTTP keep-alives and &lt;code&gt;sysctl -w net.ipv4.tcp_tw_reuse=1&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Application&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Connection pool exhaustion&lt;/td&gt;
&lt;td&gt;Shift initial inventory decrement to in-memory store (Redis)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Database&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Lock queues and connection starvation&lt;/td&gt;
&lt;td&gt;Replace pessimistic row locks with optimistic concurrency control (&lt;code&gt;version&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Discussion
&lt;/h2&gt;

&lt;p&gt;When architecting high-demand flash inventory systems, does your team rely on edge virtual waiting rooms (such as Cloudflare Waiting Room or AWS Virtual Waiting Room) to shape ingress traffic, or do you absorb spikes in-memory using distributed caching and message brokers?&lt;/p&gt;

&lt;p&gt;Share your production experiences and trade-offs in the comments below.&lt;/p&gt;




&lt;h3&gt;
  
  
  Reference &amp;amp; Video Walkthrough
&lt;/h3&gt;

&lt;p&gt;The visual animated breakdown is available on YouTube:&lt;br&gt;
&lt;a href="https://www.youtube.com/watch?v=Hj0tPk9ZG3o&amp;amp;t=1s" rel="noopener noreferrer"&gt;Watch: What Actually Happens When 100,000 Users Click 'Buy' Concurrently&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;ScaleVector publishes deep-dive engineering breakdowns covering distributed systems, systems programming, and cloud infrastructure.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>scalability</category>
      <category>systemdesign</category>
    </item>
  </channel>
</rss>
