During a high-concurrency flash sale—such as ten limited-inventory items going live at 08:00 AM—traffic does not arrive in a smooth, manageable curve. It arrives as an instantaneous impulse function: 100,000 requests hitting the ingress layer within the same millisecond window.
In a conventional architecture where an API gateway proxies requests directly to an application cluster backed by a relational database, this volume will not merely oversell inventory. It causes a cascading system collapse across network, gateway, and storage layers.
Architecture at a Glance
Why unbuffered architectures fail during sudden demand spikes:
- Network Layer (Linux Kernel): Incoming SYN packets saturate kernel socket queues, causing silent connection drops before the application runtime receives execution time.
-
Gateway Layer (Reverse Proxy): High connection churn exhausts ephemeral ports into
TIME_WAITstate, surfacing cascadingHTTP 504 Gateway Timeouterrors. -
Storage Layer (Database): Pessimistic locking (
SELECT ... FOR UPDATE) causes queueing across connection pools, stalling unrelated services across the platform.
A resilient architecture shifts concurrency management out of the relational engine:
[ 100,000 Users ]
│
▼ (1. Rate limit & In-Memory Check)
[ Edge / Ingress ] ──(Lua Atomic Decr)──► [ Redis Cluster ]
│ │
│ (Only 10 requests pass) │ (99,990 rejected instantly)
▼ ▼
[ Kafka Queue ] ────────────────────────► [ Background Workers ]
│
▼ (Optimistic Lock)
[ PostgreSQL / SQL Server ]
Failure Modes Across the Stack
1. Kernel Backlog Saturation (Network Layer)
Before an HTTP payload reaches user space, the Linux kernel manages connection handshakes across two primary queues:
Client ──► [ SYN Queue ] ──► [ Accept Queue ] ──► Application accept()
▲
└── Saturated queue: Packets dropped silently
-
The Failure Mechanism: Default kernel configurations often set
net.core.somaxconnto conservative values (such as128or4096). Under an instantaneous 100,000-connection burst, the SYN backlog overflows immediately. -
Observed Behavior: Clients experience stalled TLS handshakes and connection timeouts (
ETIMEDOUT). The application logs zero errors because packets are dropped at the kernel boundary.
2. Ephemeral Port Starvation (Gateway Layer)
Reverse proxies and load balancers (such as Envoy, Nginx, or an AWS ALB) maintain outbound connections toward backend instances.
- Outbound connections allocate an ephemeral port from the local range (typically
32768–60999, providing roughly 28,232 ports). - When connections close rapidly without connection reuse, sockets transition to the
TIME_WAITstate for the duration of2 * MSL(typically 60 seconds). -
Observed Behavior: The proxy exhausts its ephemeral port pool, unable to open new sockets to upstream backends. Callers receive
HTTP 504 Gateway Timeout.
3. Database Connection Pool and Lock Exhaustion (Storage Layer)
For requests that reach the relational database, naive inventory deduction relies on row-level pessimistic locking:
-- Anti-pattern: Naive pessimistic row lock
SELECT stock FROM inventory WHERE id = 42 FOR UPDATE;
UPDATE inventory SET stock = stock - 1 WHERE id = 42;
Warning: When row 42 is locked exclusively, the first transaction acquires the lock while subsequent concurrent transactions queue behind it. A standard application pool of 100 connections is fully saturated in milliseconds. Once the pool is depleted, every other endpoint sharing the pool—including authentication, user profiles, and product catalogs—stalls entirely.
Core Architectural Patterns for Flash Traffic
To protect the database from lock contention, the architecture must filter non-viable requests in memory at the edge.
1. Atomic In-Memory Reservation (Redis and Lua)
Stock validation and reservation are executed atomically in memory using an embedded Lua script on Redis:
-- Atomic stock reservation in Redis
local stock = redis.call('get', KEYS[1])
if stock and tonumber(stock) > 0 then
redis.call('decr', KEYS[1])
return 1 -- Reservation acquired
else
return 0 -- Insufficient stock; reject
end
- Latency: Sub-millisecond execution (< 0.5 ms).
- Throughput: Only the initial 10 requests obtain a successful return code. The remaining 99,990 requests are rejected immediately at the ingress boundary without touching persistent storage.
2. Asynchronous Buffering (Kafka / Event Queue)
Successful reservations are not persisted via synchronous writes. Instead, they are published to a durable append-only log:
[ Checkout Event ] ──► [ Kafka Topic: orders ] ──► [ Consumer Pool ]
This decouples request ingress from disk I/O, allowing persistent database writes to execute at a steady, controlled rate irrespective of the traffic burst.
3. Optimistic Concurrency Control on Settlement
When background consumers persist the order into the primary relational store, they use Optimistic Concurrency Control (OCC) rather than exclusive locks:
-- Recommended: Optimistic concurrency control
UPDATE inventory
SET allocated = allocated + 1, version = version + 1
WHERE product_id = 42 AND version = 5;
If the version check fails due to a concurrent write, zero rows are updated. The worker backs off with jitter and retries. This eliminates database-level row lock waits and prevents thread pool starvation.
Production Tuning Reference
| Layer | Failure Mode | Mitigation Strategy |
|---|---|---|
| Linux Kernel | SYN backlog queue saturation | Tune sysctl -w net.core.somaxconn=65535 and net.ipv4.tcp_max_syn_backlog=65535
|
| Reverse Proxy | Ephemeral port exhaustion (TIME_WAIT) |
Enable HTTP keep-alives and sysctl -w net.ipv4.tcp_tw_reuse=1
|
| Application | Connection pool exhaustion | Shift initial inventory decrement to in-memory store (Redis) |
| Database | Lock queues and connection starvation | Replace pessimistic row locks with optimistic concurrency control (version) |
Discussion
When architecting high-demand flash inventory systems, does your team rely on edge virtual waiting rooms (such as Cloudflare Waiting Room or AWS Virtual Waiting Room) to shape ingress traffic, or do you absorb spikes in-memory using distributed caching and message brokers?
Share your production experiences and trade-offs in the comments below.
Reference & Video Walkthrough
The visual animated breakdown is available on YouTube:
Watch: What Actually Happens When 100,000 Users Click 'Buy' Concurrently
ScaleVector publishes deep-dive engineering breakdowns covering distributed systems, systems programming, and cloud infrastructure.
Top comments (0)