<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: MANGESH MANDLIK</title>
    <description>The latest articles on DEV Community by MANGESH MANDLIK (@mangeshmandlik).</description>
    <link>https://dev.to/mangeshmandlik</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4131208%2Fa9ac2096-0793-425e-8cf2-63f73f694fe7.jpg</url>
      <title>DEV Community: MANGESH MANDLIK</title>
      <link>https://dev.to/mangeshmandlik</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mangeshmandlik"/>
    <language>en</language>
    <item>
      <title>Designing Dropbox: File Storage &amp; Sync System Design</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Sun, 11 Oct 2026 06:55:50 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/designing-dropbox-file-storage-sync-system-design-o2k</link>
      <guid>https://dev.to/mangeshmandlik/designing-dropbox-file-storage-sync-system-design-o2k</guid>
      <description>&lt;p&gt;&lt;em&gt;Uploading a file is straightforward. Keeping its bytes, metadata, permissions, history, and copies on several intermittently connected devices consistent is the real system-design problem.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frzv4a8mz0k9yzfwy059f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frzv4a8mz0k9yzfwy059f.png" alt="complete architecture" width="800" height="367"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Consider a document named &lt;code&gt;report.docx&lt;/code&gt;. You edit it on your laptop, close the lid halfway through an upload, rename its folder on another device, and then open your phone after several days offline. Meanwhile, a collaborator may have edited the same document or lost access to it.&lt;/p&gt;

&lt;p&gt;A file storage and synchronization system must answer more than “where do we store the bytes?” It must know &lt;strong&gt;which version is current, which changes each device has observed, what to do when histories diverge, and whether a user is still authorized to download a particular version&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We'll design a Dropbox- or Google Drive-like system from a simple upload API to a reliable sync architecture. This is a conceptual design, not a claim about either company's internal implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Requirements and scale
&lt;/h2&gt;

&lt;p&gt;The system supports uploading and downloading large files, resumable transfers, folders, rename and move, automatic synchronization across devices, offline edits, version history, deletion recovery, and viewer/editor sharing.&lt;/p&gt;

&lt;p&gt;The priorities are durability, integrity, secure access, availability despite individual failures, reasonable synchronization freshness, and bandwidth efficiency. “Instant sync” and “never lose a file” are not defensible unconditional guarantees: network partitions, retention rules, disasters, and client availability impose real limits.&lt;/p&gt;

&lt;p&gt;For a scale exercise, suppose we have &lt;strong&gt;500 million registered users, 100 million daily active users, and 200 million uploads per day&lt;/strong&gt;. With an illustrative average upload of &lt;strong&gt;5 MB&lt;/strong&gt;, that's approximately &lt;strong&gt;1 PB of new uploaded data per day&lt;/strong&gt; before compression, deduplication, replication, or retention effects. The mean arrival rate is about &lt;strong&gt;2,315 uploads/second&lt;/strong&gt;, but peak capacity must be sized above the mean. Assume individual files can reach &lt;strong&gt;50 GB&lt;/strong&gt;. These are invented design inputs, not published statistics for Dropbox or Google Drive.&lt;/p&gt;

&lt;p&gt;Notice that upload &lt;em&gt;requests&lt;/em&gt;, bytes ingested, metadata operations, download traffic, and device sync connections scale differently. We should not expect one service or database to solve them all.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Start with the simplest upload architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client ── file bytes ──&amp;gt; Application server ──&amp;gt; Local disk
                             │
                             └───────────────&amp;gt; Metadata database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdskdfx0m63qeg8zvm4ca.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdskdfx0m63qeg8zvm4ca.png" alt="start simple" width="800" height="279"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This works for small workloads. But when a large upload is interrupted, a one-shot upload may have to restart. Application servers also spend bandwidth and connection capacity relaying bytes rather than handling application logic. Local disks complicate replication, capacity expansion, and recovery.&lt;/p&gt;

&lt;p&gt;Proxying bytes is not inherently wrong: inspection, transformations, restricted networks, or specialized clients can justify it. For general large-file transfers at our assumed scale, however, &lt;strong&gt;separating the control plane from the byte-transfer path&lt;/strong&gt; is a useful next step.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Separate file metadata from object storage
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                       ┌──────────────────────┐
Client ── API calls ──&amp;gt; │ File / Metadata API  │ ──&amp;gt; Metadata DB
                       └──────────────────────┘      identity, folders,
                                  │                 ACLs, versions, state
                                  │ signed capability
                                  ▼
Client ── file bytes directly ──&amp;gt; Object storage
                                 immutable version objects
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxp3g1kz97e1fd6po2ibr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxp3g1kz97e1fd6po2ibr.png" width="800" height="305"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A metadata record might contain &lt;code&gt;file_id&lt;/code&gt;, &lt;code&gt;owner_id&lt;/code&gt;, &lt;code&gt;parent_folder_id&lt;/code&gt;, &lt;code&gt;filename&lt;/code&gt;, &lt;code&gt;current_version_id&lt;/code&gt;, &lt;code&gt;lifecycle_state&lt;/code&gt;, and a concurrency token. A separate versions table stores &lt;code&gt;version_id&lt;/code&gt;, &lt;code&gt;file_id&lt;/code&gt;, &lt;code&gt;parent_version_id&lt;/code&gt;, content checksum, size, and object key. The file's &lt;strong&gt;stable ID is independent of its path&lt;/strong&gt;: renaming &lt;code&gt;report.docx&lt;/code&gt; or moving it into another folder changes metadata, not the file's identity or necessarily its stored bytes.&lt;/p&gt;

&lt;p&gt;Folders also have stable IDs and parent relationships. Listing a folder means querying its immediate children with pagination. A folder move can change the folder's parent pointer without rewriting every descendant, although subtree permissions, path caches, and asynchronous indexes may still require careful handling. If sibling names must be unique, enforce that invariant within the correct namespace, including the root-folder case.&lt;/p&gt;

&lt;p&gt;Object storage holds large binary objects; the metadata store handles queryable relationships and conditional updates. This is a &lt;strong&gt;workload-driven decision&lt;/strong&gt;, not a rule that databases can never store small blobs.&lt;/p&gt;

&lt;h3&gt;
  
  
  The metadata–object consistency gap
&lt;/h3&gt;

&lt;p&gt;The database transaction that publishes a file version cannot ordinarily atomically include the object-storage upload. Therefore, use an explicit lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CREATED → UPLOADING → VERIFYING → AVAILABLE
                ↘ FAILED / EXPIRED
AVAILABLE → DELETING → DELETED (subject to retention)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A safe publication sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create a metadata record and an upload session in a non-available state.&lt;/li&gt;
&lt;li&gt;Upload bytes to a server-selected object key.&lt;/li&gt;
&lt;li&gt;Complete the object upload and verify the integrity evidence supported by the storage API.&lt;/li&gt;
&lt;li&gt;Atomically update the file's current-version pointer &lt;strong&gt;and record its sync change&lt;/strong&gt; using a transaction or another durable, correctly ordered publication mechanism.&lt;/li&gt;
&lt;li&gt;Reconcile abandoned sessions and orphan objects asynchronously.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The fourth step matters. If a file becomes visible in metadata but its corresponding change event is lost, other devices may never discover it. A &lt;strong&gt;transactional outbox&lt;/strong&gt; or a change log committed alongside metadata can bridge that boundary; publishing to an unrelated queue after the database commit without recovery is not enough.&lt;/p&gt;

&lt;p&gt;If storage succeeds but the metadata commit fails, the object may be orphaned. A reconciler checks the session, object, checksum, and current metadata state before safely retrying publication or scheduling deletion after a grace period. If metadata says &lt;code&gt;AVAILABLE&lt;/code&gt; but the object is missing, fail visibly and investigate or recover from a replica/backup; never serve an empty object as success.&lt;/p&gt;

&lt;p&gt;For deeper treatment, see &lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/architecture-deep-dive" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Architecture Deep Dive&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Design the upload API and idempotency contracts
&lt;/h2&gt;

&lt;p&gt;An illustrative API is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;POST /files&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Create file metadata; client supplies an idempotency key&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;POST /files/{id}/uploads&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Start or resume an upload session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;POST /files/{id}/complete&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Complete and verify a session; safe to retry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /files/{id}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Return metadata and current version&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /files/{id}/download&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Authorize and issue a short-lived download capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;PATCH /files/{id}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Rename/move with a version precondition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GET /sync/changes?cursor=...&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Fetch durable changes for a device&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are example contracts, not a mandated REST standard. A production implementation must also define deletion, restore, sharing, version retrieval, and session-status endpoints.&lt;/p&gt;

&lt;p&gt;A retry of &lt;code&gt;POST /files&lt;/code&gt; must not accidentally create two files. An &lt;strong&gt;idempotency key scoped to the caller and operation&lt;/strong&gt; allows the server to return the same logical result. Completing an upload should similarly be retry-safe: a timeout may mean “completion succeeded but the response was lost.” Repeating that request must not publish a duplicate version.&lt;/p&gt;

&lt;p&gt;Idempotency handles retries of the &lt;em&gt;same&lt;/em&gt; operation. It does not solve two genuinely concurrent edits; those require version preconditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Direct uploads and presigned URL security
&lt;/h2&gt;

&lt;p&gt;For the normal large-file path, the application authenticates the caller, checks permission and quota, creates the session, and returns a short-lived presigned upload URL. The client sends bytes directly to object storage.&lt;/p&gt;

&lt;p&gt;A presigned URL is generally a &lt;strong&gt;bearer capability&lt;/strong&gt;. Anyone who obtains it can potentially use it until expiration, subject to signature and storage-policy constraints. It is &lt;strong&gt;not automatically bound to the original user's identity&lt;/strong&gt;. Therefore:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authorize &lt;em&gt;before&lt;/em&gt; issuing it; generate object keys on the server.&lt;/li&gt;
&lt;li&gt;Scope it to a specific operation and object or multipart part.&lt;/li&gt;
&lt;li&gt;Use HTTPS, short expirations, appropriate request constraints where supported, and avoid leaking URLs in logs.&lt;/li&gt;
&lt;li&gt;Configure browser CORS separately; CORS is not a replacement for authorization.&lt;/li&gt;
&lt;li&gt;Recheck authorization and the expected version before publishing a completed upload, because permissions or file state may have changed while bytes were in flight.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Revoking a user's ACL does &lt;strong&gt;not automatically revoke a previously issued, unexpired URL&lt;/strong&gt;. If near-immediate revocation is mandatory, an edge or proxy path that validates current authorization on access may be necessary. Even then, bytes already downloaded cannot be recalled.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Resumable multipart uploads
&lt;/h2&gt;

&lt;p&gt;A single 10 GB PUT still has an expensive failure mode: a dropped connection can force a complete restart. Native object-storage multipart upload lets the client send separately acknowledged parts.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Upload session U17
  part 1 ──────── ✓
  part 2 ──────── ✓
  part 3 ── X     retry part 3
  part 4 ──────── ✓
                ↓
  complete known parts → verify → publish new version
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmecayjmhuughcq10nnri.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmecayjmhuughcq10nnri.png" width="800" height="328"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A durable session tracks its owner, file ID, expected base version, storage upload ID, expiration, and status. The client can upload parts concurrently and, after reconnecting, &lt;strong&gt;reconcile confirmed parts against the storage provider's actual state&lt;/strong&gt; where supported. Its own local progress counter may be stale if a part succeeded but the response was lost.&lt;/p&gt;

&lt;p&gt;Part size is a trade-off: smaller parts reduce failed-retry cost but increase request and tracking overhead; larger parts reduce coordination but increase retry cost and buffering. Provider minimums, maximum part counts, file size, client memory, and network behavior determine the right choice. Don't treat a single part size as universal across providers.&lt;/p&gt;

&lt;p&gt;An object-store multipart part is not necessarily the same abstraction as an application-level delta-sync chunk. Native multipart simplifies upload assembly; a content-addressed chunk store is a separate, substantially more complex design.&lt;/p&gt;

&lt;h3&gt;
  
  
  Integrity is not the same as upload completion
&lt;/h3&gt;

&lt;p&gt;TLS protects bytes in transit, but it does not prove that the application selected the correct file. Where supported, validate per-part checksums and the assembled object's expected checksum. A mismatch in the final whole-object hash shows that something differs; &lt;strong&gt;it does not identify the corrupted part&lt;/strong&gt; without finer-grained evidence.&lt;/p&gt;

&lt;p&gt;Do not assume an object-storage &lt;code&gt;ETag&lt;/code&gt; is always an MD5 digest of the entire file. ETag semantics vary by provider and upload/encryption method, especially for multipart objects. Use explicit checksum mechanisms and clearly specify what has been verified before publication.&lt;/p&gt;

&lt;p&gt;Expired sessions and abandoned multipart uploads require cleanup. Unassembled parts may still consume storage and cost money; expiry policies and periodic reconciliation are part of the design, not optional housekeeping.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Downloads, range requests, and private CDN delivery
&lt;/h2&gt;

&lt;p&gt;For downloads, the file API checks the current ACL and returns a short-lived signed capability for a specific immutable file version. The client downloads bytes from storage, or through an appropriately secured CDN where caching is useful.&lt;/p&gt;

&lt;p&gt;HTTP &lt;code&gt;Range&lt;/code&gt; requests can resume an interrupted download or retrieve part of a large file. A robust client should ensure the resumed ranges belong to the &lt;strong&gt;same immutable object version&lt;/strong&gt; and verify the completed file's expected checksum. Simply appending bytes from a mutable URL can mix different versions.&lt;/p&gt;

&lt;p&gt;CDN caching is workload-dependent. Frequently requested, immutable shared objects may benefit; private files downloaded once per user may not. A private CDN needs &lt;strong&gt;its own edge authorization and private-origin configuration&lt;/strong&gt;. A presigned origin-storage URL does not magically make a CDN path private or authorized. Already-issued CDN capabilities have similar revocation windows to signed origin URLs.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Multi-device sync is a reconciliation protocol
&lt;/h2&gt;

&lt;p&gt;A sync client needs persistent local state, not merely a filesystem watcher. For each file it tracks the stable remote ID, local path, last synchronized server version, local fingerprint, pending operations, and its durable checkpoint.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Laptop: edit report.docx
  local state: LOCAL_DIRTY
           ↓
  upload candidate version based on v5
           ↓
  server checks current version == v5
           ↓
  publish v6 + durable change record
           ↓
  push hint to other online devices
           ↓
Phone: GET /sync/changes?cursor=C
           ↓
  apply v6, then advance local cursor
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Local filesystem notifications can be missed, so a periodic reconciliation scan is a useful safety net. Server push via WebSocket, mobile push, or long polling is &lt;strong&gt;only a wake-up hint&lt;/strong&gt;. The durable, cursor-based change log is the source of truth.&lt;/p&gt;

&lt;p&gt;A cursor identifies a position in an ordered change stream &lt;strong&gt;within its defined scope&lt;/strong&gt;. It is not simply the client's wall-clock timestamp. The client must not persist a cursor beyond changes it has durably applied; otherwise a crash can permanently skip an update. Applying duplicate events must be safe through operation IDs, version checks, and idempotent local updates.&lt;/p&gt;

&lt;p&gt;One subtlety: when changes are paginated or generated concurrently, the server must define a consistent ordering/checkpoint contract so clients cannot skip a late-committed entry with an earlier cursor position. A snapshot-plus-change-stream handoff likewise needs a defined boundary; “take a snapshot, then start polling” without such a boundary can miss mutations.&lt;/p&gt;

&lt;p&gt;If a device has been offline longer than change-log retention, its old cursor may be invalid. It must obtain a consistent namespace snapshot, compare it with local state, reconcile offline operations, and establish a fresh cursor without losing changes during the transition.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Whole-file hashing is not delta synchronization
&lt;/h2&gt;

&lt;p&gt;A content hash answers: &lt;strong&gt;are these bytes equal?&lt;/strong&gt; It does not answer: &lt;strong&gt;which blocks changed?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are three practical strategies:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Strategy&lt;/th&gt;
&lt;th&gt;Works well when&lt;/th&gt;
&lt;th&gt;Main limitation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Whole-file replacement&lt;/td&gt;
&lt;td&gt;Files are small or binary deltas are ineffective&lt;/td&gt;
&lt;td&gt;Transfers the entire new file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fixed-size block comparison&lt;/td&gt;
&lt;td&gt;Large files change in place without shifting offsets&lt;/td&gt;
&lt;td&gt;Insertions shift later block boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rolling-signature or content-defined chunking&lt;/td&gt;
&lt;td&gt;Large files have localized edits and reusable regions&lt;/td&gt;
&lt;td&gt;More CPU, signatures, chunk metadata, and GC complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;An rsync-style rolling matching algorithm and content-defined chunking are related approaches, &lt;strong&gt;not identical algorithms&lt;/strong&gt;. Both can locate reusable regions despite offset shifts, but their chunking and matching procedures differ.&lt;/p&gt;

&lt;p&gt;For compressed archives, encrypted data, and some document formats, a tiny logical edit can change much of the serialized byte stream. In those cases, a delta algorithm may save little bandwidth. The client should choose based on measured transfer savings versus computation and coordination cost, not promise that every one-paragraph edit uploads only a few kilobytes.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Detect conflicts using version lineage
&lt;/h2&gt;

&lt;p&gt;Suppose both laptop and phone begin with &lt;code&gt;report.docx&lt;/code&gt; at version &lt;code&gt;v5&lt;/code&gt;. The laptop edits and successfully publishes &lt;code&gt;v6&lt;/code&gt;. The phone was offline and independently edited its copy based on &lt;code&gt;v5&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             v5 (shared base)
             /            \
   laptop candidate     phone candidate
        accepted v6       base=v5
             │              │
       current=v6      compare-and-set fails
                            │
                      preserve as conflict
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server accepts a new current version only when its &lt;strong&gt;expected base version matches the current version&lt;/strong&gt;. A stale writer must not silently overwrite the newer version. This check must be atomic at the metadata commit boundary.&lt;/p&gt;

&lt;p&gt;Possible policies include preserving a conflict copy, last-write-wins, or a three-way merge for supported text formats. Conflict copies reduce silent overwrites but require user intervention and do not guarantee zero data loss if the underlying storage or backup system fails. Last-write-wins is simple but can discard edits. Automatic merging is format-dependent; it is not a general solution for arbitrary binary files. Real-time collaborative editing with OT or CRDTs is a different problem from background file synchronization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Metadata edits also race.&lt;/strong&gt; Rename and move operate on stable IDs; delete should produce a tombstone that offline devices can observe. When a deleted file receives a stale offline edit, the product must explicitly choose to reject it, restore a conflict copy, or create a recovered file. A naive sync client that sees a missing file and uploads its old local copy risks resurrecting a deletion.&lt;/p&gt;

&lt;p&gt;Content version checks and metadata-operation concurrency tokens may need separate treatment: a rename need not create a new content blob, yet it must still be protected against conflicting namespace mutations.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Sharing and access control
&lt;/h2&gt;

&lt;p&gt;The metadata layer records ownership, direct shares, and possibly inherited folder permissions. Each operation checks whether the caller may read, edit, rename, delete, share, or restore the target. Folder inheritance and override rules must be defined rather than guessed.&lt;/p&gt;

&lt;p&gt;For a file shared by Alice with Bob, Bob requests a download through the application, which checks Bob's access and issues a scoped signed URL. Direct object-store access &lt;strong&gt;after&lt;/strong&gt; that check is intentional. The file bytes need not pass through the authorization service.&lt;/p&gt;

&lt;p&gt;A recipient-oriented index for “shared with me” avoids scanning every owner's ACLs as metadata is sharded. That index must be maintained carefully, but &lt;strong&gt;an eventually updated discovery index must not become the sole authority for granting access&lt;/strong&gt;. Authorization-sensitive paths should check authoritative permission state or an explicitly safe consistency mechanism. If current authorization cannot be determined, fail closed.&lt;/p&gt;

&lt;p&gt;Malware scanning can also be a security gate. If the product requires scanning before sharing or downloading, a failed scanner must leave the file quarantined or pending, not silently mark it usable. Thumbnail generation and search indexing, by contrast, can often fail without making the file itself unavailable.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Scaling the metadata and sync planes
&lt;/h2&gt;

&lt;p&gt;Object storage handles bulk bytes; the metadata database handles frequently changing file, folder, permission, and version records. Read replicas may help appropriate read workloads, but stale replicas are unsafe for decisions requiring the latest ACL or concurrency state.&lt;/p&gt;

&lt;p&gt;Owner-based sharding keeps a user's namespace relatively local. Its weakness is a very large or active user/tenant, and cross-owner “shared with me” queries. Tenant-based partitioning improves isolation but can produce giant-tenant hotspots. File-ID sharding spreads load but requires secondary indexes for folder and owner queries. There is no universally best shard key.&lt;/p&gt;

&lt;p&gt;A huge folder should be listed with pagination, not enumerated in a single response. Fan-out notifications and derived indexes should be processed asynchronously through durable jobs. Per-tenant fairness and backpressure protect other users when one tenant generates extreme traffic.&lt;/p&gt;

&lt;p&gt;The sync change stream also has independent limits: write throughput, cursor retention, ordering scope, connected-device count, and notification fan-out. Partitioning by user can help localize ordered processing &lt;strong&gt;if the underlying stream actually provides that ordering&lt;/strong&gt;, but it does not replace operation idempotency or atomic version preconditions.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. Failure scenarios that reveal correctness bugs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure&lt;/th&gt;
&lt;th&gt;Safe behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Upload interrupted&lt;/td&gt;
&lt;td&gt;Reconcile storage-confirmed parts; retry missing parts only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Upload succeeds, metadata commit fails&lt;/td&gt;
&lt;td&gt;Idempotent completion plus reconciliation; never publish unverified bytes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Presigned URL expires&lt;/td&gt;
&lt;td&gt;Check session status and refresh URLs for remaining parts if still valid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metadata primary unavailable&lt;/td&gt;
&lt;td&gt;Fail over safely or reject mutations; don't acknowledge writes buffered only in memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Storage temporarily unavailable&lt;/td&gt;
&lt;td&gt;Bound retries with backoff and jitter; expose degraded byte-transfer state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sync push unavailable&lt;/td&gt;
&lt;td&gt;Poll durable change log on reconnect or schedule&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicate/out-of-order sync delivery&lt;/td&gt;
&lt;td&gt;Use operation IDs, version preconditions, and safe cursor checkpoints&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Device offline beyond cursor retention&lt;/td&gt;
&lt;td&gt;Full snapshot reconciliation with a safe change-stream handoff&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ACL revoked after URL issuance&lt;/td&gt;
&lt;td&gt;Recognize capability expiry window; stricter validation if required&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Abandoned multipart sessions&lt;/td&gt;
&lt;td&gt;Abort expired sessions and reclaim parts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Garbage collector thinks a live object is orphaned&lt;/td&gt;
&lt;td&gt;Grace period, durable deletion state, reference checks, and independent audit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regional outage&lt;/td&gt;
&lt;td&gt;Follow defined RPO/RTO; prevent split-brain writes during failover&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The garbage collector deserves special attention. An object may appear unreferenced because a metadata query failed, a replication lagged, or a deduplication reference has not yet been observed. &lt;strong&gt;Never hard-delete a blob solely because one transient lookup returns no references.&lt;/strong&gt; Use retention windows, tombstones, durable GC state, rechecks, and recoverable backups.&lt;/p&gt;

&lt;p&gt;Cross-user content deduplication is optional, not a free storage win. It complicates reference counting and deletion, and a naive “does this hash already exist?” endpoint can leak information about another user's private data. Start without global deduplication unless measured storage savings justify its security and operational complexity.&lt;/p&gt;

&lt;p&gt;For a scenario-by-scenario breakdown, see &lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/failure-scale-scenarios" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Failure &amp;amp; Scale Scenarios&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  14. Architectural trade-offs worth defending
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Direct transfer or application proxy?&lt;/strong&gt; Direct signed upload/download reduces application bandwidth at scale. Proxying remains appropriate for mandatory inline inspection, transformation, or real-time authorization controls.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Native multipart or application-managed chunk store?&lt;/strong&gt; Native multipart simplifies reliable large uploads. A content-addressed chunk store can support cross-version reuse and deduplication but introduces reference-counting, garbage-collection, and privacy complexity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Push or polling?&lt;/strong&gt; Push reduces notification delay for connected clients. Polling is a necessary fallback. Neither replaces the durable change log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strong consistency or eventual consistency?&lt;/strong&gt; The &lt;em&gt;acceptance of a new current version&lt;/em&gt; needs a correct atomic precondition. Other devices may observe that committed version later. These are separate consistency boundaries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Single-region or multi-region?&lt;/strong&gt; A single-region primary with tested backups and a defined disaster-recovery plan is a defensible baseline. Cross-region replication can improve recovery objectives but requires handling replication lag and preventing two primaries from accepting conflicting writes during a partition. Object durability is not the same as an RPO or RTO guarantee.&lt;/p&gt;

&lt;p&gt;For alternatives and decision criteria, see &lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/tradeoffs-decisions" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Trade-offs &amp;amp; Decisions&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  15. System design interview follow-ups
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why not keep file bytes in PostgreSQL?&lt;/strong&gt; Small blobs may be fine there. At the stated multi-GB-file and petabyte-scale ingest workload, separate object storage better fits byte capacity, transfer, and lifecycle requirements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when a 50 GB upload fails at 90%?&lt;/strong&gt; Recover the durable multipart session, reconcile parts confirmed by storage, refresh expired upload capabilities if the session is valid, and retry only missing parts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can the client trust ETag as an MD5?&lt;/strong&gt; No. Provider, multipart, and encryption behavior affect its meaning. Specify explicit checksum verification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do devices discover missed changes?&lt;/strong&gt; Fetch from the durable change log using a cursor. Push is a hint; polling and reconnect reconciliation provide recovery.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does SHA-256 identify changed blocks?&lt;/strong&gt; A single whole-file SHA-256 does not. Delta algorithms need block-level or rolling/content-defined signatures.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if two offline devices edit the same file?&lt;/strong&gt; Both carry the base version they edited. Atomic compare-and-set accepts one current version; the divergent edit is handled by an explicit conflict policy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can you instantly revoke an issued presigned URL?&lt;/strong&gt; Not by changing only the file ACL. The already-issued capability may remain usable until expiry unless the access path supports further revocation checks or invalidation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if a device reconnects after six months?&lt;/strong&gt; If its cursor is older than retained history, take a consistent snapshot, reconcile queued local changes, and rejoin the change stream at a safe checkpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you prevent sync from losing a metadata mutation?&lt;/strong&gt; Persist the mutation and its change-log/outbox record atomically or use an equivalent recoverable publication protocol. A best-effort queue publish after commit can lose events.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you recover from accidental object deletion?&lt;/strong&gt; Recovery depends on version retention, backups, and provider replication capabilities. A design that promises recovery must test its restore procedure and define acceptable loss and recovery time.&lt;/p&gt;

&lt;p&gt;A larger question bank is available in &lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/interview-perspective" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Interview Perspective&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         ┌──────────────────────────┐
Laptop / Phone / Web ───&amp;gt; │ API / Auth / File Service│ ──&amp;gt; Metadata DB
        │                └──────────┬───────────────┘     file IDs, ACLs,
        │                           │                     versions, folders
        │                    signed upload/download
        │                           │
        └──────── file bytes ───────┴───────────────&amp;gt; Object Storage
        │                                           immutable version blobs
        │
        └── Sync client ──&amp;gt; Sync API ──&amp;gt; Durable change log / outbox
                               │                    │
                               └── cursor reads &amp;lt;───┘
                                      │
                           Push hints / polling fallback

              Background workers: scanning, previews, indexing,
                    session reconciliation, retention / GC
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important insight is not the number of boxes. It is the &lt;strong&gt;boundaries between guarantees&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Object storage protects and serves bytes; metadata identifies which bytes belong to which logical file version.&lt;/li&gt;
&lt;li&gt;A verified upload is not visible until its version is safely published.&lt;/li&gt;
&lt;li&gt;An accepted edit is conditionally committed; its arrival on other devices is asynchronous.&lt;/li&gt;
&lt;li&gt;Push improves freshness, but durable replay preserves synchronization correctness.&lt;/li&gt;
&lt;li&gt;Authorization gates capability issuance; signed URLs have explicit lifetime and revocation limitations.&lt;/li&gt;
&lt;li&gt;Garbage collection must not destroy data that another component still references.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A good file storage and sync design makes those boundaries explicit. That is what turns a basic upload feature into a system that can survive interrupted transfers, offline edits, permission changes, retries, and operational failures without silently corrupting user state.&lt;/p&gt;

&lt;h2&gt;
  
  
  References and further reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/complete-design" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Complete Design&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/architecture-deep-dive" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Architecture Deep Dive&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/tradeoffs-decisions" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Trade-offs &amp;amp; Decisions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/failure-scale-scenarios" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Failure &amp;amp; Scale Scenarios&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/file-storage-sync/interview-perspective" rel="noopener noreferrer"&gt;File Storage &amp;amp; Sync — Interview Perspective&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>systemdesign</category>
      <category>architecture</category>
      <category>backend</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>Designing Search &amp; Autocomplete: Inverted Index, Trie, Ranking &amp; Scale</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Sat, 10 Oct 2026 04:10:33 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/designing-search-autocomplete-inverted-index-trie-ranking-scale-13k5</link>
      <guid>https://dev.to/mangeshmandlik/designing-search-autocomplete-inverted-index-trie-ranking-scale-13k5</guid>
      <description>&lt;p&gt;&lt;em&gt;The difficult part of search isn't finding a word. It's returning the right results quickly while documents change, users type, and parts of the search infrastructure fail.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4jwfynfnwc7l3u3899s2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4jwfynfnwc7l3u3899s2.png" alt="complete architecture" width="800" height="386"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Imagine an e-commerce catalog with millions of products. A user types &lt;code&gt;iph&lt;/code&gt;, sees suggestions, completes &lt;code&gt;iphone case&lt;/code&gt;, applies a price filter, and expects relevant results almost immediately.&lt;/p&gt;

&lt;p&gt;A database query can support this at small scale. At larger scale, the system must solve several distinct problems: &lt;strong&gt;text retrieval, suggestion generation, ranking, freshness, filtering, and distributed execution&lt;/strong&gt;. Treating all of them as one database lookup makes the design harder to reason about.&lt;/p&gt;

&lt;p&gt;This article develops a search and autocomplete system from requirements through production failure modes. The example is a product catalog, but the underlying techniques also apply to document search, content discovery, and entity search.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Requirements: What Are We Actually Building?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Functional requirements
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Search products by keywords across titles, descriptions, and selected attributes.&lt;/li&gt;
&lt;li&gt;Show autocomplete suggestions as users type.&lt;/li&gt;
&lt;li&gt;Rank results by relevance, with optional popularity and recency signals.&lt;/li&gt;
&lt;li&gt;Support filters (brand, category, price, availability) and facets (counts by brand or category).&lt;/li&gt;
&lt;li&gt;Handle pagination and, where useful, typo tolerance.&lt;/li&gt;
&lt;li&gt;Reflect product additions, edits, and deletions within an agreed freshness window.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Suggestions can be &lt;strong&gt;query completions&lt;/strong&gt; (&lt;code&gt;iphone 15 case&lt;/code&gt;), &lt;strong&gt;entity suggestions&lt;/strong&gt; (a particular product), or a mixture. This decision affects indexing and ranking; a trie of words is not automatically a complete autocomplete product.&lt;/p&gt;

&lt;h3&gt;
  
  
  Non-functional requirements
&lt;/h3&gt;

&lt;p&gt;For discussion, assume:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Illustrative target or assumption&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Catalog&lt;/td&gt;
&lt;td&gt;10 million searchable products&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full search&lt;/td&gt;
&lt;td&gt;2,000 requests/second at peak&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autocomplete&lt;/td&gt;
&lt;td&gt;10,000 requests/second at peak&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Full-search latency&lt;/td&gt;
&lt;td&gt;P99 below 200 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autocomplete latency&lt;/td&gt;
&lt;td&gt;P99 below 100 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Index freshness&lt;/td&gt;
&lt;td&gt;Most updates visible within several seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Availability&lt;/td&gt;
&lt;td&gt;Continue serving when individual nodes fail, subject to product correctness rules&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;These are design assumptions, not universal search-industry benchmarks.&lt;/strong&gt; The actual requirements depend on corpus size, languages, query mix, hardware, and product expectations.&lt;/p&gt;

&lt;p&gt;A crucial distinction: &lt;strong&gt;search correctness, ranking quality, freshness, and availability are different properties&lt;/strong&gt;. A query can return valid matches in the wrong order; an index can be available but stale; a partial result can be fast but incomplete.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Why Not Just Use SQL &lt;code&gt;LIKE&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;The simplest implementation queries the transactional database directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;title&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="k"&gt;LOWER&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;LIKE&lt;/span&gt; &lt;span class="s1"&gt;'%iphone%'&lt;/span&gt;
&lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A leading-wildcard pattern such as &lt;code&gt;%iphone%&lt;/code&gt; generally cannot use a conventional B-tree index as an efficient prefix range lookup. On a large table, it may require expensive scanning.&lt;/p&gt;

&lt;p&gt;But the correct conclusion is &lt;strong&gt;not&lt;/strong&gt; that relational databases cannot do full-text search. PostgreSQL, for example, supports full-text indexes and trigram-based approaches. Database-native search may be the right choice for a smaller catalog or modest query volume.&lt;/p&gt;

&lt;p&gt;A dedicated search engine becomes attractive when we need several capabilities together: inverted indexes, language-aware analysis, flexible ranking, fuzzy matching, faceting, and independent scaling of read-heavy search traffic. The trade-off is another data system and an eventually consistent copy of source data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Design decision:&lt;/strong&gt; keep the transactional database authoritative and maintain a separate, query-optimized search index for this example.&lt;/p&gt;

&lt;p&gt;For a deeper comparison, see &lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/tradeoffs-decisions" rel="noopener noreferrer"&gt;Database vs Dedicated Search Index&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. High-Level Search &amp;amp; Autocomplete Architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                        WRITE / INDEXING PATH

  Product Service
        |
        v
  Source Database  (source of truth)
        |
        v
  CDC / Reliable Change Events
        |
        v
  Durable Stream / Queue
        |
        v
  Indexing Workers ----&amp;gt; Transform + Analyze + Version Check
        |                              |
        +------------------------------+
        |                              |
        v                              v
  Full-Text Search Index       Autocomplete Suggestion Index
  (inverted index)            (prefix / n-gram / precomputed)

                        READ / QUERY PATH

  Browser / App
        |
        v
  Search API / Gateway
        |
        +---- /suggest ---&amp;gt; Suggestion Cache ---&amp;gt; Suggestion Service
        |                                             |
        |                                             v
        |                                      Prefix Candidates
        |                                             |
        |                                      Suggestion Ranking
        |
        +---- /search ----&amp;gt; Result Cache ----&amp;gt; Search Coordinator
                                                  |
                                      +-----------+-----------+
                                      |           |           |
                                   Shard A     Shard B     Shard C
                                      |           |           |
                                      +-----------+-----------+
                                                  |
                                          Merge / Rank Top-K
                                                  |
                                             Search Results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The diagram separates &lt;strong&gt;indexing&lt;/strong&gt; from &lt;strong&gt;serving&lt;/strong&gt;. The search engine does not need to query the source database for every request. That separation protects transactional workloads and allows search capacity to scale independently.&lt;/p&gt;

&lt;p&gt;The diagram also shows a dedicated suggestion index. This is a &lt;strong&gt;choice for our assumed low-latency, high-autocomplete-QPS workload&lt;/strong&gt;, not a requirement for every implementation. A unified search engine with prefix queries can be simpler and entirely sufficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. How an Inverted Index Works
&lt;/h2&gt;

&lt;p&gt;A full-text search engine typically uses an &lt;strong&gt;inverted index&lt;/strong&gt;, mapping terms to the documents containing them.&lt;/p&gt;

&lt;p&gt;Suppose the catalog contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;D1: "Wireless iPhone case"
D2: "iPhone fast charger"
D3: "Wireless charging pad"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After suitable text analysis, a simplified index is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"iphone"   -&amp;gt; [D1, D2]
"wireless" -&amp;gt; [D1, D3]
"case"     -&amp;gt; [D1]
"charger"  -&amp;gt; [D2]
"charging" -&amp;gt; [D3]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A query for &lt;code&gt;wireless iphone&lt;/code&gt; can retrieve candidate documents by intersecting or combining the posting lists, depending on AND/OR query semantics. The search engine then scores the candidates.&lt;/p&gt;

&lt;p&gt;Real &lt;strong&gt;posting lists&lt;/strong&gt; may also store term frequencies and positions. Positions allow phrase queries to distinguish &lt;code&gt;wireless iphone case&lt;/code&gt; from the same words appearing far apart.&lt;/p&gt;

&lt;p&gt;Term lookup can be fast, but &lt;strong&gt;processing the posting list is not constant-time regardless of corpus size&lt;/strong&gt;. A common term can appear in millions of documents. Compression, skip structures, and top-k pruning help control that cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  Text analysis must be compatible
&lt;/h3&gt;

&lt;p&gt;Indexing and querying must use compatible analyzers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Raw text -&amp;gt; Tokenize -&amp;gt; Normalize -&amp;gt; Optional stemming/synonyms -&amp;gt; Terms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Normalization:&lt;/strong&gt; case folding and, when appropriate, accent handling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tokenization:&lt;/strong&gt; splitting text into meaningful units; language-dependent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stemming/lemmatization:&lt;/strong&gt; optional normalization of related word forms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synonyms:&lt;/strong&gt; domain-specific mappings such as &lt;code&gt;tv&lt;/code&gt; and &lt;code&gt;television&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Field-specific behavior:&lt;/strong&gt; product titles may be analyzed; SKUs and IDs usually require exact matching.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Applying aggressive stemming to an SKU or brand field can reduce correctness rather than improve it. Likewise, changing an analyzer can require &lt;strong&gt;reindexing existing documents&lt;/strong&gt;, because the stored terms were generated under the previous rules.&lt;/p&gt;

&lt;p&gt;For the internals, see &lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/architecture-deep-dive" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Architecture Deep Dive&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Autocomplete System Design: Generating Suggestions
&lt;/h2&gt;

&lt;p&gt;Autocomplete is a different workload from full search. A user may issue several suggestion requests before submitting one search, and the usefulness of a response drops quickly as they type the next character.&lt;/p&gt;

&lt;p&gt;The serving pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Typed prefix -&amp;gt; Candidate generation -&amp;gt; Suggestion ranking -&amp;gt; Top suggestions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There are several legitimate candidate-generation approaches.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Strength&lt;/th&gt;
&lt;th&gt;Main trade-off&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Search-engine prefix query&lt;/td&gt;
&lt;td&gt;Reuses existing search infrastructure&lt;/td&gt;
&lt;td&gt;Shares resources with full search; performance depends on engine and load&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edge n-grams&lt;/td&gt;
&lt;td&gt;Prefixes are indexed as terms for fast lookup&lt;/td&gt;
&lt;td&gt;More index storage and write amplification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Trie / FST-like structure&lt;/td&gt;
&lt;td&gt;Purpose-built prefix traversal; can precompute top suggestions&lt;/td&gt;
&lt;td&gt;Additional update, memory, and operational complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precomputed query-log suggestions&lt;/td&gt;
&lt;td&gt;Efficient for frequently searched queries&lt;/td&gt;
&lt;td&gt;Requires freshness, spam filtering, and safe handling of query logs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Trie example
&lt;/h3&gt;

&lt;p&gt;For the terms &lt;code&gt;car&lt;/code&gt;, &lt;code&gt;card&lt;/code&gt;, and &lt;code&gt;care&lt;/code&gt;, a conceptual trie looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(root)
   |
   c
   |
   a
   |
   r  ("car")
  / \
 d   e
("card") ("care")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looking up &lt;code&gt;car&lt;/code&gt; finds the node for that prefix. Retrieving the best suggestions can still require traversing many descendants &lt;strong&gt;unless the structure maintains precomputed top-k suggestions at nodes&lt;/strong&gt;. Prefix lookup and ranking are separate operations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Edge n-grams example
&lt;/h3&gt;

&lt;p&gt;For &lt;code&gt;iphone&lt;/code&gt;, index prefixes such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ip, iph, ipho, iphon, iphone
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then a query for &lt;code&gt;iph&lt;/code&gt; can match the corresponding prefix term. In practice, minimum and maximum prefix lengths should be configured deliberately: indexing every single-character prefix may be wasteful, and very short prefixes often produce noisy suggestions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ranking suggestions
&lt;/h3&gt;

&lt;p&gt;Prefix matching only produces candidates. Suggestions can then be ordered by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Historical search or selection frequency.&lt;/li&gt;
&lt;li&gt;Recent trending activity.&lt;/li&gt;
&lt;li&gt;Locale, language, and category context.&lt;/li&gt;
&lt;li&gt;Optional personalization.&lt;/li&gt;
&lt;li&gt;Safety, policy, and abuse controls.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Never treat raw query-log popularity as automatically trustworthy.&lt;/strong&gt; Bots, spam, inappropriate queries, and old trends can distort the list. Query logs may also contain sensitive user input, so aggregation and retention need appropriate privacy controls.&lt;/p&gt;

&lt;p&gt;For this example, a dedicated suggestion index helps isolate autocomplete latency from expensive full-text searches. It remains a workload-driven decision, not a universally superior architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Preventing Autocomplete Request Storms
&lt;/h2&gt;

&lt;p&gt;Without client-side controls, a user typing &lt;code&gt;iphone&lt;/code&gt; could issue six requests in quick succession. Across many active users, this is a meaningful traffic multiplier.&lt;/p&gt;

&lt;p&gt;Use a combination of:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Debouncing:&lt;/strong&gt; wait a short, tuned interval after typing before requesting suggestions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Minimum prefix length:&lt;/strong&gt; avoid expensive, low-value one-character requests where the product allows it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cancellation or stale-response suppression:&lt;/strong&gt; prevent a slow response for &lt;code&gt;ip&lt;/code&gt; from replacing newer results for &lt;code&gt;iphone&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Server-side rate limits:&lt;/strong&gt; protect shared resources from abusive clients or bursts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Caching hot prefixes:&lt;/strong&gt; popular prefixes often have excellent reuse.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Cancellation is an optimization, not a correctness guarantee: a request may already have reached the server. The client should still ignore responses that do not correspond to its latest input.&lt;/p&gt;

&lt;p&gt;A practical API might be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /suggest?q=iph&amp;amp;locale=en-US&amp;amp;limit=8
GET /search?q=iphone%20case&amp;amp;category=accessories&amp;amp;sort=relevance&amp;amp;limit=20
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service should bound input length, expansion complexity, and returned suggestion count. The cache key must include any locale, market, or other context that changes the response.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Search Ranking: Retrieval Is Not Relevance
&lt;/h2&gt;

&lt;p&gt;An inverted index answers &lt;strong&gt;which documents might match&lt;/strong&gt;. It does not, by itself, determine which matches are most useful.&lt;/p&gt;

&lt;p&gt;A practical ranking pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Query
  |
  v
Candidate retrieval + eligibility filters
  |
  v
First-stage lexical ranking (for example, BM25)
  |
  v
Optional re-ranking of a bounded top-N set
  |
  v
Final eligibility / business rules + top-K results
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;BM25&lt;/strong&gt; is a common lexical baseline. It rewards matches on distinctive terms, considers term frequency with diminishing returns, and normalizes for document length. It is more nuanced than simply counting keyword occurrences.&lt;/p&gt;

&lt;p&gt;An optional second-stage ranker can incorporate popularity, recency, category relevance, or personalization. Running an expensive model on a bounded candidate set controls &lt;strong&gt;re-ranking cost&lt;/strong&gt;, not the cost of retrieving candidates from the index.&lt;/p&gt;

&lt;p&gt;Business rules need care. If an out-of-stock product must never be shown, filtering it &lt;strong&gt;only after selecting the top 10&lt;/strong&gt; could leave too few results even when valid products exist lower down. Apply hard eligibility constraints early enough to preserve the requested result count, or retrieve enough additional candidates to satisfy them.&lt;/p&gt;

&lt;p&gt;When the optional re-ranker fails, the system can fall back to lexical ranking, preserving search availability with reduced relevance quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Filters, Facets, and Typo Tolerance
&lt;/h2&gt;

&lt;p&gt;These features are related but not interchangeable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Filters&lt;/strong&gt; constrain eligibility: &lt;code&gt;brand = Apple&lt;/code&gt;, &lt;code&gt;price &amp;lt;= 500&lt;/code&gt;, or &lt;code&gt;in_stock = true&lt;/code&gt;. They generally should not change textual relevance scores unless that is a deliberate ranking rule.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Facets&lt;/strong&gt; aggregate the matching set: for example, how many eligible products fall into each brand or price range. Facets may be expensive on large result sets or high-cardinality fields. Their semantics must be defined carefully, especially whether a facet's own filter is included when computing its counts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Typo tolerance&lt;/strong&gt; expands potential matches. Common techniques include edit-distance matching, n-grams, and spelling correction. They can improve recall, but broad fuzzy matching on a two-character query may expand into an enormous set of candidate terms.&lt;/p&gt;

&lt;p&gt;Use bounded edit distance, minimum query length, and query-complexity limits. Exact-match identifiers such as SKUs should generally not receive the same fuzzy treatment as natural-language titles.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Keeping the Search Index Up to Date
&lt;/h2&gt;

&lt;p&gt;The source database is authoritative; the search index is a &lt;strong&gt;derived view&lt;/strong&gt;. Product updates need to reach the index reliably.&lt;/p&gt;

&lt;p&gt;A typical pipeline is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Committed DB change
        |
        v
CDC or reliably published event
        |
        v
Durable log / queue
        |
        v
Indexer: validate -&amp;gt; transform -&amp;gt; version-check
        |
        v
Search index (+ suggestion index when applicable)
        |
        v
Refresh / visibility to queries
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CDC is one option, not the only one. Application-emitted events and scheduled imports can also work. If an application writes the database and separately publishes an event, it must address the &lt;strong&gt;dual-write failure gap&lt;/strong&gt; (for example, through a transactional outbox).&lt;/p&gt;

&lt;h3&gt;
  
  
  Duplicate and out-of-order updates
&lt;/h3&gt;

&lt;p&gt;Consider the same product receiving:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Update v11: price = 120
Update v12: price = 100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If v12 is applied first and a delayed v11 is then applied without checking versions, search shows the wrong price. The indexer must reject stale updates based on a trustworthy, monotonically increasing &lt;strong&gt;per-document version&lt;/strong&gt;. A wall-clock &lt;code&gt;updated_at&lt;/code&gt; timestamp alone may not be reliable enough if timestamps collide or clocks are inconsistent.&lt;/p&gt;

&lt;p&gt;Deletes require the same care. A delayed pre-delete update must not resurrect a deleted product. The system needs version-aware tombstone handling or an equivalent mechanism that prevents stale resurrection.&lt;/p&gt;

&lt;h3&gt;
  
  
  Indexing lag is not just queue lag
&lt;/h3&gt;

&lt;p&gt;Even after an indexer successfully writes a document, the search engine may not expose it to normal queries until a refresh. Observed freshness includes both &lt;strong&gt;pipeline delay&lt;/strong&gt; and &lt;strong&gt;index visibility delay&lt;/strong&gt;, plus any applicable result-cache staleness.&lt;/p&gt;

&lt;p&gt;Reducing refresh intervals can improve freshness but increases indexing/segment-management work. A freshness target should therefore be an explicit product requirement.&lt;/p&gt;

&lt;h3&gt;
  
  
  Full reindexing without taking search offline
&lt;/h3&gt;

&lt;p&gt;Changing analyzers, mappings, or transformations can require rebuilding the corpus. A safer approach is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create a new versioned index.&lt;/li&gt;
&lt;li&gt;Bulk-load a consistent source snapshot.&lt;/li&gt;
&lt;li&gt;Apply and catch up on changes made during the rebuild.&lt;/li&gt;
&lt;li&gt;Validate counts, sample documents, deletes, and search quality.&lt;/li&gt;
&lt;li&gt;Switch the serving alias or routing layer to the new index.&lt;/li&gt;
&lt;li&gt;Retain the old index briefly for rollback.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The snapshot and change stream need a well-defined handoff position; otherwise updates made during the rebuild can be missed. Do not destructively rebuild the only live index and assume rollback will be easy.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Caching Search Results Correctly
&lt;/h2&gt;

&lt;p&gt;Popular queries repeat, making caching attractive. But a cache key based only on the raw query string is often &lt;strong&gt;incorrect&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two requests for &lt;code&gt;iphone&lt;/code&gt; can produce different responses depending on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Filters and sort order.&lt;/li&gt;
&lt;li&gt;Locale, region, or tenant.&lt;/li&gt;
&lt;li&gt;Pagination cursor or page.&lt;/li&gt;
&lt;li&gt;Ranking/index version.&lt;/li&gt;
&lt;li&gt;Personalization or user-specific eligibility.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cache only across requests that are genuinely allowed to share the same response. Authorization or tenant boundaries must never be lost through cache-key normalization.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cache key = normalized query + filters + sort + locale
          + page/cursor + index/ranker version
          + relevant access/personalization context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choose TTLs according to query reuse and freshness requirements; there is no universally correct &lt;code&gt;60-second&lt;/code&gt; TTL. For personalized queries with little full-response reuse, caching candidates or reusable components may work better.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cache stampede
&lt;/h3&gt;

&lt;p&gt;Suppose the cache absorbs 80% of incoming search requests. The index normally receives 20%. If the cache disappears and incoming demand stays the same, index-facing traffic can rise by approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 / (1 - 0.80) = 5x
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is &lt;strong&gt;an example derived from an assumed hit rate&lt;/strong&gt;, not a universal multiplier. Request coalescing, TTL jitter, stale-while-revalidate where acceptable, capacity headroom, and load shedding can limit the damage.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Sharding, Replication, and Scatter-Gather
&lt;/h2&gt;

&lt;p&gt;When one search node cannot efficiently hold or serve the corpus, distribute documents across &lt;strong&gt;shards&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hash-based document sharding
&lt;/h3&gt;

&lt;p&gt;A document ID hash can spread documents relatively evenly. A broad keyword query may then need to reach every relevant shard, because matching documents could be anywhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  Routed sharding
&lt;/h3&gt;

&lt;p&gt;Routing by tenant, region, or category can reduce query fan-out &lt;strong&gt;when the query specifies that routing dimension&lt;/strong&gt;. It can also create skew: one very large tenant may produce a hot shard.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scatter-gather query execution
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             Search Coordinator
             /       |        \
          Shard A  Shard B   Shard C
          top-K     top-K     top-K
             \       |        /
              Merge + Global Top-K
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each shard produces its best local candidates; the coordinator merges them. For a straightforward global top-K under a consistent comparable scoring function, retrieving &lt;strong&gt;K from each shard is sufficient&lt;/strong&gt; before global merging. Additional candidates may be required by a subsequent global re-ranking stage, pagination strategy, or other engine-specific behavior; oversampling is not inherently necessary just because results are sharded.&lt;/p&gt;

&lt;p&gt;Distributed search is sensitive to &lt;strong&gt;tail latency&lt;/strong&gt;: a single slow shard can delay the response. Replicas improve availability and read capacity, but do not guarantee uninterrupted or perfectly fresh service. If no replica of a required shard responds, the system must explicitly choose to fail the query or return &lt;strong&gt;clearly marked partial results&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For a product catalog, partial results may be acceptable. For legal discovery or financial reporting, silently omitting a shard's results could be unacceptable.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Pagination: Why Deep OFFSET Becomes Expensive
&lt;/h2&gt;

&lt;p&gt;Simple pagination is easy to explain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- Conceptual syntax; actual search-engine APIs differ.&lt;/span&gt;
&lt;span class="k"&gt;OFFSET&lt;/span&gt; &lt;span class="mi"&gt;10000&lt;/span&gt; &lt;span class="k"&gt;LIMIT&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a distributed search, deep offset pagination may require shards to collect many more results than the 20 shown to the user, increasing work with page depth. Changes between requests can also cause duplicates or omissions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Search-after/cursor pagination&lt;/strong&gt; carries the sort position of the last result, including a deterministic tie-breaker such as document ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"query"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"iphone case"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"limit"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"search_after"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;8.42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"product_991823"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The example assumes a score-descending, ID-tiebroken sort; the actual cursor and comparison rules must match the engine's configured ordering. A cursor avoids repeatedly skipping earlier results, but it does &lt;strong&gt;not&lt;/strong&gt; freeze the underlying result set. If stable multi-page results matter, combine it with point-in-time/snapshot semantics supported by the engine.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. Production Failure Scenarios
&lt;/h2&gt;

&lt;p&gt;A design is incomplete if it only describes the happy path. These are the failures most likely to change our architectural choices.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure&lt;/th&gt;
&lt;th&gt;User-visible impact&lt;/th&gt;
&lt;th&gt;Response&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Search shard unavailable&lt;/td&gt;
&lt;td&gt;Errors, higher latency, or incomplete results&lt;/td&gt;
&lt;td&gt;Route to healthy replica; otherwise follow explicit fail/partial policy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache outage&lt;/td&gt;
&lt;td&gt;Search cluster receives sudden extra traffic&lt;/td&gt;
&lt;td&gt;Request coalescing, headroom, load shedding&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Indexer backlog&lt;/td&gt;
&lt;td&gt;Newly changed products remain stale in search&lt;/td&gt;
&lt;td&gt;Monitor oldest event age, retain backlog, scale catch-up capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lost CDC events / retention gap&lt;/td&gt;
&lt;td&gt;Index may permanently diverge&lt;/td&gt;
&lt;td&gt;Reconcile against source; rebuild when replay is impossible&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bad analyzer/index deployment&lt;/td&gt;
&lt;td&gt;Wrong matches or degraded relevance despite healthy servers&lt;/td&gt;
&lt;td&gt;Versioned indexes, validation, alias rollback&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Old update arrives after delete&lt;/td&gt;
&lt;td&gt;Deleted product reappears&lt;/td&gt;
&lt;td&gt;Source-version checks and durable tombstone semantics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hot prefix or expensive fuzzy query&lt;/td&gt;
&lt;td&gt;CPU spikes, increased tail latency&lt;/td&gt;
&lt;td&gt;Cache hot prefixes; bound expansions, deadlines, and query complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ranking dependency fails&lt;/td&gt;
&lt;td&gt;Search relevance degrades&lt;/td&gt;
&lt;td&gt;Fall back to first-stage lexical ranking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autocomplete index lags&lt;/td&gt;
&lt;td&gt;Search finds a product that suggestions omit&lt;/td&gt;
&lt;td&gt;Separate freshness monitoring and defined suggestion-update cadence&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Monitor more than availability
&lt;/h3&gt;

&lt;p&gt;A search service can return HTTP 200 while its results are wrong. Useful measurements include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Search and autocomplete P50/P95/P99 latency, error rate, and QPS.&lt;/li&gt;
&lt;li&gt;Per-shard latency, timeouts, load, and replica health.&lt;/li&gt;
&lt;li&gt;Cache hit rate, hot-key pressure, and upstream request coalescing.&lt;/li&gt;
&lt;li&gt;Indexing backlog, &lt;strong&gt;oldest unprocessed event age&lt;/strong&gt;, and end-to-end freshness.&lt;/li&gt;
&lt;li&gt;Source/index divergence checks, delete correctness, and failed indexing events.&lt;/li&gt;
&lt;li&gt;Zero-result rate, relevance regression tests, and other quality signals interpreted in context.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For more failure walkthroughs, see &lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/failure-scale-scenarios" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Failure &amp;amp; Scale Scenarios&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  14. Major Architecture Trade-offs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision&lt;/th&gt;
&lt;th&gt;Simpler option&lt;/th&gt;
&lt;th&gt;When added complexity may be justified&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Search engine&lt;/td&gt;
&lt;td&gt;Database-native full-text search&lt;/td&gt;
&lt;td&gt;High QPS, sophisticated ranking/facets, independent scaling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Autocomplete&lt;/td&gt;
&lt;td&gt;Search-engine prefix queries&lt;/td&gt;
&lt;td&gt;Dedicated suggestion index for latency isolation or specialized ranking&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefix representation&lt;/td&gt;
&lt;td&gt;Edge n-grams&lt;/td&gt;
&lt;td&gt;Trie/FST or precomputed suggestions for specific memory/update/latency goals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Change propagation&lt;/td&gt;
&lt;td&gt;Batch import or polling&lt;/td&gt;
&lt;td&gt;CDC or reliable domain events for tighter freshness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ranking&lt;/td&gt;
&lt;td&gt;Lexical relevance only&lt;/td&gt;
&lt;td&gt;Re-rank bounded candidates for richer relevance signals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache&lt;/td&gt;
&lt;td&gt;Cache common final responses&lt;/td&gt;
&lt;td&gt;Cache candidates/components when full responses have low reuse&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sharding&lt;/td&gt;
&lt;td&gt;Hash by document ID&lt;/td&gt;
&lt;td&gt;Routing by tenant/domain when queries can exploit it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pagination&lt;/td&gt;
&lt;td&gt;Shallow offset&lt;/td&gt;
&lt;td&gt;Search-after for deep pages; snapshot for stable multi-page views&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shard failure&lt;/td&gt;
&lt;td&gt;Fail the query&lt;/td&gt;
&lt;td&gt;Marked partial results when the product tolerates incompleteness&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The correct design is the &lt;strong&gt;simplest architecture that meets the stated requirements&lt;/strong&gt;. A trie, CDC pipeline, separate ranking service, and multiple caching layers should each have a reason to exist. Adding all of them by default is not a sign of a stronger design.&lt;/p&gt;

&lt;p&gt;See the full &lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/tradeoffs-decisions" rel="noopener noreferrer"&gt;Trade-offs &amp;amp; Decisions&lt;/a&gt; discussion for the alternatives behind these choices.&lt;/p&gt;

&lt;h2&gt;
  
  
  15. Search &amp;amp; Autocomplete System Design Interview Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why not use the primary database for search?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For small workloads, we may. At higher scale, a dedicated index provides search-specific structures and independent scaling. Avoid claiming SQL databases cannot perform full-text search.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What's the difference between an inverted index and a trie?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An inverted index maps terms to matching documents. A trie organizes prefixes for efficient completion lookup. They answer different questions, and autocomplete does not necessarily require a trie.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you handle typos?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Use bounded fuzzy matching, n-grams, or spelling correction where appropriate. Keep exact identifiers exact and prevent broad short-prefix expansions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How does a new product become searchable?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A committed source change reaches an indexer through CDC or reliable events, is applied with version checks, and becomes visible after the search engine's refresh. The process is usually eventually consistent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens if indexing stops for 30 minutes?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Search can continue serving stale results. Recovery requires the missing events to remain available and the indexer to catch up faster than new events arrive. If event history is lost, reconciliation or reindexing is necessary.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why use BM25 and then a re-ranker?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;BM25 gives an efficient lexical baseline. A bounded re-ranking stage can use richer signals without scoring the entire corpus with an expensive model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you prevent autocomplete from overwhelming the backend?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Debounce, ignore stale responses, enforce minimum prefix lengths where suitable, cache common prefixes, and apply server-side rate and complexity limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when a shard is down?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Try a healthy replica. If no copy is available, either fail the query or return results explicitly marked as partial, depending on whether incompleteness is acceptable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you change analyzers without downtime?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Build a new index from a snapshot, replay intervening changes, validate, switch the serving alias, and retain the old version for rollback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why not use OFFSET for page 500?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Deep offset work grows with depth and results can shift between requests. Search-after avoids repeated skipping; point-in-time semantics can provide a stable view when required.&lt;/p&gt;

&lt;p&gt;For a broader question bank, see &lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/interview-perspective" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Interview Perspective&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Takeaways
&lt;/h2&gt;

&lt;p&gt;A reliable search system separates &lt;strong&gt;the source of truth from the search-optimized view&lt;/strong&gt;, and separates &lt;strong&gt;candidate retrieval from ranking&lt;/strong&gt;. Autocomplete introduces another latency-sensitive workload whose candidate generation, ranking, and request controls deserve explicit design decisions.&lt;/p&gt;

&lt;p&gt;The most important production details are easy to overlook: compatible text analyzers, version-aware indexing, safe delete handling, refresh lag, correct cache keys, bounded expensive queries, shard failure semantics, and rebuilds that do not destroy the live index.&lt;/p&gt;

&lt;p&gt;None of these mechanisms is a universal prescription. Start with the actual search experience, traffic, freshness, and correctness requirements; introduce complexity only where those requirements justify it.&lt;/p&gt;

&lt;h2&gt;
  
  
  References and Further Reading
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/complete-design" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Complete Design&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/architecture-deep-dive" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Architecture Deep Dive&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/tradeoffs-decisions" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Trade-offs &amp;amp; Decisions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/failure-scale-scenarios" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Failure &amp;amp; Scale Scenarios&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/system-design/foundations/search-autocomplete/interview-perspective" rel="noopener noreferrer"&gt;Search &amp;amp; Autocomplete — Interview Perspective&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://seeitflow.com/" rel="noopener noreferrer"&gt;SeeItFlow — Visual Software Engineering Learning&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>systemdesign</category>
      <category>architecture</category>
      <category>backend</category>
      <category>search</category>
    </item>
    <item>
      <title>Designing a Notification System: Queues, Retries, Rate Limiting &amp; Scale</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Thu, 08 Oct 2026 10:44:37 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/designing-a-notification-system-queues-retries-rate-limiting-scale-4bg1</link>
      <guid>https://dev.to/mangeshmandlik/designing-a-notification-system-queues-retries-rate-limiting-scale-4bg1</guid>
      <description>&lt;p&gt;&lt;em&gt;The difficult part isn't sending an email. It's making sure a slow provider, a viral campaign, or a retry storm cannot take down the services that need to send one.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffexlso17cnv15h2dw7oh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffexlso17cnv15h2dw7oh.png" alt="complete architecture" width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Imagine an e-commerce platform just completed an order. The order service needs to send a confirmation email, the payment service needs to send a receipt, and a security service may need to send an urgent login alert.&lt;/p&gt;

&lt;p&gt;Now imagine the email provider starts taking eight seconds to respond.&lt;/p&gt;

&lt;p&gt;If those services call the provider synchronously, an external notification dependency suddenly becomes part of the checkout and login latency budget. If a promotional campaign simultaneously produces millions of messages, critical security alerts can end up behind a mountain of marketing traffic.&lt;/p&gt;

&lt;p&gt;A production notification system has to solve a more interesting problem than sending messages: &lt;strong&gt;accept notification intent durably, then deliver it independently at the rate each downstream provider can sustain.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Let's build that architecture from first principles, and examine where seemingly sensible implementations fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Requirements: What Are We Designing?
&lt;/h2&gt;

&lt;p&gt;Our service accepts requests from many producer applications and delivers notifications through &lt;strong&gt;push, email, and SMS&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It should support:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Transactional messages such as order confirmations and payment updates.&lt;/li&gt;
&lt;li&gt;Time-sensitive messages such as security alerts and one-time passcodes.&lt;/li&gt;
&lt;li&gt;Promotional campaigns, potentially targeting millions of recipients.&lt;/li&gt;
&lt;li&gt;Recipient preferences, opt-outs, quiet hours, and valid destination checks.&lt;/li&gt;
&lt;li&gt;Delivery status, retries, and investigation of failed messages.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important non-functional requirements are &lt;strong&gt;fast producer acknowledgment, durable processing, channel isolation, preferential latency for urgent traffic, and observability&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;One distinction defines our API contract:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;ACCEPTED does not mean DELIVERED.&lt;/strong&gt; It means the system has durably recorded the notification for asynchronous processing. Delivery may later succeed, fail permanently, expire, or be suppressed by policy.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That contract is the foundation of the &lt;a href="https://seeitflow.com/system-design/foundations/notification/complete-design" rel="noopener noreferrer"&gt;Notification System Complete Design&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Capacity Estimation: Ingestion and Delivery Are Different Rates
&lt;/h2&gt;

&lt;p&gt;Consider an illustrative traffic burst:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Example workload&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Incoming notifications&lt;/td&gt;
&lt;td&gt;5,000/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider's sustainable send capacity&lt;/td&gt;
&lt;td&gt;800/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Excess arriving during the burst&lt;/td&gt;
&lt;td&gt;4,200/second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Backlog after a 60-second burst&lt;/td&gt;
&lt;td&gt;252,000 notifications&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If incoming traffic drops to zero and the provider continues delivering at 800/second, that backlog alone takes &lt;strong&gt;315 seconds&lt;/strong&gt;, or about &lt;strong&gt;5 minutes 15 seconds&lt;/strong&gt;, to drain.&lt;/p&gt;

&lt;p&gt;In reality, new traffic may continue arriving, reducing the spare capacity available to clear old messages.&lt;/p&gt;

&lt;p&gt;This leads to a rule that is easy to overlook:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A queue absorbs a temporary difference between arrival and delivery rates. It does not manufacture provider capacity.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If 5,000 messages arrive every second indefinitely and the provider can only deliver 800, the backlog grows indefinitely. We eventually need more provider capacity, slower campaign ingestion, traffic shedding, or an explicitly longer delivery window for less urgent notifications.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Start With the Simplest Architecture — and Watch It Break
&lt;/h2&gt;

&lt;p&gt;The first implementation might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Order Service
     |
     v
Notification API
     |
     v
Email / SMS / Push Provider
     |
     v
Response to Order Service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp5y89ee2svljq48e7ys4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp5y89ee2svljq48e7ys4.png" width="800" height="281"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It's simple, but the calling service is now waiting for a provider it does not control.&lt;/p&gt;

&lt;p&gt;If the provider slows to an eight-second p99 response time, the producer's request may inherit that delay. If the provider is unavailable, notifications may fail alongside otherwise healthy business operations.&lt;/p&gt;

&lt;p&gt;The first meaningful architectural change is not adding ten microservices. It's &lt;strong&gt;putting a durable asynchronous boundary between notification creation and delivery&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Producer
   |
   v
Notification API
   |
   v
Durable Queue  -----&amp;gt;  Delivery Workers  -----&amp;gt;  Provider
   ^                         |
   |                         v
Fast acceptance         Retry / failure handling
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz156fwmisaekuk2hy9wd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz156fwmisaekuk2hy9wd.png" width="800" height="278"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now the producer does not wait for the provider. Workers can process messages at a controlled pace, and a temporary provider outage becomes a backlog rather than an immediate failure of every calling application.&lt;/p&gt;

&lt;p&gt;But the phrase “durable queue” hides a subtle consistency problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Durable Ingestion: The Database–Queue Dual-Write Problem
&lt;/h2&gt;

&lt;p&gt;A typical notification API must record the notification's state in a database and publish work to a message broker.&lt;/p&gt;

&lt;p&gt;Suppose we implement:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Insert notification into the database.&lt;/li&gt;
&lt;li&gt;Publish the notification to Kafka, SQS, or RabbitMQ.&lt;/li&gt;
&lt;li&gt;Return success.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What if step 1 succeeds but step 2 fails?&lt;/p&gt;

&lt;p&gt;The database says the notification exists, but no worker ever receives it.&lt;/p&gt;

&lt;p&gt;Reversing the order doesn't solve it. The broker may accept the message while the database write fails, leaving workers processing a notification with no durable tracking record.&lt;/p&gt;

&lt;h3&gt;
  
  
  Transactional outbox
&lt;/h3&gt;

&lt;p&gt;A practical solution is to write &lt;strong&gt;the notification record and an outbox event in the same database transaction&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Notification API
       |
       v
Single DB transaction
  +------------------------+
  | Notification record    |
  | Outbox event           |
  +------------------------+
       |
       v
Outbox Publisher  ---&amp;gt;  Message Broker  ---&amp;gt;  Workers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API can acknowledge once that durable transaction commits, while the outbox publisher continues retrying publication if the broker is temporarily unavailable.&lt;/p&gt;

&lt;p&gt;This closes the dual-write gap, but &lt;strong&gt;does not guarantee exactly-once processing&lt;/strong&gt;. A publisher can send an event successfully and crash before marking it published, causing it to publish again after restart. Consumers still need idempotent handling.&lt;/p&gt;

&lt;p&gt;The exact failure windows and guarantees are covered in &lt;a href="https://seeitflow.com/system-design/foundations/notification/architecture-deep-dive" rel="noopener noreferrer"&gt;Notification System Architecture Deep Dive — Durable Ingestion&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. API Design: Accept Intent, Not Delivery Promises
&lt;/h2&gt;

&lt;p&gt;A minimal API might expose:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /notifications
Content-Type: application/json
Idempotency-Key: order-8421-confirmation
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recipient_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user-123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"channel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"priority"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"normal"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"template_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"order-confirmed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"template_variables"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"order_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"8421"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful response returns a notification ID and &lt;code&gt;ACCEPTED&lt;/code&gt; status — not &lt;code&gt;DELIVERED&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"notification_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"notif-9271"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ACCEPTED"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Other useful endpoints include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;GET /notifications/{id}/status&lt;/code&gt; — inspect the lifecycle state.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;POST /notifications/batch&lt;/code&gt; — accept a campaign definition and return a batch ID without synchronously expanding millions of recipients.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The producer-supplied idempotency key matters. If the producer times out after the API successfully commits, it may retry. The same key should return the existing logical notification instead of creating another one.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. High-Level Architecture: Isolate Channels and Priorities
&lt;/h2&gt;

&lt;p&gt;A single queue with one pool of workers looks attractive until the email provider fails while push notifications are healthy.&lt;/p&gt;

&lt;p&gt;We want the channels to fail and scale independently.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                       Producer Services
                              |
                              v
                       Notification API
                              |
                        Durable Ingestion
                              |
                              v
                   Policy / Channel Router
                              |
                  +-----------+-----------+
                  |           |           |
               Push Queue  Email Queue  SMS Queue
                  |           |           |
               Workers     Workers     Workers
                  |           |           |
             Rate Limiter Rate Limiter Rate Limiter
                  |           |           |
                FCM/APNs   Email Provider SMS Provider
                  \           |           /
                   \          |          /
                    Delivery State &amp;amp; Reconciliation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv85mk1jgesjq9rxssoek.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv85mk1jgesjq9rxssoek.png" width="799" height="319"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each channel can have its own backlog, retry policy, worker concurrency, provider adapter, and rate-limit budget.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffexlso17cnv15h2dw7oh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffexlso17cnv15h2dw7oh.png" alt="complete architecture" width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The architecture also needs &lt;strong&gt;priority-aware scheduling&lt;/strong&gt; within or across channel queues. We'll address that next.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Priority Queues: Security Alerts Must Not Wait Behind Campaigns
&lt;/h2&gt;

&lt;p&gt;Imagine a promotional campaign enqueues 10 million notifications. A security alert arrives one second later.&lt;/p&gt;

&lt;p&gt;If everything shares one FIFO queue, that alert could wait behind enormous amounts of lower-value traffic.&lt;/p&gt;

&lt;p&gt;Separate priority classes help:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Priority&lt;/th&gt;
&lt;th&gt;Typical traffic&lt;/th&gt;
&lt;th&gt;Desired behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Security alerts, payment failures&lt;/td&gt;
&lt;td&gt;Preferential low latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Normal&lt;/td&gt;
&lt;td&gt;Order confirmations, shipping updates&lt;/td&gt;
&lt;td&gt;Predictable progress&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Promotions, newsletters&lt;/td&gt;
&lt;td&gt;Can tolerate delay&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;But a naive rule — &lt;em&gt;always drain high priority before touching normal or low&lt;/em&gt; — creates &lt;strong&gt;starvation&lt;/strong&gt;. If high-priority traffic never fully stops, lower-priority notifications may never be served.&lt;/p&gt;

&lt;p&gt;Production alternatives include &lt;strong&gt;weighted polling&lt;/strong&gt;, &lt;strong&gt;reserved capacity&lt;/strong&gt;, and &lt;strong&gt;aging&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example, a scheduler might use an illustrative ratio of eight high-priority messages, two normal, and one low per round. The exact weights are workload decisions, not universal defaults.&lt;/p&gt;

&lt;p&gt;The objective is to protect urgent traffic &lt;strong&gt;while guaranteeing that every eligible class makes progress&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;See &lt;a href="https://seeitflow.com/system-design/foundations/notification/tradeoffs-decisions" rel="noopener noreferrer"&gt;Trade-offs &amp;amp; Decisions — Strict vs Fair Priority&lt;/a&gt; for the alternatives and their costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Recipient Preferences, Push Tokens, and Templates
&lt;/h2&gt;

&lt;p&gt;Before a notification reaches a provider, the router must decide whether it is &lt;strong&gt;eligible&lt;/strong&gt; to send.&lt;/p&gt;

&lt;p&gt;That means checking the recipient's category preferences, channel permissions, opt-out status, quiet hours, and whether the destination is valid.&lt;/p&gt;

&lt;p&gt;A promotional email should not be sent simply because it was queued before the user unsubscribed. For delayed messages, cached preferences may need to be checked again close to delivery time.&lt;/p&gt;

&lt;p&gt;Push adds another complication: &lt;strong&gt;users have devices, and devices have tokens&lt;/strong&gt;. One user may have several active tokens. A provider-reported invalid token should deactivate that specific destination, not trigger endless retries or disable the user's other devices.&lt;/p&gt;

&lt;p&gt;Templates also need versioning. If an order confirmation sits in a queue for an hour and the template changes, should the message use the original or the new content? Referencing a specific template version at creation time avoids silently changing in-flight notifications.&lt;/p&gt;

&lt;p&gt;These concerns are explored separately in the &lt;a href="https://seeitflow.com/system-design/foundations/notification/architecture-deep-dive" rel="noopener noreferrer"&gt;Architecture Deep Dive&lt;/a&gt;, including recipient eligibility, token lifecycle, localization, and render-before-queue versus render-at-delivery trade-offs.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Rate Limiting and Backpressure: More Workers Won't Fix a Provider Quota
&lt;/h2&gt;

&lt;p&gt;Suppose workers are capable of making 5,000 provider requests per second, but the provider account allows only 1,000.&lt;/p&gt;

&lt;p&gt;Adding workers does not raise that account limit. It may simply generate more &lt;code&gt;429 Too Many Requests&lt;/code&gt; responses.&lt;/p&gt;

&lt;p&gt;The system needs &lt;strong&gt;provider-level rate limiting shared across distributed workers&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Two common designs are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Centralized token bucket:&lt;/strong&gt; All workers coordinate through a shared rate limiter. This provides a global quota view but adds network calls and contention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leased quota:&lt;/strong&gt; Workers receive bounded slices of the total quota and spend them locally. This reduces coordination overhead but can leave unused quota stranded with idle workers.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When a provider returns &lt;code&gt;429&lt;/code&gt; or &lt;code&gt;Retry-After&lt;/code&gt;, the limiter should adjust rather than repeatedly sending at the same rejected rate.&lt;/p&gt;

&lt;p&gt;The queue provides &lt;strong&gt;backpressure absorption&lt;/strong&gt; while workers respect the provider's sustainable throughput. It does not remove the need for admission control when traffic stays above capacity.&lt;/p&gt;

&lt;p&gt;For a detailed comparison, see &lt;a href="https://seeitflow.com/system-design/foundations/notification/tradeoffs-decisions" rel="noopener noreferrer"&gt;Notification System Trade-offs &amp;amp; Decisions&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Retries: Why Exponential Backoff Needs Jitter and Deadlines
&lt;/h2&gt;

&lt;p&gt;A provider outage often triggers a second incident caused by the notification system itself.&lt;/p&gt;

&lt;p&gt;Picture thousands of workers receiving failures at the same moment. If they all retry one second later, then two seconds later, then four seconds later, the recovering provider gets hit by synchronized waves of traffic.&lt;/p&gt;

&lt;p&gt;This is a &lt;strong&gt;retry storm&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Exponential backoff reduces the frequency of retries. &lt;strong&gt;Jitter&lt;/strong&gt; randomizes retry timing so workers don't all return simultaneously.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Without jitter:
  Workers fail together
       | 1s
       +----&amp;gt; retry spike
       | 2s
       +----&amp;gt; retry spike

With jitter:
  Workers fail together
       |  randomized delays
       +--&amp;gt; --&amp;gt; ---&amp;gt; ----&amp;gt; spread-out retries
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Not every failure should retry:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure&lt;/th&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Temporary timeout or provider 5xx&lt;/td&gt;
&lt;td&gt;Retry with backoff and jitter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Provider 429&lt;/td&gt;
&lt;td&gt;Honor retry guidance; lower send rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Invalid push token&lt;/td&gt;
&lt;td&gt;Deactivate token; do not retry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Malformed destination&lt;/td&gt;
&lt;td&gt;Permanent failure or appropriate terminal handling&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum retries exhausted&lt;/td&gt;
&lt;td&gt;Move to DLQ when appropriate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Notification deadline passed&lt;/td&gt;
&lt;td&gt;Expire instead of delivering stale content&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Retries should be &lt;strong&gt;scheduled&lt;/strong&gt; using delayed queues or retry topics, not by keeping worker threads asleep.&lt;/p&gt;

&lt;p&gt;And every time-sensitive notification needs a deadline. A one-time passcode delivered six hours late is not a successful user experience, even if the provider eventually accepts it.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Dead-Letter Queues: Failure Must Be Visible and Recoverable
&lt;/h2&gt;

&lt;p&gt;After a bounded number of attempts, unresolved failures can be moved to a &lt;strong&gt;dead-letter queue (DLQ)&lt;/strong&gt; for diagnosis.&lt;/p&gt;

&lt;p&gt;The DLQ is not a trash can. It is an operational tool that preserves context such as the notification ID, error, attempt count, and relevant timestamps.&lt;/p&gt;

&lt;p&gt;A growing DLQ should be monitored by &lt;strong&gt;channel, provider, error category, and message age&lt;/strong&gt;, not only by one global threshold.&lt;/p&gt;

&lt;p&gt;Replaying messages requires care:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Fix the underlying failure first.&lt;/li&gt;
&lt;li&gt;Re-check expiry and recipient preferences.&lt;/li&gt;
&lt;li&gt;Replay at a controlled pace.&lt;/li&gt;
&lt;li&gt;Retain prior attempt history and observe whether the problem recurs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Dumping the entire DLQ back into the main queue can recreate the original overload.&lt;/p&gt;

&lt;p&gt;See &lt;a href="https://seeitflow.com/system-design/foundations/notification/failure-scale-scenarios" rel="noopener noreferrer"&gt;Failure &amp;amp; Scale Scenarios — DLQ Growth and Backlog Recovery&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Idempotency: Can We Guarantee Exactly-Once Notifications?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;No — not universally across an external notification provider.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There are three different boundaries to protect:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Producer → API:&lt;/strong&gt; A retried API call should not create a second logical notification. Use a producer idempotency key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Queue → Worker:&lt;/strong&gt; Brokers may redeliver messages. Make worker state transitions idempotent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Worker → Provider:&lt;/strong&gt; The external side effect can become ambiguous after a timeout or crash.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Consider this sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Worker sends notification to provider
                  |
                  v
Provider accepts request
                  |
                  v
Worker crashes before recording success
                  |
                  v
Broker redelivers the message
                  |
                  v
New worker cannot tell whether it was sent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the new worker retries, the user might receive the notification twice. If it doesn't, the notification might never arrive.&lt;/p&gt;

&lt;p&gt;A pre-send deduplication check cannot resolve this ambiguity because the successful external action was never recorded.&lt;/p&gt;

&lt;p&gt;Where supported, &lt;strong&gt;provider-side idempotency keys&lt;/strong&gt; reduce the risk. Persisting provider message IDs and reconciling against callbacks or status APIs also helps.&lt;/p&gt;

&lt;p&gt;The honest contract is &lt;strong&gt;durable at-least-once processing with strong duplicate reduction&lt;/strong&gt;, not an unsupported promise of exactly-once end-user delivery.&lt;/p&gt;

&lt;p&gt;The full crash-window analysis is in &lt;a href="https://seeitflow.com/system-design/foundations/notification/architecture-deep-dive" rel="noopener noreferrer"&gt;Notification System Architecture Deep Dive — Idempotency &amp;amp; Duplicate Suppression&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. Provider Failover: A Timeout Is Not a Rejection
&lt;/h2&gt;

&lt;p&gt;Suppose the primary SMS provider times out. Should we immediately send through a backup provider?&lt;/p&gt;

&lt;p&gt;Not necessarily.&lt;/p&gt;

&lt;p&gt;The primary may have already accepted the request but lost the response. Sending through a second provider can produce two real messages.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;known rejection&lt;/strong&gt; is different from an &lt;strong&gt;ambiguous outcome&lt;/strong&gt;. Failover policy should consider the signal, message urgency, duplicate tolerance, and the backup provider's capacity.&lt;/p&gt;

&lt;p&gt;For a security alert, accepting some duplicate risk might be reasonable to improve the chance of timely delivery. For a promotional notification, conservative retry or reconciliation may be preferable.&lt;/p&gt;

&lt;p&gt;Provider failover is a trade-off, not a universal on/off switch. The &lt;a href="https://seeitflow.com/system-design/foundations/notification/tradeoffs-decisions" rel="noopener noreferrer"&gt;Trade-offs &amp;amp; Decisions&lt;/a&gt; page compares failover versus retrying the primary in detail.&lt;/p&gt;

&lt;h2&gt;
  
  
  14. Delivery Tracking: SENT Is Not DELIVERED
&lt;/h2&gt;

&lt;p&gt;The system should maintain an observable notification lifecycle:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ACCEPTED → QUEUED → PROCESSING → PROVIDER_ACCEPTED
                                     |
                                     v
                                  DELIVERED

Alternative outcomes:
  RETRY_SCHEDULED | FAILED_PERMANENT | EXPIRED
  SUPPRESSED | BOUNCED | UNDELIVERABLE | DLQ
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;PROVIDER_ACCEPTED&lt;/code&gt; means the provider took the request. It does &lt;strong&gt;not&lt;/strong&gt; prove that the recipient received it.&lt;/p&gt;

&lt;p&gt;Providers may send delivery-status webhooks later, but those callbacks can arrive duplicated, delayed, out of order, or not at all.&lt;/p&gt;

&lt;p&gt;A robust reconciliation pipeline should:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authenticate provider callbacks.&lt;/li&gt;
&lt;li&gt;Deduplicate callback events.&lt;/li&gt;
&lt;li&gt;Apply idempotent, monotonic state transitions.&lt;/li&gt;
&lt;li&gt;Prevent a late &lt;code&gt;SENT&lt;/code&gt; callback from overwriting &lt;code&gt;DELIVERED&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Query provider status APIs when callbacks are missing and such APIs are available.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not every channel provides a reliable end-device delivery receipt. The status model must describe what the system actually knows, rather than claim more than a provider can prove.&lt;/p&gt;

&lt;h2&gt;
  
  
  15. Production Failure Scenarios: Where the Architecture Gets Tested
&lt;/h2&gt;

&lt;p&gt;A notification system's quality becomes clear when something breaks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Email provider outage
&lt;/h3&gt;

&lt;p&gt;Email workers back off and accumulate a durable backlog. Push and SMS continue independently. Recovery is paced according to provider capacity and message deadlines.&lt;/p&gt;

&lt;h3&gt;
  
  
  Broker outage
&lt;/h3&gt;

&lt;p&gt;With a transactional outbox, the API can continue durably recording notification intent while the broker is down, subject to database capacity and the API's acceptance contract. Publishing resumes when the broker recovers. Without a durable fallback, the API should not falsely report successful acceptance.&lt;/p&gt;

&lt;h3&gt;
  
  
  Massive promotional campaign
&lt;/h3&gt;

&lt;p&gt;A campaign targeting millions of recipients should be expanded into &lt;strong&gt;bounded fanout chunks&lt;/strong&gt;, paced to sustainable capacity. Transactional and security traffic must retain protected capacity. The campaign should support cancellation and expiry.&lt;/p&gt;

&lt;h3&gt;
  
  
  Priority starvation
&lt;/h3&gt;

&lt;p&gt;A strict high-first scheduler can leave normal and low-priority notifications waiting forever. Monitor &lt;strong&gt;oldest-message age by priority&lt;/strong&gt;, not just queue depth, and use fair scheduling or reserved capacity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Delivery tracker outage
&lt;/h3&gt;

&lt;p&gt;A well-isolated tracker failure should degrade status visibility rather than stop all delivery. But provider IDs and state events must be preserved for later reconciliation; simply losing them creates a permanent audit gap.&lt;/p&gt;

&lt;h3&gt;
  
  
  Backlog recovery storm
&lt;/h3&gt;

&lt;p&gt;After a long outage, releasing every queued notification at once can overwhelm a recovering provider. Drain gradually, respect priority, and discard or expire messages that are no longer useful.&lt;/p&gt;

&lt;p&gt;These are part of the &lt;strong&gt;14 concrete incidents&lt;/strong&gt; covered in &lt;a href="https://seeitflow.com/system-design/foundations/notification/failure-scale-scenarios" rel="noopener noreferrer"&gt;Notification System Failure &amp;amp; Scale Scenarios&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  16. How Would We Evolve This Architecture?
&lt;/h2&gt;

&lt;p&gt;A good system design grows in response to evidence, not because a diagram looks more impressive with more components.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Architecture&lt;/th&gt;
&lt;th&gt;Problem it solves&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Simple notification service&lt;/td&gt;
&lt;td&gt;Basic low-volume delivery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Durable ingestion and asynchronous workers&lt;/td&gt;
&lt;td&gt;Provider latency and failure isolation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Per-channel queues, preferences, priority scheduling&lt;/td&gt;
&lt;td&gt;Multi-channel correctness and urgent-message latency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Shared provider rate limits, retry scheduling, DLQ&lt;/td&gt;
&lt;td&gt;Throttling, transient failures, and observability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Idempotency, reconciliation, campaign pacing&lt;/td&gt;
&lt;td&gt;Duplicate reduction, trustworthy state, and large-scale fanout&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The precise order depends on product requirements. For example, recipient opt-outs and basic idempotency may be mandatory from the beginning rather than later enhancements.&lt;/p&gt;

&lt;h2&gt;
  
  
  17. Notification System Design Interview Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Why use a queue in a notification system?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
To decouple the rate of notification creation from the rate of provider delivery, keep producer latency independent of provider health, and buffer temporary bursts durably.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can Kafka or SQS guarantee exactly-once notification delivery?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. Internal broker semantics cannot eliminate the ambiguous external provider-call window. Use at-least-once processing, idempotent internal transitions, provider idempotency where available, and reconciliation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do security alerts bypass promotional notifications?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Use priority classes with protected capacity or fair scheduling, so urgent traffic gets preferential latency without starving other classes indefinitely.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens if the provider allows only 1,000 messages per second?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Coordinate a rate-limit budget across workers. Additional workers cannot bypass the provider's account-level ceiling. Pace incoming campaigns and monitor backlog age.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why do retries need jitter?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Exponential backoff alone can leave workers synchronized. Jitter spreads attempts across time and reduces retry spikes during recovery.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should we fail over to a second provider after a timeout?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Only after considering whether the outcome is ambiguous, how urgent the notification is, and whether duplicate delivery is acceptable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do we send 100 million promotional notifications without delaying OTPs?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Accept the campaign quickly, enumerate recipients asynchronously in bounded chunks, pace fanout, and reserve downstream capacity for time-sensitive traffic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between SENT and DELIVERED?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
SENT or PROVIDER_ACCEPTED means the external provider accepted the request; DELIVERED requires a stronger provider or downstream confirmation, where supported.&lt;/p&gt;

&lt;p&gt;For more advanced follow-ups on outbox consistency, webhook ordering, push-token invalidation, and dedup-store failures, see the &lt;a href="https://seeitflow.com/system-design/foundations/notification/interview-perspective" rel="noopener noreferrer"&gt;Notification System Interview Perspective&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;The hardest part of a notification system isn't choosing Kafka over RabbitMQ or integrating an SMS SDK.&lt;/p&gt;

&lt;p&gt;It's recognizing that &lt;strong&gt;notification creation, delivery, provider acceptance, and confirmed receipt are different events with different guarantees&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A durable queue protects producers from provider slowness. Channel isolation contains failures. Fair scheduling protects critical traffic. Rate limiting respects real provider ceilings. Backoff, jitter, deadlines, and DLQs make recovery controlled rather than chaotic. Idempotency and reconciliation reduce duplicate risk without pretending external side effects can be made universally exactly-once.&lt;/p&gt;

&lt;p&gt;The architecture becomes easier to reason about when each component exists to solve a specific failure mode.&lt;/p&gt;

&lt;p&gt;That's the deeper lesson of notification system design: &lt;strong&gt;reliability comes from making failure boundaries explicit, not from adding more infrastructure.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  References &amp;amp; Further Reading
&lt;/h2&gt;

&lt;p&gt;These SeeItFlow resources explore the same system in more detail:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/notification/complete-design" rel="noopener noreferrer"&gt;Notification System — Complete Design&lt;/a&gt;&lt;/strong&gt; — Requirements, API contract, queue-based evolution, channel routing, priorities, retries, DLQ, and final architecture.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/notification/architecture-deep-dive" rel="noopener noreferrer"&gt;Notification System — Architecture Deep Dive&lt;/a&gt;&lt;/strong&gt; — Transactional outbox, three idempotency boundaries, fair scheduling, recipient policy, push tokens, templates, distributed rate limiting, provider adapters, webhooks, campaign fanout, and delivery-state storage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/notification/tradeoffs-decisions" rel="noopener noreferrer"&gt;Notification System — Trade-offs &amp;amp; Decisions&lt;/a&gt;&lt;/strong&gt; — Sync versus async, queue topology, retry policies, rate-limit coordination, provider failover, and template-rendering decisions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/notification/failure-scale-scenarios" rel="noopener noreferrer"&gt;Notification System — Failure &amp;amp; Scale Scenarios&lt;/a&gt;&lt;/strong&gt; — Fourteen operational scenarios covering provider outages, retry storms, queue failures, throttling, duplicate events, starvation, and backlog recovery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/notification/interview-perspective" rel="noopener noreferrer"&gt;Notification System — Interview Perspective&lt;/a&gt;&lt;/strong&gt; — Interview answer structure, common mistakes, and advanced follow-up questions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Explore more free, visual explanations of distributed systems, backend engineering, and system design at &lt;strong&gt;&lt;a href="https://seeitflow.com/" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>backend</category>
      <category>architecture</category>
      <category>distributedsystems</category>
    </item>
    <item>
      <title>Designing a URL Shortener: Architecture, Base62, Redis Caching &amp; Scaling</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Thu, 08 Oct 2026 07:51:19 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/designing-a-url-shortener-architecture-base62-redis-caching-scaling-576n</link>
      <guid>https://dev.to/mangeshmandlik/designing-a-url-shortener-architecture-base62-redis-caching-scaling-576n</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9al1yl32ow14ihfqsol.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu9al1yl32ow14ihfqsol.png" alt="final architecture" width="800" height="352"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A URL shortener sounds like one of the simplest backend systems to build.&lt;/p&gt;

&lt;p&gt;Take a long URL:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;https://example.com/articles/distributed-systems/introduction&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Generate a short code:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;https://short.ly/aB7xK2p&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Store the mapping, and redirect anyone who visits the short URL.&lt;/p&gt;

&lt;p&gt;Three operations. One database table. A few API endpoints.&lt;/p&gt;

&lt;p&gt;So why is URL Shortener System Design such a useful distributed systems interview problem?&lt;/p&gt;

&lt;p&gt;Because generating the short URL is rarely the difficult part.&lt;/p&gt;

&lt;p&gt;The interesting engineering begins when a link gets shared with millions of people, thousands of requests hit the same database record, a cache restarts during peak traffic, or the analytics pipeline starts slowing down redirects.&lt;/p&gt;

&lt;p&gt;A production-grade URL shortener is fundamentally a problem of &lt;strong&gt;read-heavy architecture, uniqueness, caching, consistency, and failure isolation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Let's design one from first principles.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What Are We Actually Building?
&lt;/h2&gt;

&lt;p&gt;Before choosing Redis, Kafka, or a particular database, we need to understand what the system must do.&lt;/p&gt;

&lt;h3&gt;
  
  
  Functional requirements
&lt;/h3&gt;

&lt;p&gt;Our URL shortening service should support:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Creating a unique short URL from a long URL.&lt;/li&gt;
&lt;li&gt;Redirecting visitors to the original destination.&lt;/li&gt;
&lt;li&gt;Optional custom aliases, such as &lt;code&gt;short.ly/summer-sale&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Expiring links after a specified duration.&lt;/li&gt;
&lt;li&gt;Tracking click analytics, including counts and trends.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We will keep authentication, billing, QR codes, and advanced link-management features outside the initial scope.&lt;/p&gt;

&lt;h3&gt;
  
  
  Non-functional requirements
&lt;/h3&gt;

&lt;p&gt;The system should provide low-latency redirects, high availability, durable URL mappings, and the ability to absorb sudden traffic spikes.&lt;/p&gt;

&lt;p&gt;One requirement is especially important: &lt;strong&gt;redirect availability matters more than URL creation availability&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If users cannot create new links for a few minutes, that is inconvenient. If existing links stop redirecting, every application, email, or document containing those links can be affected.&lt;/p&gt;

&lt;p&gt;That difference will influence how we design the entire system.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Capacity Estimation: Why Reads Dominate the Architecture
&lt;/h2&gt;

&lt;p&gt;Let's establish an illustrative workload.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Design assumption&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Stored URL mappings&lt;/td&gt;
&lt;td&gt;100 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Daily redirects&lt;/td&gt;
&lt;td&gt;100 million&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read-to-write ratio&lt;/td&gt;
&lt;td&gt;Approximately 100:1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Average mapping storage&lt;/td&gt;
&lt;td&gt;500 bytes including overhead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Peak traffic&lt;/td&gt;
&lt;td&gt;Up to 100× average redirect rate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Redirect latency&lt;/td&gt;
&lt;td&gt;Low milliseconds on the cache-hit path&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These are assumptions for our design exercise, not universal URL-shortener benchmarks.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much storage do we need?
&lt;/h3&gt;

&lt;p&gt;With 100 million mappings at approximately 500 bytes each:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;100,000,000 × 500 bytes ≈ 50 GB&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That is a manageable amount of mapping data for a modern database deployment, although indexes, replicas, backups, and future growth require additional capacity.&lt;/p&gt;

&lt;p&gt;Storage is not necessarily our first scaling problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  How many redirects must we serve?
&lt;/h3&gt;

&lt;p&gt;100 million redirects per day translates to approximately:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;100,000,000 / 86,400 ≈ 1,157 redirects/second&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Now imagine a major traffic event pushes peak demand toward 100,000 redirects per second.&lt;/p&gt;

&lt;p&gt;If every redirect queries the database, the database repeatedly performs lookups for the same popular links.&lt;/p&gt;

&lt;p&gt;The system isn't doing complicated work. It is doing simple, repetitive work far too often.&lt;/p&gt;

&lt;p&gt;That distinction matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The first architectural optimization should target repeated reads, not immediately introduce database sharding.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For a more complete capacity model, see &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/complete-design" rel="noopener noreferrer"&gt;URL Shortener — Complete System Design&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Start With the Simplest Possible Architecture
&lt;/h2&gt;

&lt;p&gt;Our first version needs only three components:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
   |
   v
URL Shortener API
   |
   v
Relational Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application exposes two core operations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /shorten
Content-Type: application/json

{
  "long_url": "https://example.com/article"
}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"short_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://short.ly/aB7xK2p"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And for redirects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /aB7xK2p
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The service looks up the mapping and returns an appropriate HTTP redirect response with a &lt;code&gt;Location&lt;/code&gt; header.&lt;/p&gt;

&lt;p&gt;Our initial database schema can be remarkably small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;url_mappings&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;short_code&lt;/span&gt; &lt;span class="nb"&gt;VARCHAR&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;long_url&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="n"&gt;TIMESTAMPTZ&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;NOW&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="n"&gt;expires_at&lt;/span&gt; &lt;span class="n"&gt;TIMESTAMPTZ&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The primary-key index supports fast lookups by short code and enforces uniqueness.&lt;/p&gt;

&lt;p&gt;This architecture is perfectly reasonable for an early-stage product.&lt;/p&gt;

&lt;p&gt;There is no reason to operate Redis clusters, distributed ID generators, or multiple database regions before the workload requires them.&lt;/p&gt;

&lt;p&gt;But once popular URLs receive large numbers of repeated requests, our database begins doing unnecessary work.&lt;/p&gt;

&lt;p&gt;That is the first bottleneck we should solve.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. How Do URL Shorteners Generate Unique Short Codes?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1wirhufoj3a1iun8ajoc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1wirhufoj3a1iun8ajoc.png" alt="approaches to generate short code" width="799" height="244"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The short code is the identity of a URL mapping.&lt;/p&gt;

&lt;p&gt;We need it to be compact, unique, and ideally difficult to guess if enumeration is a concern.&lt;/p&gt;

&lt;p&gt;There are several approaches.&lt;/p&gt;

&lt;h3&gt;
  
  
  Approach A: Auto-increment IDs with Base62
&lt;/h3&gt;

&lt;p&gt;Suppose the database generates an integer ID.&lt;/p&gt;

&lt;p&gt;We convert that integer into Base62, using:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;0-9&lt;/code&gt;, &lt;code&gt;a-z&lt;/code&gt;, and &lt;code&gt;A-Z&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;This produces compact, URL-safe representations.&lt;/p&gt;

&lt;p&gt;Advantages include straightforward uniqueness and minimal collision handling.&lt;/p&gt;

&lt;p&gt;The disadvantage is predictability. Sequential identifiers can be enumerated, even when encoded in Base62.&lt;/p&gt;

&lt;h3&gt;
  
  
  Approach B: Hash the original URL
&lt;/h3&gt;

&lt;p&gt;We could hash the long URL and use part of the resulting value as the short code.&lt;/p&gt;

&lt;p&gt;This can support deterministic mappings, where the same input produces the same output.&lt;/p&gt;

&lt;p&gt;But truncated hashes can collide.&lt;/p&gt;

&lt;p&gt;It also creates a product question: should two different users shortening the same destination receive the same code?&lt;/p&gt;

&lt;p&gt;For campaign tracking, ownership, or independent expiry policies, they may need different codes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Approach C: Random Base62
&lt;/h3&gt;

&lt;p&gt;Generate a random code from the Base62 alphabet.&lt;/p&gt;

&lt;p&gt;The available namespace grows exponentially with code length.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Code length&lt;/th&gt;
&lt;th&gt;Possible codes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;6 characters&lt;/td&gt;
&lt;td&gt;56.8 billion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;7 characters&lt;/td&gt;
&lt;td&gt;3.52 trillion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8 characters&lt;/td&gt;
&lt;td&gt;218 trillion&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A seven-character code offers a large namespace for our illustrative workload.&lt;/p&gt;

&lt;p&gt;But large does not mean collision-free.&lt;/p&gt;

&lt;p&gt;With 100 million occupied codes in a seven-character namespace, the probability that the next uniformly random attempt collides is approximately:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;100,000,000 / 62^7 ≈ 0.00284%&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That is a small per-attempt probability, but collisions will still occur over a sufficiently large number of creations.&lt;/p&gt;

&lt;p&gt;The solution is not to pretend collisions are impossible.&lt;/p&gt;

&lt;p&gt;It is to make them harmless.&lt;/p&gt;

&lt;p&gt;Use a database uniqueness constraint, and retry with a new random code when a collision occurs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Generate random Base62 code
           |
           v
Attempt database INSERT
           |
      +----+----+
      |         |
   Success   Collision
      |         |
    Return   Generate again
              (bounded retries)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The retry count must be bounded. An unexpected increase in collisions should be observable rather than producing an infinite loop.&lt;/p&gt;

&lt;h3&gt;
  
  
  Approach D: Snowflake-style distributed IDs
&lt;/h3&gt;

&lt;p&gt;A Snowflake-style generator combines information such as a timestamp, worker identifier, and sequence counter to generate unique IDs without coordinating with the database on every allocation.&lt;/p&gt;

&lt;p&gt;It can be useful for very high creation throughput or independent generation across many nodes.&lt;/p&gt;

&lt;p&gt;However, it introduces clock-handling and worker-ID management requirements.&lt;/p&gt;

&lt;p&gt;And importantly, Base62-encoding a Snowflake ID does not make it unpredictable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Our choice:&lt;/strong&gt; For this workload, random Base62 with a database uniqueness constraint and bounded retries is a sensible starting point.&lt;/p&gt;

&lt;p&gt;Snowflake is an alternative when decentralized ID allocation becomes a genuine requirement, not an automatic upgrade merely because multiple application servers exist.&lt;/p&gt;

&lt;p&gt;For collision mathematics, namespace sizing, and distributed ID mechanics, see &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/architecture-deep-dive" rel="noopener noreferrer"&gt;URL Shortener — Architecture Deep Dive&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Choosing the Database: SQL or NoSQL?
&lt;/h2&gt;

&lt;p&gt;A common system design question is:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Should a URL shortener use PostgreSQL, MySQL, DynamoDB, or another distributed key-value store?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The answer depends on the workload.&lt;/p&gt;

&lt;p&gt;Our dominant query is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="n"&gt;long_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expires_at&lt;/span&gt;
&lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;url_mappings&lt;/span&gt;
&lt;span class="k"&gt;WHERE&lt;/span&gt; &lt;span class="n"&gt;short_code&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s1"&gt;'aB7xK2p'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We need exact-key lookups, durable writes, and reliable uniqueness enforcement.&lt;/p&gt;

&lt;p&gt;A relational database such as PostgreSQL or MySQL handles these requirements well.&lt;/p&gt;

&lt;p&gt;A distributed key-value database can also be appropriate, especially when horizontal write scaling, global distribution, or multi-region availability becomes a stronger requirement.&lt;/p&gt;

&lt;p&gt;But choosing NoSQL solely because the application might eventually serve millions of redirects is not sufficient reasoning.&lt;/p&gt;

&lt;p&gt;If most redirects are served from cache, the database never sees most of that traffic.&lt;/p&gt;

&lt;p&gt;We should scale the database based on its actual workload, not the public request rate.&lt;/p&gt;

&lt;p&gt;There is another important trade-off: database replicas.&lt;/p&gt;

&lt;p&gt;Routing cache misses to read replicas protects the primary's write capacity. However, asynchronously replicated data may be temporarily stale.&lt;/p&gt;

&lt;p&gt;A short URL created moments ago might exist on the primary but not yet on a replica.&lt;/p&gt;

&lt;p&gt;If the cache misses during that window, the system could incorrectly return 404.&lt;/p&gt;

&lt;p&gt;Eager cache population after a successful write, read-after-write routing, or stronger replication guarantees can mitigate this, depending on the consistency requirement.&lt;/p&gt;

&lt;p&gt;The detailed alternatives are compared in &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/tradeoffs-decisions" rel="noopener noreferrer"&gt;URL Shortener — Trade-offs &amp;amp; Decisions&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Redis Caching: The Most Important Scaling Decision
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4nnmgejqdztcrxemzl25.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4nnmgejqdztcrxemzl25.png" alt="with caching" width="800" height="261"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now we address our main bottleneck.&lt;/p&gt;

&lt;p&gt;Without caching, every redirect requires a database lookup.&lt;/p&gt;

&lt;p&gt;With Redis, popular mappings can be served from memory.&lt;/p&gt;

&lt;p&gt;The request path becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
   |
   v
Load Balancer
   |
   v
URL Service
   |
   +----&amp;gt; Redis
   |        |
   |     Cache Hit ----&amp;gt; Redirect
   |
   +----&amp;gt; Cache Miss
              |
              v
         Database
              |
              v
        Populate Redis
              |
              v
           Redirect
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the cache-aside pattern.&lt;/p&gt;

&lt;p&gt;On URL creation, we can also populate Redis immediately after the database successfully commits the mapping.&lt;/p&gt;

&lt;p&gt;That reduces the chance that a newly created link experiences a cold-cache miss when its first visitors arrive.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much database traffic can caching eliminate?
&lt;/h3&gt;

&lt;p&gt;Assume 100,000 redirects per second and a measured or estimated cache-hit ratio of 92%.&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Database reads = 100,000 × (1 - 0.92)&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Database reads ≈ 8,000 per second&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The cache absorbs approximately 92,000 reads per second.&lt;/p&gt;

&lt;p&gt;This illustrates why caching can dramatically reduce database pressure. The actual hit ratio must be measured; it depends on traffic distribution, memory capacity, and expiry policies.&lt;/p&gt;

&lt;h3&gt;
  
  
  The subtle problem: cache expiry correctness
&lt;/h3&gt;

&lt;p&gt;Suppose a short URL expires in 10 minutes.&lt;/p&gt;

&lt;p&gt;But Redis stores the mapping with a fixed TTL of 24 hours.&lt;/p&gt;

&lt;p&gt;After 10 minutes, the authoritative mapping is expired. Yet Redis may continue redirecting users for almost another day.&lt;/p&gt;

&lt;p&gt;This is a correctness bug, not just a caching inefficiency.&lt;/p&gt;

&lt;p&gt;A safer TTL rule is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cache_ttl = min(
    configured_cache_ttl,
    expires_at - current_time
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For non-expiring mappings, use the configured cache TTL.&lt;/p&gt;

&lt;p&gt;The redirect handler should also check logical expiry when that metadata is available.&lt;/p&gt;

&lt;p&gt;Explicit updates or deletions require cache invalidation rather than relying exclusively on TTL expiry.&lt;/p&gt;

&lt;p&gt;The complete cache hierarchy, expiry rules, and invalidation mechanisms are covered in the &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/architecture-deep-dive" rel="noopener noreferrer"&gt;Architecture Deep Dive&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. What Happens When a Short URL Goes Viral?
&lt;/h2&gt;

&lt;p&gt;Imagine a short link appears in a major social-media post.&lt;/p&gt;

&lt;p&gt;Within seconds, hundreds of thousands of people click it.&lt;/p&gt;

&lt;p&gt;A warm Redis cache helps because most requests avoid the database.&lt;/p&gt;

&lt;p&gt;But a different problem can appear.&lt;/p&gt;

&lt;p&gt;Every request for the same short code may hit the same Redis key.&lt;/p&gt;

&lt;p&gt;Even in a sharded Redis cluster, that key is generally owned by one shard.&lt;/p&gt;

&lt;p&gt;Adding more shards distributes different keys. It does not automatically distribute reads for one extremely popular key.&lt;/p&gt;

&lt;p&gt;This is called a &lt;strong&gt;hot-key problem&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do we solve it?
&lt;/h3&gt;

&lt;p&gt;One option is a small process-local cache inside each URL service instance.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                  Load Balancer
                       |
          +------------+------------+
          |            |            |
       Service A    Service B    Service C
          |            |            |
       Local Cache  Local Cache  Local Cache
          |            |            |
          +------------+------------+
                       |
                     Redis
                       |
                    Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If each application instance caches the extremely popular mapping locally, many requests no longer reach Redis.&lt;/p&gt;

&lt;p&gt;Other options include serving reads through cache replicas or using CDN edge caching for genuinely immutable redirects.&lt;/p&gt;

&lt;p&gt;Each choice adds different consistency and invalidation considerations.&lt;/p&gt;

&lt;h3&gt;
  
  
  What if Redis restarts during the viral spike?
&lt;/h3&gt;

&lt;p&gt;Now the cache is empty.&lt;/p&gt;

&lt;p&gt;Thousands of requests for the same URL can simultaneously miss and query the database.&lt;/p&gt;

&lt;p&gt;That is a &lt;strong&gt;cache stampede&lt;/strong&gt;, also called a thundering herd.&lt;/p&gt;

&lt;p&gt;Request coalescing solves an important part of the problem: concurrent misses for the same code share one in-flight database lookup rather than issuing duplicate queries.&lt;/p&gt;

&lt;p&gt;TTL jitter and proactive warming of known hot keys can reduce the risk further.&lt;/p&gt;

&lt;p&gt;A viral URL, a Redis hot key, and a cold-cache stampede are related but distinct operational scenarios.&lt;/p&gt;

&lt;p&gt;See &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/failure-scale-scenarios" rel="noopener noreferrer"&gt;URL Shortener — Failure &amp;amp; Scale Scenarios&lt;/a&gt; for the incident-by-incident breakdown.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. 301 vs 302 Redirect: A Small Decision With Big Consequences
&lt;/h2&gt;

&lt;p&gt;Should a URL shortener return &lt;code&gt;301 Moved Permanently&lt;/code&gt; or &lt;code&gt;302 Found&lt;/code&gt;?&lt;/p&gt;

&lt;p&gt;There is no universally correct answer.&lt;/p&gt;

&lt;p&gt;A permanent redirect communicates that the destination is intended to remain stable.&lt;/p&gt;

&lt;p&gt;A temporary redirect communicates that the destination may change.&lt;/p&gt;

&lt;p&gt;The choice affects browser caching, infrastructure load, and analytics visibility.&lt;/p&gt;

&lt;p&gt;For example, if a browser caches a permanent redirect, later visits may bypass the URL shortener entirely.&lt;/p&gt;

&lt;p&gt;That reduces load on our infrastructure.&lt;/p&gt;

&lt;p&gt;But those visits may also become invisible to origin-side click analytics.&lt;/p&gt;

&lt;p&gt;If a short URL's destination can be edited, long-lived client caching becomes more problematic because visitors may continue reaching an old destination.&lt;/p&gt;

&lt;p&gt;Temporary redirects can keep requests visible to the service, provided the caching headers and client behavior are configured accordingly.&lt;/p&gt;

&lt;p&gt;Remember that HTTP caching depends on &lt;strong&gt;status code, response headers, and client behavior together&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It is incorrect to assume that every 301 is cached forever or that a 302 can never be cached.&lt;/p&gt;

&lt;p&gt;For mutable destinations or products prioritizing click visibility, a temporary redirect posture is often easier to manage.&lt;/p&gt;

&lt;p&gt;For immutable, heavily accessed links, permanent redirects with appropriate caching policies may reduce origin traffic significantly.&lt;/p&gt;

&lt;p&gt;This decision is explored further in &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/tradeoffs-decisions" rel="noopener noreferrer"&gt;Trade-offs &amp;amp; Decisions — Permanent vs Temporary Redirect&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Click Analytics Should Not Slow Down Redirects
&lt;/h2&gt;

&lt;p&gt;A URL shortener may need to record every click for dashboards, reporting, or campaign analytics.&lt;/p&gt;

&lt;p&gt;The naive approach is to insert an analytics record before returning each redirect.&lt;/p&gt;

&lt;p&gt;But that makes redirect latency depend on the analytics database.&lt;/p&gt;

&lt;p&gt;At 100,000 redirects per second, we could generate 100,000 analytics events per second.&lt;/p&gt;

&lt;p&gt;Worse, if the analytics store slows down, the redirect path slows down too.&lt;/p&gt;

&lt;p&gt;These are two workloads with different requirements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Redirect serving needs low latency and high availability.&lt;/li&gt;
&lt;li&gt;Analytics processing needs high throughput and can often tolerate delayed aggregation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We should isolate them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                   URL Service
                   /         \
                  /           \
           Return Redirect   Async Event
                                  |
                                  v
                            Kafka / SQS
                                  |
                                  v
                          Analytics Workers
                                  |
                                  v
                           Analytics Store
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The redirect service emits an event through a bounded, asynchronous mechanism.&lt;/p&gt;

&lt;p&gt;Workers consume events, batch them, and write them into an analytics store such as ClickHouse.&lt;/p&gt;

&lt;p&gt;But asynchronous processing is not automatically safe.&lt;/p&gt;

&lt;p&gt;What happens if Kafka or SQS becomes unavailable?&lt;/p&gt;

&lt;p&gt;The service must have an explicit policy: drop, sample, or temporarily buffer events within strict limits, depending on whether analytics loss is acceptable.&lt;/p&gt;

&lt;p&gt;The redirect should not block indefinitely waiting for analytics.&lt;/p&gt;

&lt;p&gt;If the product requires exact, lossless click counts, the design needs stronger event durability and deduplication guarantees.&lt;/p&gt;

&lt;h3&gt;
  
  
  Another subtle scaling issue: analytics partitioning
&lt;/h3&gt;

&lt;p&gt;If we partition events by short code, all clicks for one viral URL may land on the same partition.&lt;/p&gt;

&lt;p&gt;That creates a hot partition even if the rest of the event-streaming cluster has available capacity.&lt;/p&gt;

&lt;p&gt;If the analytics workload primarily computes counts and aggregates, distributing events more broadly and merging counts downstream may be preferable.&lt;/p&gt;

&lt;p&gt;Strict per-URL event ordering is not always necessary.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/architecture-deep-dive" rel="noopener noreferrer"&gt;Architecture Deep Dive&lt;/a&gt; covers event-delivery guarantees, partitioning alternatives, and duplicate-count handling.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. The Production Architecture
&lt;/h2&gt;

&lt;p&gt;We can now assemble the architecture based on the bottlenecks we have actually identified.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     Clients
                        |
                        v
                   Load Balancer
                        |
                        v
              Stateless URL Services
              /          |         \
             /           |          \
       Local Cache     Redis     Analytics Events
             |           |              |
             +-----+-----+              v
                   |                 Kafka / SQS
                   v                     |
             Database Layer             v
              /         \         Analytics Workers
          Primary     Replicas           |
                                        v
                                  Analytics Store
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each component has a specific responsibility.&lt;/p&gt;

&lt;p&gt;The load balancer distributes requests across stateless application instances.&lt;/p&gt;

&lt;p&gt;The URL services handle validation, short-code generation, cache lookups, expiry checks, and redirects.&lt;/p&gt;

&lt;p&gt;Redis absorbs repeated mapping reads, while optional local caches protect extremely hot keys.&lt;/p&gt;

&lt;p&gt;The relational database remains the authoritative source of URL mappings and uniqueness.&lt;/p&gt;

&lt;p&gt;The event pipeline processes analytics independently.&lt;/p&gt;

&lt;p&gt;This is a logical architecture, not a prescription for a fixed number of servers, replicas, or cache nodes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The goal is not to add as many distributed components as possible. It is to introduce each component when it solves a demonstrated problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To understand how the simple database-only design evolves into this architecture, explore the &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/visual-design-journey" rel="noopener noreferrer"&gt;interactive URL Shortener Visual Design Journey&lt;/a&gt; on SeeItFlow.&lt;/p&gt;

&lt;p&gt;It walks through the architecture's evolution with animated, narrated explanations rather than relying only on a static diagram.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Production Failures That Reveal Whether the Design Is Correct
&lt;/h2&gt;

&lt;p&gt;A system design is incomplete until we understand what happens when dependencies fail.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario A: Redis becomes unavailable
&lt;/h3&gt;

&lt;p&gt;All requests that previously hit Redis may fall through to the database.&lt;/p&gt;

&lt;p&gt;Redirects can continue if the database has sufficient capacity.&lt;/p&gt;

&lt;p&gt;But if it cannot absorb the sudden read increase, a cache outage can cascade into a database outage.&lt;/p&gt;

&lt;p&gt;Circuit breakers, bounded database concurrency, request coalescing, and local caching help contain the damage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario B: The database primary fails
&lt;/h3&gt;

&lt;p&gt;New URL creation becomes unavailable until a writable primary is restored or a replica is promoted.&lt;/p&gt;

&lt;p&gt;Existing redirects may continue through Redis and healthy read replicas.&lt;/p&gt;

&lt;p&gt;The exact impact depends on the replication topology and the freshness of available cached data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario C: A newly created URL returns 404
&lt;/h3&gt;

&lt;p&gt;The write succeeded on the primary, but a cache miss reached a lagging read replica.&lt;/p&gt;

&lt;p&gt;The replica has not yet received the new mapping.&lt;/p&gt;

&lt;p&gt;This is a read-after-write consistency problem.&lt;/p&gt;

&lt;p&gt;Eager cache population, primary reads for recently created mappings, or stronger replication guarantees can mitigate it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario D: Two users request the same custom alias
&lt;/h3&gt;

&lt;p&gt;Both requests check availability and initially find the alias unused.&lt;/p&gt;

&lt;p&gt;Both attempt to create it.&lt;/p&gt;

&lt;p&gt;An application-level availability check cannot safely resolve this race.&lt;/p&gt;

&lt;p&gt;The database uniqueness constraint must arbitrate the conflict atomically. One request succeeds, while the other receives a conflict response.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario E: The analytics broker fails
&lt;/h3&gt;

&lt;p&gt;If event publishing blocks indefinitely, analytics infrastructure can take down redirect serving.&lt;/p&gt;

&lt;p&gt;A bounded, explicitly defined producer failure policy keeps the two availability domains separated.&lt;/p&gt;

&lt;p&gt;These are only some of the important cases.&lt;/p&gt;

&lt;p&gt;The full &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/failure-scale-scenarios" rel="noopener noreferrer"&gt;Failure &amp;amp; Scale Scenarios&lt;/a&gt; reference covers Redis failures, cache stampedes, viral spikes, replica lag, namespace pressure, expired mappings, broker outages, and abuse traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. How Would This Architecture Evolve Over Time?
&lt;/h2&gt;

&lt;p&gt;We should not deploy the final architecture on day one.&lt;/p&gt;

&lt;p&gt;A more sensible progression is:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Stage&lt;/th&gt;
&lt;th&gt;Architecture&lt;/th&gt;
&lt;th&gt;Reason to evolve&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Application + relational database&lt;/td&gt;
&lt;td&gt;Initial product with manageable traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Add Redis cache&lt;/td&gt;
&lt;td&gt;Repeated redirect reads begin stressing the database&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Stateless services, replicas, async analytics&lt;/td&gt;
&lt;td&gt;Application throughput and analytics isolation become important&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Local hot-key caching and additional capacity&lt;/td&gt;
&lt;td&gt;Concentrated viral traffic exceeds shared-cache or application capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Multi-region reads or writes&lt;/td&gt;
&lt;td&gt;Geographic latency, availability, or independent regional write requirements justify the complexity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The trigger for each transition should be a measured bottleneck or a new product requirement.&lt;/p&gt;

&lt;p&gt;Multi-region writes, for example, introduce global uniqueness, replication consistency, and cross-region invalidation problems.&lt;/p&gt;

&lt;p&gt;A single write region with geographically distributed reads may be sufficient for many products.&lt;/p&gt;

&lt;p&gt;Adding active-active writes merely because the system is large would create complexity without necessarily improving the user experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. URL Shortener System Design Interview Questions
&lt;/h2&gt;

&lt;p&gt;Here are some of the most important questions to prepare for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is Redis useful in a URL shortener?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Because a read-heavy workload with skewed access patterns repeatedly requests the same mappings. Redis can serve popular mappings from memory, reducing database load and redirect latency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you prevent duplicate short URLs?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For randomly generated codes, enforce uniqueness at the database layer and retry generation on a conflict. If the product requires deterministic deduplication of destination URLs, that is a separate rule with its own identity and concurrency requirements.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Snowflake required for distributed ID generation?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. Random Base62 with database uniqueness enforcement or a database sequence can work across multiple application instances. Snowflake becomes useful when decentralized generation offers a concrete advantage.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens if Redis fails?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Reads can fall back to the database, but the additional load must be bounded. Otherwise, the cache failure can trigger a cascading database overload.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Should the URL mapping database be SQL or NoSQL?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A relational database is a strong starting choice for exact-key lookups and uniqueness constraints. A distributed key-value store becomes more attractive when global distribution or horizontal write scalability justifies it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you handle millions of clicks on one short URL?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Keep the mapping cached, identify hot-key concentration, consider process-local caching or read replicas, and ensure the application tier and analytics producer can handle the request volume. The first bottleneck depends on the actual deployment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you ensure expired links stop redirecting?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Enforce logical expiry on reads and ensure cache entries cannot outlive the authoritative mapping's expiry. Background row deletion alone is insufficient.&lt;/p&gt;

&lt;p&gt;For a much larger collection of follow-up questions and reasoning strategies, see &lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/interview-perspective" rel="noopener noreferrer"&gt;URL Shortener — Interview Perspective&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;A URL shortener teaches an important lesson about distributed systems.&lt;/p&gt;

&lt;p&gt;The difficult part is not generating a seven-character string or writing a database query.&lt;/p&gt;

&lt;p&gt;It is understanding how a small operation behaves when repeated millions of times under unpredictable conditions.&lt;/p&gt;

&lt;p&gt;The strongest architecture emerges from a sequence of questions:&lt;/p&gt;

&lt;p&gt;Why is the database doing the same lookup repeatedly?&lt;/p&gt;

&lt;p&gt;What happens if our cache disappears?&lt;/p&gt;

&lt;p&gt;Can a recently created URL temporarily return 404?&lt;/p&gt;

&lt;p&gt;Does our analytics pipeline have the power to break redirects?&lt;/p&gt;

&lt;p&gt;Are we solving a real scaling problem, or adding infrastructure because it looks sophisticated?&lt;/p&gt;

&lt;p&gt;Answering these questions leads to better architecture than memorizing a diagram.&lt;/p&gt;

&lt;p&gt;And that reasoning applies far beyond URL shorteners — to API gateways, caching systems, distributed databases, event-driven platforms, and other read-heavy services.&lt;/p&gt;




&lt;h2&gt;
  
  
  References &amp;amp; Further Reading
&lt;/h2&gt;

&lt;p&gt;The following SeeItFlow resources explore the design in greater depth:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/complete-design" rel="noopener noreferrer"&gt;URL Shortener — Complete Design&lt;/a&gt;&lt;/strong&gt; — Requirements, capacity estimation, ID generation, caching, data model, final architecture, and production evolution.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/visual-design-journey" rel="noopener noreferrer"&gt;URL Shortener — Visual Design Journey&lt;/a&gt;&lt;/strong&gt; — Interactive, narrated architectural evolution from a simple database-backed service to a distributed redirect system.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/architecture-deep-dive" rel="noopener noreferrer"&gt;URL Shortener — Architecture Deep Dive&lt;/a&gt;&lt;/strong&gt; — Base62 namespace sizing, collision handling, Snowflake IDs, cache hierarchy, hot-key protection, database replication, and analytics partitioning.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/tradeoffs-decisions" rel="noopener noreferrer"&gt;URL Shortener — Trade-offs &amp;amp; Decisions&lt;/a&gt;&lt;/strong&gt; — Alternative approaches and selection criteria for databases, redirects, caching, ID generation, and analytics.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/failure-scale-scenarios" rel="noopener noreferrer"&gt;URL Shortener — Failure &amp;amp; Scale Scenarios&lt;/a&gt;&lt;/strong&gt; — Twelve production failure and scaling scenarios, including triggers, detection, mitigation, and recovery.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;&lt;a href="https://seeitflow.com/system-design/foundations/url-shortener/interview-perspective" rel="noopener noreferrer"&gt;URL Shortener — Interview Perspective&lt;/a&gt;&lt;/strong&gt; — Interview preparation, architecture reasoning, common mistakes, and advanced follow-up questions.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;All six resources are part of &lt;a href="https://seeitflow.com/" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, a free visual software-engineering learning platform covering system design, distributed systems, backend engineering, databases, networking, and modern AI systems.&lt;/p&gt;

</description>
      <category>systemdesign</category>
      <category>architecture</category>
      <category>distributed</category>
      <category>backend</category>
    </item>
    <item>
      <title>Nobody Remembers Why We Chose This, So Nobody Will Change It</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:15:54 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/nobody-remembers-why-we-chose-this-so-nobody-will-change-it-3o80</link>
      <guid>https://dev.to/mangeshmandlik/nobody-remembers-why-we-chose-this-so-nobody-will-change-it-3o80</guid>
      <description>&lt;p&gt;A team picks a read replica to fix database latency under load. It works, latency drops, everyone moves on. Eighteen months later, a new engineer joins, looks at the replication setup, and asks why reads are eventually consistent instead of strongly consistent, since that's clearly adding complexity somewhere downstream. Nobody still on the team remembers the original reasoning. The engineer, reasonably, proposes removing it to simplify things. It gets removed. Three weeks later, the database is saturated again during peak hours, and the team is back to solving a problem they already solved once, except this time with no memory of how.&lt;/p&gt;

&lt;p&gt;Nothing about this story involves a bad technology choice. The read replica was the right call both times it got considered. The actual failure was that the first decision never got written down anywhere a future engineer could find it, so it looked, eighteen months later, like an arbitrary complexity someone had added for no clear reason. That's the problem architecture, as a discipline, actually exists to solve, and it's a narrower claim than it sounds like at first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture is a set of decisions, not a diagram
&lt;/h2&gt;

&lt;p&gt;It's tempting to picture "architecture" as the boxes-and-arrows diagram that gets drawn on a whiteboard or dropped into a design doc. The diagram is a byproduct. What architecture actually is: the set of decisions that are expensive to change, the constraints a team deliberately chooses to live with so the system can meet its real requirements under real-world conditions.&lt;/p&gt;

&lt;p&gt;Those decisions get shaped by three things pulling against each other. Requirements are what the system actually has to do, both the functional behavior and the non-functional expectations around it. Constraints are the hard limits already in place, budget, team size, a deadline, a compliance requirement, things that aren't negotiable no matter how elegant an alternative looks. Quality attributes describe how the system has to behave once it's running, its latency, its availability, its durability, its cost, how operable it is for the people on call for it. A decision that ignores any one of these three tends to look fine in a design review and then quietly becomes the wrong choice the moment it meets production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why undocumented decisions decay into mysteries
&lt;/h2&gt;

&lt;p&gt;Most architectural failures aren't the result of picking the wrong technology. They're the result of a decision that got made implicitly, without anyone writing down the requirement it was satisfying, the alternatives that were considered and rejected, or the specific condition under which it should eventually get revisited.&lt;/p&gt;

&lt;p&gt;Once that documentation doesn't exist, every new engineer who touches the system has to re-evaluate the decision from scratch, and they frequently end up re-selecting an option the team already tried and rejected, for reasons that are no longer visible to anyone. Explicit reasoning, written down somewhere, is what lets that shared understanding survive team turnover and the organization simply growing past the people who made the original call. The read replica story at the top is exactly this pattern: the decision was fine, the absence of a record was the actual bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reasoning loop underneath a good decision
&lt;/h2&gt;

&lt;p&gt;Stripped down, the architectural reasoning process is a short loop, and skipping a step in it tends to be where things go wrong later, not in any single step itself.&lt;/p&gt;

&lt;p&gt;Start by identifying the requirements: what the system must actually do, and which quality attributes, availability, latency, consistency, cost, matter most for this particular decision. Then identify the constraints, team size, budget, deadline, what already exists, any compliance obligations, the hard limits that rule certain options out before you've even evaluated them. Next, enumerate every viable option before judging any of them, collapsing the option space too early is one of the most common ways teams end up with a worse decision than they needed to make. Evaluate the trade-offs of each option against the criteria that actually came from the requirements and constraints, not against whichever option happens to be trendiest. Decide, and write down why, the decision itself, which alternatives were rejected and why, and a specific trigger condition for when this decision should be revisited. And finally, acknowledge the consequences explicitly, every decision introduces new constraints of its own, and naming them up front is what lets a future team actually plan around them instead of discovering them by accident.&lt;/p&gt;

&lt;p&gt;The artifact that captures most of this is usually called an Architecture Decision Record, an ADR, and the single most valuable section in one is almost always the rejected options. That section is specifically what stops a future engineer from re-opening a question the team already closed, for reasons that made sense at the time and are otherwise invisible to anyone who wasn't in the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question experienced engineers actually ask
&lt;/h2&gt;

&lt;p&gt;The best architecture isn't the most sophisticated one available. It's the one that satisfies the current requirements with the least unnecessary complexity, while still being genuinely evolvable once requirements change, which they reliably will.&lt;/p&gt;

&lt;p&gt;Principal-level engineers tend to ask one specific question more than any other: "what would change my decision?" That's the mark of conditional reasoning rather than a fixed, context-free preference for one architecture over another. A cache is the right call precisely when some amount of stale data is acceptable, not on principle. A queue is the right call when asynchronous processing is tolerable for that specific workflow. Multi-region deployment is the right call when the cost of downtime genuinely exceeds the cost of the operational complexity multi-region adds, not just because the company is now large enough that it feels like the next step. Architecture isn't really about knowing the correct answer in the abstract. It's about knowing which questions to ask, and which of the constraints in front of you actually matter most for this particular decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watching one system evolve through three stages
&lt;/h2&gt;

&lt;p&gt;A small team launches with the simplest possible setup, one application server, one PostgreSQL database, deploys running straight from someone's laptop. Response times sit comfortably under 50ms. Nothing about this is wrong, it's exactly the right amount of architecture for where the system actually is.&lt;/p&gt;

&lt;p&gt;Users grow by a factor of ten. The database's CPU starts saturating during peak hours, and read latency climbs. The team adds a read replica, writes still go to the primary, and reads now come with replication lag, eventual consistency instead of strong consistency. That's a real trade-off accepted deliberately, not a flaw slipping in unnoticed.&lt;/p&gt;

&lt;p&gt;Traffic becomes globally distributed. Users in Asia are seeing 200ms-plus latency that a read replica does nothing for, since it's a geography problem, not a database-load problem. The team adds a CDN for static assets and starts seriously considering a multi-region deployment. But multi-region brings real data-consistency challenges, meaningfully more operational complexity, and real cost. The honest question at this stage isn't "should we go multi-region," it's what specific requirement would actually make multi-region necessary right now, and which constraints make it premature at this particular size. Different teams can legitimately land on different answers here, the point isn't that there's one correct stage to make this jump, it's that the decision should be made against an actual requirement, not against a vague sense that the company has gotten big enough that it's time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade-offs that show up in nearly every decision
&lt;/h2&gt;

&lt;p&gt;A handful of tensions recur across almost any architectural decision you'll actually face. Simplicity versus scalability: a monolith is simpler to operate day to day but harder to scale independently piece by piece, while microservices scale more flexibly at the cost of a real multiplication in operational burden. Consistency versus availability: strong consistency requires coordination, which costs latency and can reduce availability during a network partition, while eventual consistency buys higher availability at the cost of application logic that now has to account for staleness. Performance versus cost: more replicas, more caching layers, more regions all improve performance, and each one adds real infrastructure cost and real operational complexity that someone has to carry. Flexibility versus commitment: delaying a decision keeps options open longer, but it can also quietly accumulate technical debt, while committing early enables real optimization at the cost of reduced room to adapt later if the assumption turns out wrong.&lt;/p&gt;

&lt;p&gt;None of these trade-offs resolve to a universal right answer. They resolve differently depending on the actual requirements and constraints in front of a given team at a given stage, which is exactly why the reasoning loop above matters more than memorizing which side of each trade-off is "correct."&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure modes that show up again and again
&lt;/h2&gt;

&lt;p&gt;A specific, recurring handful of failure patterns account for a disproportionate share of real architectural pain. Implicit decisions, a choice made with no documentation, so six months later nobody remembers why it was made, and it gets reversed, reintroducing the exact problem it was originally meant to solve, is the read-replica story from the opening, and it's extremely common. Solution-first framing starts from "we need Kafka" instead of starting from the actual requirement, collapsing the option space before any real evaluation happens, and frequently building the wrong fit as a result. No change trigger means a decision that was genuinely correct back at an earlier stage is still in place at a much later stage, with nobody having recorded when it should be revisited, so the old architecture quietly becomes a constraint on everything built after it. Over-engineering early looks like choosing microservices for a three-person team because "we'll need it eventually," and watching the resulting operational complexity consume engineering capacity that could have gone toward the product itself. Ignoring constraints means picking the technically superior option on paper, one that happens to require expertise the team doesn't actually have, and discovering that the correct choice in theory is the wrong choice in practice. And architecture by committee, requiring full consensus on every decision regardless of how reversible it actually is, treats cheap, easily-undone choices with the same ceremony as genuinely irreversible ones, and velocity collapses under the weight of that mismatch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not every decision deserves the same amount of rigor
&lt;/h2&gt;

&lt;p&gt;Applying the full reasoning loop, options, trade-offs, a written ADR, to every single decision is its own kind of failure, the architecture-by-committee pattern above. The actual skill is matching the rigor to how reversible the decision is.&lt;/p&gt;

&lt;p&gt;Full rigor earns its cost when choosing a primary datastore or data model, defining a public API contract, committing to a platform-level technology, making a call that affects multiple teams at once, or any decision where the cost of changing course later is genuinely high. Moving fast is the right call when choosing an internal library or a local caching strategy, deciding the shape of an internal API within a single service, making something that's reversible within hours or days if it turns out wrong, when the blast radius stays bounded to one team, or when the choice can just be validated with a quick spike instead of a formal process. The mismatch in either direction, heavy process on a trivial, reversible choice, or an implicit, undocumented call on something genuinely hard to undo, is where most of the pain in the failure modes above actually comes from.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sentence worth remembering
&lt;/h2&gt;

&lt;p&gt;If a startup's monolith is slowing feature delivery and a senior engineer proposes a full microservices rewrite, the most important question isn't which framework to use or how many services to split into. It's what specific constraint the monolith is actually creating right now, since that's the requirement any proposed solution, rewrite included, actually needs to satisfy, and skipping straight to the solution is the exact solution-first framing mistake from above.&lt;/p&gt;

&lt;p&gt;The single most useful sentence in any architectural decision, the one worth writing down every time, is "revisit this if...". It's what turns evolution into something intentional, a planned transition triggered by a condition someone actually named in advance, rather than a surprise rework triggered by nobody remembering why the original decision was made in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Explore It Visually
&lt;/h2&gt;

&lt;p&gt;I put together an interactive walkthrough of this entire reasoning cycle, requirements, constraints, options, trade-offs, the decision and its ADR, and then watching a system evolve through an actual change trigger, on &lt;a href="https://seeitflow.com/architecture/modules/architecture-fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to see the cycle play out scene by scene rather than read through it linearly.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>systemdesign</category>
      <category>backend</category>
      <category>software</category>
    </item>
    <item>
      <title>A CDN Isn't Just Making Things Faster. It's Protecting Your Origin From Its Own Traffic.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Thu, 01 Oct 2026 14:05:36 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/a-cdn-isnt-just-making-things-faster-its-protecting-your-origin-from-its-own-traffic-39ic</link>
      <guid>https://dev.to/mangeshmandlik/a-cdn-isnt-just-making-things-faster-its-protecting-your-origin-from-its-own-traffic-39ic</guid>
      <description>&lt;p&gt;Picture a user in Mumbai requesting content from a server sitting in Virginia. No matter how fast that server is, no matter how optimized the application code running on it happens to be, the request still has to physically travel there and back, and that round trip has a hard floor set by the speed of light over fiber, measured in hundreds of milliseconds. You cannot optimize your way out of geography with a faster CPU.&lt;/p&gt;

&lt;p&gt;That's the first reason a CDN exists, and it's the one most people already know. The second reason is less obvious and arguably more important: a single origin server is a single point of load. If a piece of content suddenly goes viral, every one of those requests would otherwise land on that one server, all at once, regardless of whether it was ever provisioned to handle that kind of traffic. A CDN turns "thousands of requests hitting the origin" into "one request hits the origin, thousands get served from a cache somewhere else entirely." Latency and origin protection are both the real answer here, not just one or the other. I put together a visual walkthrough of how this actually works, edge locations, cache hits and misses, invalidation, on &lt;a href="https://seeitflow.com/networking/advanced-networking/cdn/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to see it rather than read through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a CDN actually is
&lt;/h2&gt;

&lt;p&gt;A CDN is a globally distributed network of servers that caches content physically close to users, so a request gets served from somewhere nearby instead of making the full trip to a single distant origin every time. From a user's perspective, it's completely invisible, they request a URL and get a fast response back, with no idea that a network of edge locations, also called points of presence or PoPs, is sitting between them and the server that actually owns the content.&lt;/p&gt;

&lt;p&gt;The CDN's job is to answer as many requests as it possibly can directly from a cache at the edge, and only fall back to the origin when it genuinely has no other choice. Because it sits on every single request at the network's edge, it also ends up being a natural place to handle DDoS protection, TLS termination, and request routing, and increasingly even small amounts of actual compute. Almost every production CDN is wearing several of these hats simultaneously, not just the caching one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Hit or miss, and why the ratio is the metric that matters
&lt;/h2&gt;

&lt;p&gt;Every request reaching an edge location resolves to one of two outcomes. A cache hit means the requested object already exists at that edge location and hasn't expired yet, so the response gets served directly, no trip to the origin, minimal latency. A cache miss means the object isn't there, or it's expired, so the edge has to fetch it from the origin first, store a copy locally, and only then return it to the user. The very first request for anything new is always a miss by definition. Every request after that, until the cached copy expires, is a hit.&lt;/p&gt;

&lt;p&gt;Cache-hit ratio, the percentage of requests served from cache versus the total, is the single most-watched CDN health metric that exists, and for good reason. A dropping hit ratio is usually one of the earliest signals that something's gone wrong, a misconfigured cache key, a TTL set too short, a sudden spike in genuinely unique, uncacheable requests, and it tends to show up on a dashboard well before origin load becomes visibly dangerous.&lt;/p&gt;

&lt;h2&gt;
  
  
  TTL: the number that decides how fresh is fresh enough
&lt;/h2&gt;

&lt;p&gt;Time-to-live determines how long a cached object is considered fresh before the edge needs to re-check with the origin. It's usually set through HTTP response headers, most commonly &lt;code&gt;Cache-Control: max-age=&amp;lt;seconds&amp;gt;&lt;/code&gt;, or the older &lt;code&gt;Expires&lt;/code&gt; header. Once that TTL runs out, the object is considered stale, and the next request for it triggers either a fresh fetch from the origin or a conditional request using &lt;code&gt;ETag&lt;/code&gt;/&lt;code&gt;If-None-Match&lt;/code&gt; to check whether it actually changed at all.&lt;/p&gt;

&lt;p&gt;There's no universally correct TTL, it's a genuine trade-off made per asset. A short TTL means fresher content but more load reaching the origin. A long TTL means far less origin load but more risk of serving something outdated if it changes unexpectedly. A hashed, versioned JS bundle can reasonably cache for a year, since its filename changes the moment its content does. A homepage's HTML might only be safe to cache for a few seconds, or not at all. Two patterns soften this trade-off further: stale-while-revalidate serves the stale cached copy immediately while refreshing it in the background, so no single user ever waits on the refresh, and stale-if-error keeps serving the last known-good cached copy if the origin becomes unreachable or starts erroring, rather than failing the request outright just because the origin is having a bad moment.&lt;/p&gt;

&lt;h2&gt;
  
  
  When you can't wait for the TTL to run out
&lt;/h2&gt;

&lt;p&gt;Sometimes a TTL expiring naturally isn't fast enough. Cache invalidation, or purging, forcibly removes a stale object from edge caches immediately, typically triggered right after a deployment pipeline publishes new content, "purge &lt;code&gt;/app.js&lt;/code&gt; and &lt;code&gt;/styles.css&lt;/code&gt;", so users don't have to sit through a long TTL window just to see an update that already shipped.&lt;/p&gt;

&lt;p&gt;The mistake worth avoiding here is assuming a purge is instantaneous everywhere at once. Propagating a purge across hundreds of globally distributed PoPs takes measurable time, seconds at minimum, sometimes longer for wildcard purges, so there's genuinely a short window where some regions are still serving the old version while others have already updated. A common alternative sidesteps the whole problem: bake a content hash directly into the filename, &lt;code&gt;app.a1b2c3.js&lt;/code&gt;, so a new deployment automatically produces a new cache key with no purge step required at all, and the old version can safely be cached forever since nothing ever references it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Not everything deserves to be cached
&lt;/h2&gt;

&lt;p&gt;Static content, images, CSS, JS bundles, videos, fonts, is identical for every single user and only changes on deployment, which makes it ideal for long TTLs and aggressive edge caching. Dynamic content, personalized API responses, authenticated pages, real-time data, differs per user or per request, and usually bypasses the cache entirely via &lt;code&gt;Cache-Control: no-store&lt;/code&gt;, getting forwarded straight through to the origin instead.&lt;/p&gt;

&lt;p&gt;But the line between those two categories is softer than it first looks. Some dynamic-looking content is genuinely cacheable for short windows, a product listing that updates every few minutes can still benefit meaningfully from a 30-second edge cache, cutting a large share of origin load without serving anything a user would notice as stale. The more useful question for any given endpoint isn't "is this static or dynamic," it's "how stale can this be before someone actually notices or cares." That answer, not the content type on paper, is what should actually drive the TTL.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cache key decides what counts as "the same request"
&lt;/h2&gt;

&lt;p&gt;By default, most CDNs build a cache key from the request path plus, often, the full query string. Two requests differing only by &lt;code&gt;?color=red&lt;/code&gt; versus &lt;code&gt;?color=blue&lt;/code&gt; get treated as entirely separate cached objects unless you configure it otherwise. CDNs let you customize this, stripping out tracking parameters like &lt;code&gt;utm_source&lt;/code&gt; so semantically identical requests share one cache entry instead of fragmenting into thousands of near-duplicate ones, or including specific headers via the &lt;code&gt;Vary&lt;/code&gt; header when a response genuinely differs based on something like &lt;code&gt;Accept-Language&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is also where one of the more serious CDN mistakes lives. Caching a response that varies by &lt;code&gt;Authorization&lt;/code&gt; or a session cookie, without including that in the cache key or declaring it in &lt;code&gt;Vary&lt;/code&gt;, can leak one user's personalized or authenticated response straight to a completely different user who happens to request the same URL next. It's the single most damaging category of CDN misconfiguration in production, not because it's exotic, but because it's extremely easy to miss during a routine setup and invisible until someone notices they're looking at a stranger's data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping the origin from getting hit all at once
&lt;/h2&gt;

&lt;p&gt;Without any special handling, every regional edge location independently experiences a cache miss the first time a newly popular object shows up, and all of them hit the origin simultaneously, a thundering herd from the origin's point of view even though no single PoP did anything wrong. Origin shielding designates one edge location as the only one allowed to talk directly to the origin. Every other PoP routes its misses through that shield location instead, which fetches and caches the object exactly once and then serves every other region from there. It's one of the lowest-effort, highest-impact changes available for any origin that isn't trivially horizontally scalable, especially during a cold-cache event like a fresh deploy or a sudden viral spike.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security that happens before traffic ever reaches you
&lt;/h2&gt;

&lt;p&gt;Sitting on every single request makes a CDN a natural security enforcement point, not just a performance layer bolted on top. TLS termination happens right at the edge, closest to the user, cutting handshake latency compared to a distant origin doing the same negotiation itself. A Web Application Firewall can filter known malicious patterns, SQL injection attempts, XSS payloads, known bad bot signatures, before any of it reaches an actual application server. Signed URLs restrict access to protected content, paid media, private downloads, using time-limited tokens validated right at the edge, with no need to hit the origin just to authorize each individual request.&lt;/p&gt;

&lt;p&gt;Because a CDN spreads traffic across every PoP globally via anycast, a volumetric DDoS attack gets naturally distributed across the whole network rather than concentrating entirely on one target. Rate limiting at the edge blocks abusive clients before they ever reach the origin, and challenge pages or JS-based bot mitigation can filter automated attacks while letting legitimate users straight through. None of this matters much, though, if the origin's real IP address is still publicly reachable, an attacker who can hit the origin directly bypasses every one of these protections entirely, which is why locking the origin down to only accept traffic from the CDN's own IP ranges is such a disproportionately valuable, low-effort step.&lt;/p&gt;

&lt;h2&gt;
  
  
  A CDN isn't a load balancer, even though they overlap
&lt;/h2&gt;

&lt;p&gt;These two get confused constantly because they both sit somewhere in the request path distributing traffic, but they're solving genuinely different problems. A CDN sits between the user and the entire origin infrastructure, deciding whether a request even needs to reach the origin at all. A load balancer sits in front of a specific pool of backend servers, typically at or near the origin itself, deciding which one of those servers should actually handle a request that does make it through.&lt;/p&gt;

&lt;p&gt;They compose rather than compete. A fairly typical architecture looks like user, to CDN edge, to load balancer, to application servers, the CDN absorbs everything cacheable before it ever becomes the load balancer's problem, and the load balancer distributes whatever's left, the genuinely dynamic, uncacheable remainder, across the application fleet. A CDN can't replace a load balancer here, it has no way to meaningfully distribute per-request, uncacheable dynamic traffic across a server pool, that's precisely the job a load balancer exists to do.&lt;/p&gt;

&lt;h2&gt;
  
  
  A CDN is really a reverse proxy, scaled globally
&lt;/h2&gt;

&lt;p&gt;The cleanest way to understand a CDN's actual mechanism is as a globally distributed, managed reverse proxy. A reverse proxy like Nginx or Varnish sits in front of an origin, forwarding and optionally caching requests, usually from one location, close to or inside the origin's own infrastructure. A CDN takes that exact same forwarding-and-caching idea and spreads it across hundreds of geographically dispersed edge locations instead of one, adding global routing, DDoS absorption, and origin protection at a scale a single reverse proxy was never meant to operate at. Many CDN vendors are, quite literally, running reverse-proxy software, or a heavily modified fork of one, at every single PoP. The products differ in scale and geographic distribution, not in the underlying mechanism.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this goes wrong in production
&lt;/h2&gt;

&lt;p&gt;A handful of mistakes account for most of the real incidents once a CDN is actually running in front of production traffic. Treating it as fire-and-forget, enabling it with default settings and never revisiting cache-control headers, TTLs, or cache keys as the application evolves, is probably the most common one, caching policy that lives entirely in a dashboard rather than version-controlled code is invisible during code review and easy to silently break. Caching a personalized response by accident, forgetting to mark an authenticated page as &lt;code&gt;private&lt;/code&gt; or &lt;code&gt;no-store&lt;/code&gt;, causes exactly the user-data-leak scenario described above. Not locking down the origin leaves its real IP reachable and every CDN-level protection trivially bypassable. And ignoring cache-hit ratio until an actual incident forces attention onto it means a caching regression only gets discovered during a traffic spike, instead of being caught early through routine monitoring that was watching for it all along.&lt;/p&gt;

&lt;p&gt;One more worth calling out specifically because the naming is actively misleading: &lt;code&gt;no-cache&lt;/code&gt; does not mean "don't cache." It means "cache this, but always revalidate with the origin before serving it." &lt;code&gt;no-store&lt;/code&gt; is the directive that actually forbids caching outright. Mixing these two up is a genuinely common production gotcha, not just an interview trick question.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question worth asking about any caching strategy
&lt;/h2&gt;

&lt;p&gt;A useful test for whether a caching setup is actually protecting the origin, rather than just happening to look fast most of the time: what happens to the origin if the CDN's cache were completely cold right now, this second? If the honest answer is "it falls over," the caching strategy is under-protecting the origin no matter how good the measured latency numbers currently look. Treating origin protection as an explicit design goal, not an incidental side effect of caching for speed, changes real decisions, it's what justifies micro-caching a dynamic endpoint for even a few seconds, or paying for origin shielding, even at the cost of a little extra staleness, because the alternative is an origin that can't survive its own traffic the moment the cache goes cold.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this is heading
&lt;/h2&gt;

&lt;p&gt;The newest layer on top of all of this is edge computing, running actual application logic at the same distributed locations that used to only cache static files. Edge functions can handle request rewriting, A/B test bucketing, authentication checks, even full server-side rendering, physically close to the user, cutting out round trips to the origin for logic that never needed the full backend stack in the first place. This is steadily blurring the line between "CDN" and "application platform," and it's worth knowing the trend exists even if a given system isn't using it yet, since it comes up constantly in forward-looking system design conversations.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model worth keeping
&lt;/h2&gt;

&lt;p&gt;A CDN caches content at globally distributed edge locations for two reasons at once, cutting latency by physically moving content closer to users, and shielding the origin from load it was never provisioned to handle directly. Every request resolves to a hit or a miss, TTL controls how long something stays fresh before that distinction gets re-evaluated, and invalidation exists for the moments when waiting out a TTL isn't fast enough. Not everything belongs in a cache, and the real question for any given piece of content is how stale it can get away with being, not whether it's technically static or dynamic. And the cache key, quietly, is deciding what counts as "the same request" the entire time, get it wrong and you're either wasting cache space or handing one user's data to another.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;This post covers the core caching mechanics and production trade-offs of a CDN. The full guide on &lt;a href="https://seeitflow.com/networking/advanced-networking/cdn/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; covers the fundamentals in more depth across edge locations, TTL, and cache keys, there's a dedicated &lt;a href="https://seeitflow.com/networking/advanced-networking/cdn/production-engineering" rel="noopener noreferrer"&gt;production engineering guide&lt;/a&gt; covering cache strategy, origin shielding, DDoS protection, and a full debugging checklist, and an &lt;a href="https://seeitflow.com/networking/advanced-networking/cdn/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on CDN versus load balancer, CDN versus reverse proxy, and the real trade-offs behind invalidation and origin protection.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>cdn</category>
      <category>systemdesign</category>
      <category>webdev</category>
    </item>
    <item>
      <title>A Reverse Proxy Isn't One Feature. It's Five Cross-Cutting Concerns Living in One Place.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Wed, 30 Sep 2026 14:54:19 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/a-reverse-proxy-isnt-one-feature-its-five-cross-cutting-concerns-living-in-one-place-1had</link>
      <guid>https://dev.to/mangeshmandlik/a-reverse-proxy-isnt-one-feature-its-five-cross-cutting-concerns-living-in-one-place-1had</guid>
      <description>&lt;p&gt;Start with the simplest possible setup: a client talking directly to a single application server. No proxy, no extra hop, just a browser hitting one address and getting an answer back. That server ends up doing a surprising amount on its own once you look closely, it's terminating its own TLS certificate, serving its own static files like images and stylesheets alongside its actual application logic, and it's directly exposed to the internet with no layer in front of it filtering anything out.&lt;/p&gt;

&lt;p&gt;Add a second server and every one of those responsibilities gets duplicated. Two TLS certificates to renew instead of one, two copies of the same static-file-serving logic, two separate attack surfaces facing the internet. Add an admin panel, a public API, and a static marketing site, and now routing logic starts getting tangled into the application itself, or worse, hardcoded into the clients that call it. None of these problems are hard individually. They're expensive specifically because they're the same problem, repeated once per server, forever. I put together a visual walkthrough of exactly this evolution, from direct access to a full proxy tier, on &lt;a href="https://seeitflow.com/networking/traffic-routing/reverse-proxy/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, building it up the way I wish someone had explained it to me, one problem at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a reverse proxy actually is
&lt;/h2&gt;

&lt;p&gt;A reverse proxy is a server that sits in front of your application servers, receives every client request on their behalf, and forwards it to whichever backend should actually handle it, then relays the response back. From the client's point of view, the reverse proxy &lt;em&gt;is&lt;/em&gt; the application. The browser connects to one address, sends a request, gets a response back, and has no idea whether there's one app server behind that address or fifty, whether the response came from a cache, or whether a file was served straight off the proxy's own disk without an application ever getting involved. That indirection is the entire point.&lt;/p&gt;

&lt;p&gt;Because the proxy sits on every single request without exception, it becomes the natural home for exactly the concerns that shouldn't live in individual applications: terminating HTTPS, serving static files, routing by URL, caching responses, and filtering out hostile traffic. Pull those five things out of the app and into one shared edge layer, and every backend server gets simpler for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two proxies, opposite directions
&lt;/h2&gt;

&lt;p&gt;The word "proxy" covers two roles that face in opposite directions, and mixing them up is an easy mistake to make once, since the underlying mechanics look similar.&lt;/p&gt;

&lt;p&gt;A forward proxy sits in front of a group of clients. When those clients want to reach the internet, their requests go through the forward proxy first, which can filter, log, cache, or anonymize that outbound traffic. A corporate web filter or a VPN's egress point is a forward proxy, it works on behalf of the client, and the destination server on the other end usually has no idea the proxy was ever involved.&lt;/p&gt;

&lt;p&gt;A reverse proxy sits in front of a group of servers instead. When clients out on the internet want to reach your application, they hit the reverse proxy first, which forwards the request to the right backend. It works on behalf of the server, and the client usually has no idea it's there. Same basic machinery, pointed in the opposite direction, serving the opposite side's interests.&lt;/p&gt;

&lt;h2&gt;
  
  
  It overlaps with a load balancer without being one
&lt;/h2&gt;

&lt;p&gt;A reverse proxy can spread traffic across backends, and that shared capability is exactly why people conflate it with a dedicated load balancer. But load balancing is one job among several for a reverse proxy, and it's the &lt;em&gt;only&lt;/em&gt; job for a dedicated load balancer.&lt;/p&gt;

&lt;p&gt;A dedicated load balancer is deliberately narrow: distribute requests across identical replicas of a service, health-check them, do all of that at very high throughput, typically operating at Layer 4, connections, or a fairly basic Layer 7, with almost no understanding of the application itself. Its narrowness is what lets it be extremely fast and extremely reliable. A reverse proxy load-balancing across replicas is doing that same distribution, but as one feature bundled alongside HTTPS termination, response caching, static file serving, and path-based routing, all running on the same box. In most real production stacks, you actually run both, a dedicated L4 load balancer in front of a cluster of L7 reverse proxies, letting each layer stay narrowly focused and scale independently of the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  It's also not the same thing as an API gateway
&lt;/h2&gt;

&lt;p&gt;An API gateway can be thought of as a reverse proxy that specialized specifically for APIs and cross-cutting policy. A reverse proxy routes requests to backends, often identical replicas of the same application, and handles edge concerns like SSL, caching, and static files. It's content-aware enough to route by path, but it doesn't inherently understand your API's identity or enforce business-level rules.&lt;/p&gt;

&lt;p&gt;An API gateway routes fundamentally &lt;em&gt;different&lt;/em&gt; requests to &lt;em&gt;different&lt;/em&gt; services, and layers on API-specific policy on top of that routing: authentication, authorization, per-key rate limiting, request aggregation, versioning. Under the hood a gateway is still built on reverse-proxy technology, but its center of gravity is policy and coordinating many distinct services, not edge plumbing for one application. Nginx, for what it's worth, can genuinely wear either hat, the question in any given deployment is which one it's actually configured to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happens to one request, in order
&lt;/h2&gt;

&lt;p&gt;A single request through a reverse proxy is really two independent connections and four distinct hops, with the proxy sitting as the pivot in the middle. The client opens a connection to the proxy and sends its request, that's hop one. The proxy then opens its &lt;em&gt;own, separate&lt;/em&gt; connection to a backend and forwards the request along, hop two. The backend does the actual work and replies to the proxy, hop three. The proxy relays that response back down the original client-side connection, hop four.&lt;/p&gt;

&lt;p&gt;That two-connection design is what unlocks essentially everything else a reverse proxy does. Because the client-side and backend-side connections are genuinely independent of each other, the proxy can terminate TLS on the client side while speaking plain HTTP to the backend, pool and reuse backend connections across many different clients, buffer a slow client without tying up a backend worker waiting on them, and even swap which backend it's talking to entirely without the client ever needing to reconnect. It's a full intermediary sitting on both halves of the exchange, not a redirect that steps out of the way after the first hop.&lt;/p&gt;

&lt;h2&gt;
  
  
  Terminating TLS once instead of everywhere
&lt;/h2&gt;

&lt;p&gt;HTTPS has to be decrypted somewhere, and doing that once at the proxy instead of separately on every backend server is one of the largest operational wins a reverse proxy provides. With SSL termination, also called TLS offload, the client negotiates HTTPS with the proxy and only the proxy. The certificate and private key live in exactly one place. From the proxy inward, on your trusted private network, traffic runs as plain HTTP, so backend servers never touch a certificate and never pay the CPU cost of encryption themselves.&lt;/p&gt;

&lt;p&gt;The operational payoff compounds quickly: one certificate to install, monitor, and rotate instead of one per server, a new backend inheriting HTTPS automatically the moment it's added behind the proxy, and upgrading cipher suites or TLS versions becoming a single config change instead of a fleet-wide rollout. The trade-off worth knowing: the leg from proxy to backend really is unencrypted plain HTTP. That's a completely reasonable choice on a trusted private network, but in zero-trust or regulated environments, that hop typically gets re-encrypted too, sometimes called TLS re-encryption or end-to-end TLS, via mTLS between the proxy and the app. You keep the centralized-certificate win on the public side while still encrypting the internal hop.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;server&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kn"&gt;listen&lt;/span&gt; &lt;span class="mi"&gt;443&lt;/span&gt; &lt;span class="s"&gt;ssl&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kn"&gt;server_name&lt;/span&gt; &lt;span class="s"&gt;example.com&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kn"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://app_servers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;Host&lt;/span&gt; &lt;span class="nv"&gt;$host&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_set_header&lt;/span&gt; &lt;span class="s"&gt;X-Real-IP&lt;/span&gt; &lt;span class="nv"&gt;$remote_addr&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Serving static files without waking up the application
&lt;/h2&gt;

&lt;p&gt;A large fraction of web traffic is static files that never change per request, a logo, a compiled JS bundle, a stylesheet. Making a full application runtime generate those on every request is genuinely wasteful. When a request for &lt;code&gt;/logo.png&lt;/code&gt; arrives, the proxy can match it against its static rules, find the file directly on its own disk, and return it without ever contacting the application server at all.&lt;/p&gt;

&lt;p&gt;A compiled proxy like Nginx serving a file straight from disk, often using &lt;code&gt;sendfile&lt;/code&gt; so the kernel copies bytes from the file cache to the socket with zero application involvement, is dramatically faster and cheaper than a language runtime doing the equivalent work. Splitting static traffic off from dynamic traffic at the proxy is frequently the single biggest reason teams put something like Nginx in front of an application in the first place, all that repetitive traffic gets absorbed at the edge, leaving the application's CPU free for the dynamic work only it can actually compute. Fingerprinted filenames, &lt;code&gt;app.4f2c.js&lt;/code&gt; rather than &lt;code&gt;app.js&lt;/code&gt;, let you set very long cache TTLs without any risk of serving stale content, since a new deploy simply produces a new filename.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/static/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kn"&gt;root&lt;/span&gt; &lt;span class="n"&gt;/var/www&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kn"&gt;expires&lt;/span&gt; &lt;span class="s"&gt;30d&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://app_servers&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Routing one domain to several independent services
&lt;/h2&gt;

&lt;p&gt;A real site is rarely a single application. The proxy can front several independent services behind one domain and route each request based on its URL path, &lt;code&gt;/api/*&lt;/code&gt; to the API service, &lt;code&gt;/admin&lt;/code&gt; to the admin panel, &lt;code&gt;/static/*&lt;/code&gt; to the asset service. To the outside world it looks like one coherent site. Behind the proxy, it's several genuinely independent services that different teams can deploy, scale, and own separately.&lt;/p&gt;

&lt;p&gt;This is exactly the point where a reverse proxy starts to resemble an API gateway, the difference is that the proxy is routing purely by host and path, without layering on authentication, rate limiting, or aggregation policy on top. Nginx resolves each request to exactly one location block, matched by a defined precedence, exact matches first, then longest-prefix, then regex in file order. Get that precedence wrong and a broad &lt;code&gt;location /&lt;/code&gt; placed above a specific &lt;code&gt;/api&lt;/code&gt; rule will silently swallow API traffic that was supposed to go somewhere else, a routing bug that stays completely invisible until a user ends up on the wrong backend.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/api/&lt;/span&gt;   &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://api_service&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/admin&lt;/span&gt;  &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://admin_service&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/static&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://asset_service&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Answering a request without waking the backend at all
&lt;/h2&gt;

&lt;p&gt;If ten thousand people ask for the same page within a minute, regenerating it ten thousand times is pure waste. On a cache miss, the proxy has nothing stored yet, so it forwards to the backend and then keeps a copy of the response, tagged with a time-to-live, before passing it along. On a subsequent request for the same thing, while that copy is still fresh, the proxy can answer immediately and the backend never gets touched.&lt;/p&gt;

&lt;p&gt;Hit ratio is the entire game here. At a 90% hit ratio, your backend sees roughly a tenth of the traffic it otherwise would, and users get responses back almost instantly. Request collapsing matters too: when a cached item finally does expire, a thundering herd of simultaneous requests for it should trigger exactly one backend fetch to repopulate the cache, not thousands of simultaneous ones. The classic mistake here is caching something that was never safe to share in the first place, per-user or rapidly changing content, under a key that doesn't account for who's actually asking. Caching a logged-in user's page under a shared key doesn't just fail to help, it leaks that user's data to the next visitor who happens to request the same URL.&lt;/p&gt;

&lt;h2&gt;
  
  
  The perimeter your backends never see
&lt;/h2&gt;

&lt;p&gt;Because the proxy is the only thing actually exposed to the internet, it's the natural place to defend the entire system at once. Backends sit on a private network with no public address, so nobody outside can even attempt to reach them directly, only the proxy can. Hiding the internal topology genuinely matters here, an attacker can't target a box they can't see or address in the first place.&lt;/p&gt;

&lt;p&gt;Because every request flows through this one choke point, it's also the natural place to inspect and reject bad traffic before it ever reaches an application that would otherwise have to parse it. TLS with modern cipher suites, security headers like HSTS and CSP, IP allow or deny lists, connection and request rate limiting, request size and timeout limits, and optionally a WAF module for known exploit patterns, all of it layers up at this one point. Blocking at the edge shrinks your actual blast radius meaningfully: the application never even parses hostile input, because the request already died at the proxy before reaching it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The proxy itself can't be the single point of failure
&lt;/h2&gt;

&lt;p&gt;A single proxy sitting on every request makes that one process a single point of failure for the entire site. If it dies, SSL, routing, and caching all go down with it, total outage, not a partial degradation. The fix is running several identical proxy instances as a stateless cluster and placing a dedicated load balancer in front of them, health-checking each instance individually. If one proxy fails, the load balancer simply stops routing to it, and users never notice anything happened.&lt;/p&gt;

&lt;p&gt;Each layer in this stack scales independently and horizontally. Under load, you add proxy instances for more edge capacity; backends scale on their own separate axis for more compute. Because proxies absorb static, cached, and blocked traffic before it ever reaches an application, the backend tier that actually needs scaling is often just the genuinely dynamic remainder of the traffic, which tends to be a lot smaller than the total request volume hitting the edge.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this looks like assembled together
&lt;/h2&gt;

&lt;p&gt;Put every piece in one place and you get the shape behind most large-scale web systems: traffic from the internet hits a dedicated load balancer, which spreads it across a resilient tier of reverse proxies. That proxy tier terminates SSL, serves static files, caches responses, routes by path, and filters hostile traffic, then forwards whatever dynamic requests actually survive all of that to the backend services behind it. Every layer scales independently, and no single layer is a point of failure for the whole system.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Internet → Load Balancer → Reverse Proxy Cluster → Backend Services
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The handful of mistakes that take sites down repeatedly
&lt;/h2&gt;

&lt;p&gt;A small set of reverse-proxy mistakes account for a disproportionate share of real outages. Running a single proxy instance with no redundancy behind it is the most basic one, one failure away from a total outage, always run a cluster of at least two behind a load balancer. A catch-all &lt;code&gt;location /&lt;/code&gt; placed above more specific routes will silently swallow traffic meant for those specific routes, order rules from specific to general and prefer exact or longest-prefix matches for anything critical. Caching an authenticated response under a shared cache key leaks one user's private data to whoever else requests that same URL next, vary the cache key on session, or simply mark authenticated responses as non-cacheable entirely. Forgetting certificate renewal turns into a hard, total HTTPS outage the moment the cert expires, browsers refuse the handshake outright, so automate renewal and alert well before expiry rather than relying on someone remembering. Dropping the client's real IP is a quieter mistake, backends end up seeing the proxy's IP address for every single request unless &lt;code&gt;X-Forwarded-For&lt;/code&gt; and &lt;code&gt;X-Real-IP&lt;/code&gt; are explicitly set and, just as importantly, trusted carefully so they can't be spoofed by a malicious client. And reloading a proxy's configuration without validating it first means a single bad config gets pushed to every instance in the cluster simultaneously, &lt;code&gt;nginx -t&lt;/code&gt; before every reload and canarying config changes catches this before it becomes an outage rather than after.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model worth keeping
&lt;/h2&gt;

&lt;p&gt;Reverse proxy, load balancer, and API gateway all sit in a similar physical position in the architecture, and it's tempting to treat them as interchangeable because of that. They're not answering the same question. A load balancer asks which replica should handle this. An API gateway asks which service should handle this, and whether the request is even allowed to happen. A reverse proxy is broader than either of those framings alone, it's the shared edge layer absorbing SSL termination, static serving, path-based routing, caching, and security, so that none of the backend services sitting behind it have to reimplement any of those five concerns on their own.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;This post covers the core concepts and request lifecycle of a reverse proxy. The full guide on &lt;a href="https://seeitflow.com/networking/traffic-routing/reverse-proxy/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; covers the fundamentals in more depth across 13 chapters, there's a dedicated &lt;a href="https://seeitflow.com/networking/traffic-routing/reverse-proxy/production-engineering" rel="noopener noreferrer"&gt;production engineering guide&lt;/a&gt; covering forwarding internals, TLS termination flow, and a full failure-and-scale playbook, and an &lt;a href="https://seeitflow.com/networking/traffic-routing/reverse-proxy/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on where to cache, where to terminate TLS, and the comparisons against load balancers and API gateways in more detail.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>backend</category>
      <category>systemdesign</category>
      <category>webdev</category>
    </item>
    <item>
      <title>An API Gateway Isn't a Router. It's Every Cross-Cutting Concern Your Services Would Otherwise Duplicate.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:37:31 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/an-api-gateway-isnt-a-router-its-every-cross-cutting-concern-your-services-would-otherwise-414b</link>
      <guid>https://dev.to/mangeshmandlik/an-api-gateway-isnt-a-router-its-every-cross-cutting-concern-your-services-would-otherwise-414b</guid>
      <description>&lt;p&gt;With two services in a system, letting clients call each of them directly feels completely reasonable, there's not much to get wrong about a client hitting the users service here and the orders service there. Then the system grows. A payments service shows up, then search, then several different client applications, each needing authentication, staying under a rate limit, targeting the right API version, and all of it ideally traceable when something goes wrong. Instances of every one of those services are also constantly scaling up and down, so their actual addresses keep changing underneath everything.&lt;/p&gt;

&lt;p&gt;At that point, "which server should receive this request" stops being the real question. The real question becomes which &lt;em&gt;service&lt;/em&gt; should receive this request, whether the caller is even allowed to make it, whether they're within their quota, and how you're supposed to observe the whole journey once it's scattered across half a dozen services. That's the specific problem an API gateway exists to solve. I put together a visual walkthrough of the request flow through a gateway, routing, auth, rate limiting, the whole pipeline, on &lt;a href="https://seeitflow.com/networking/traffic-routing/api-gateway/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to see it step by step rather than read it linearly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The basic shift: one door instead of many
&lt;/h2&gt;

&lt;p&gt;Instead of exposing every backend service directly to clients:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client ──→ Users Service
       ──→ Orders Service
       ──→ Payments Service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;a gateway sits in front of all of them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         ┌──→ Users Service
Client ──→ API Gateway ──┼──→ Orders Service
                         └──→ Payments Service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client now sees exactly one API endpoint. The gateway is the thing that actually knows the internal topology, so a request like &lt;code&gt;GET /orders/42&lt;/code&gt; gets routed internally to wherever the orders service actually lives, without the client ever needing to know or care.&lt;/p&gt;

&lt;p&gt;Routing alone would already be a reasonable thing to centralize, but it's really just the entry point into a bigger idea. Because literally every external request has to pass through this one layer, it becomes the natural place to put anything that would otherwise need to be reimplemented, slightly differently, inside every single service: authentication and authorization, rate limiting, API versioning, service discovery, request aggregation, tracing and metrics. Without a gateway, each service tends to end up with its own slightly-off version of the same edge logic, and "slightly off" is exactly where inconsistent security behavior creeps in.&lt;/p&gt;

&lt;h2&gt;
  
  
  This isn't the same job as a load balancer
&lt;/h2&gt;

&lt;p&gt;These two get confused constantly because they both physically sit between clients and backend systems, but they're answering genuinely different questions. A load balancer asks which &lt;em&gt;replica&lt;/em&gt; of a given service should handle this request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              ┌──→ Orders #1
Client → LB ──┼──→ Orders #2
              └──→ Orders #3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every box there is running the same service. An API gateway asks a different question entirely, which &lt;em&gt;service&lt;/em&gt; should handle this, and what policy needs to apply before it even gets there:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                         /users/*   ──→ Users
Client → API Gateway ─── /orders/*  ──→ Orders
                         /payments/*──→ Payments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In practice these two usually work together rather than substituting for each other:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Clients
   │
   ▼
Load Balancer
   │
   ▼
API Gateway Cluster
   │
   ├──→ Users
   ├──→ Orders
   └──→ Payments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The load balancer's job is spreading traffic across a pool of interchangeable gateway instances. The gateways then do the actual application-aware routing and policy enforcement on top of that.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happens to one request
&lt;/h2&gt;

&lt;p&gt;Thinking of the gateway as a pipeline, rather than a single box that "does gateway stuff," makes its behavior much easier to reason about. A request typically moves through something like this sequence: terminate TLS and parse the HTTP request, assign or propagate a trace ID, match the route, authenticate the caller, authorize the caller, apply rate limits, transform or aggregate the request or response if needed, discover a currently-healthy backend instance, forward the request, and finally record metrics before returning the response.&lt;/p&gt;

&lt;p&gt;The order here isn't arbitrary. Authentication has to happen before authorization, since you can't decide what someone's allowed to do before you know who they are. Rate limiting needs to happen before any expensive downstream work, otherwise you're doing the expensive work first and only then deciding it shouldn't have been allowed. And the trace ID needs to be assigned near the very beginning, so every later stage in the pipeline, and every downstream service the request eventually touches, can be tied back to the same originating request. This ordering is exactly why an API gateway is better understood as a structured, sequential pipeline rather than a loose bag of unrelated features that happen to live in one process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why authentication sits early in that pipeline
&lt;/h2&gt;

&lt;p&gt;Say a request arrives carrying a JWT. The gateway can verify its signature, check its expiry, confirm the issuer and audience, all before the request ever reaches a backend service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  │
  │ Bearer token
  ▼
Gateway
  │
  ├── invalid token ──→ 401
  │
  └── valid token
          │
          ▼
       Service
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Authentication is answering "who are you." Authorization is a separate question, "are you allowed to do this specific thing." A logged-in user attempting an admin-only action and someone presenting an outright invalid token are genuinely different situations, one gets rejected before their identity is even established, the other gets identified successfully and then denied for a completely different reason. Centralizing both checks in the gateway also means individual backend services don't each need their own slightly divergent implementation of the same edge policy, which is exactly the kind of duplication that quietly drifts out of sync over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Rate limiting gets genuinely interesting once you scale the gateway itself
&lt;/h2&gt;

&lt;p&gt;A token bucket is a common approach: each client gets a bucket that refills at some configured rate, requests consume tokens from it, and once the bucket's empty, the gateway responds with &lt;code&gt;429 Too Many Requests&lt;/code&gt;. Straightforward, right up until there's more than one gateway instance.&lt;/p&gt;

&lt;p&gt;Picture a limit of 100 requests per minute enforced across three separate gateway instances:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;           ┌── Gateway 1 → counter = 100
Client ────┼── Gateway 2 → counter = 100
           └── Gateway 3 → counter = 100
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If each gateway keeps its own independent counter, a client hitting all three effectively gets 300 requests per minute instead of 100, three separate quotas instead of one shared one. The fix is coordinating that rate-limit state somewhere all the gateway instances can see it, commonly Redis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gateway 1 ─┐
Gateway 2 ─┼──→ Shared rate-limit state
Gateway 3 ─┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of those details that looks completely trivial on a whiteboard, "just add a rate limiter", and turns out to be considerably more interesting the moment you actually run more than one instance of the thing enforcing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Services don't sit at fixed addresses
&lt;/h2&gt;

&lt;p&gt;Hardcoding something like &lt;code&gt;orders-service = 10.0.4.17&lt;/code&gt; works for approximately as long as nothing autoscales. Instances come up, go down, fail health checks, and generally move around constantly in any system that scales dynamically. Instead, the gateway can rely on service discovery:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gateway
   │
   ▼
Service Registry
   │
   ├── orders-1 ✓
   ├── orders-2 ✓
   └── orders-3 ✗
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The registry tells the gateway which instances are currently actually healthy, which means a brand-new instance can start receiving traffic the moment it's registered, with no client change needed and no redeploy of the gateway itself carrying a new hardcoded address.&lt;/p&gt;

&lt;h2&gt;
  
  
  Aggregation trades client round trips for gateway complexity
&lt;/h2&gt;

&lt;p&gt;Say a mobile home screen needs a user's profile, their recent orders, and their current payment status, three genuinely separate pieces of data from three separate services. The client could fire off three separate requests. Or the gateway could expose one endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /home
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and fan that single request out internally:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                  ┌──→ Users
Client → Gateway ─┼──→ Orders
                  └──→ Payments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;running those backend calls in parallel and combining the results into one response. That reduces the number of round trips the client makes, which matters more than it sounds like it should on higher-latency mobile networks. The trade-off is real, though: push too much aggregation and response-shaping logic into the gateway and it slowly turns into a place genuinely full of business logic, which is exactly the kind of responsibility a shared, cross-cutting layer shouldn't be accumulating. For anything beyond simple fan-out, a dedicated Backend-for-Frontend tends to be a cleaner home for that complexity than continually teaching the shared gateway more application-specific behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gateway can't be a single point of failure
&lt;/h2&gt;

&lt;p&gt;Putting one single gateway process in front of everything creates an obvious problem:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Clients → ONE Gateway → Everything
              💥
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that one process dies, the entire API disappears at once, every service behind it becomes unreachable regardless of how healthy they individually are. A real production setup looks more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ┌── Gateway 1
Clients → LB ───────┼── Gateway 2
                    └── Gateway 3
                           │
             ┌─────────────┼─────────────┐
             ▼             ▼             ▼
           Users         Orders       Payments
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gateway instances should generally be stateless and fully interchangeable with each other. Routes, keys, and policy configuration get pulled from a shared control plane rather than baked into any one instance, and any runtime state that genuinely needs to be coordinated across instances, distributed rate limits being the obvious example from earlier, lives outside the individual gateway processes entirely. With that in place, any single gateway instance can disappear and the load balancer in front of the cluster simply routes around it, no different in principle from losing one replica of any other stateless service.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gateway is also a natural place to see the whole request
&lt;/h2&gt;

&lt;p&gt;Because every external request enters through this one layer, it's a genuinely good place to originate distributed tracing rather than trying to stitch it together afterward from scattered logs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  │
  ▼
Gateway  trace=8f2a
  │
  ├──→ Users   trace=8f2a
  └──→ Orders  trace=8f2a
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same trace ID follows the request through every downstream call it triggers, so instead of separately eyeballing unrelated log lines from several different services and trying to guess which ones belong to the same user action, you can reconstruct the entire request as one coherent story. The gateway is also well positioned to expose useful per-route metrics directly, request rate, error rate, latency percentiles, broken down by route rather than averaged across the whole system.&lt;/p&gt;

&lt;h2&gt;
  
  
  None of this is free
&lt;/h2&gt;

&lt;p&gt;Every request now takes an extra hop through the gateway before it reaches the service that actually handles it. That's added latency, one more system that has to be operated and kept healthy, and one more thing sitting in the critical path that the whole API now depends on. For a small system with a couple of services, calling them directly may genuinely be simpler and entirely sufficient, there's no rule saying every architecture needs a gateway on principle.&lt;/p&gt;

&lt;p&gt;The gateway earns its cost as the number of services, clients, and shared policies grows, specifically once those things start needing to be enforced consistently across many services rather than once. The more useful framing isn't "should every microservice architecture have an API gateway." It's whether the cost of duplicating routing, authentication, rate limiting, versioning, and observability across every individual service has become more expensive than the cost of operating one shared gateway layer that does all of it consistently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model worth keeping
&lt;/h2&gt;

&lt;p&gt;Three separate layers are answering three separate questions here, and keeping them separate is what keeps the whole architecture reasoning-friendly rather than tangled. The load balancer is asking which replica should handle this. The API gateway is asking which service should handle this, and whether the request is even allowed to happen. And the individual service is asking what business operation should actually occur now that it's arrived. Once those three questions stop being conflated into one blurry "handle the request" step, a surprising amount of what makes distributed systems hard to reason about gets a lot more tractable.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;This post covers the core request pipeline and production concerns of an API gateway. The full visual walkthrough on &lt;a href="https://seeitflow.com/networking/traffic-routing/api-gateway/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; covers the fundamentals in more depth, there's a dedicated &lt;a href="https://seeitflow.com/networking/traffic-routing/api-gateway/production-engineering" rel="noopener noreferrer"&gt;production engineering guide&lt;/a&gt; covering gateway clustering, distributed rate limiting, and service discovery in production, and an &lt;a href="https://seeitflow.com/networking/traffic-routing/api-gateway/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on failure modes and debugging a gateway layer under real traffic.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>backend</category>
      <category>systemdesign</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Load Balancing Isn't Round Robin. It's the Control Point for Everything Behind It.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:27:59 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/load-balancing-isnt-round-robin-its-the-control-point-for-everything-behind-it-4dd</link>
      <guid>https://dev.to/mangeshmandlik/load-balancing-isnt-round-robin-its-the-control-point-for-everything-behind-it-4dd</guid>
      <description>&lt;p&gt;Put three application servers behind a load balancer and the architecture looks almost too simple to write an article about:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 ┌── Server A
Client → LB ─────┼── Server B
                 └── Server C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A request arrives, the load balancer picks a server, done. That's genuinely the least interesting part of what a production load balancer does, though. The real questions start right after: is Server B actually healthy, not just running? What happens to the 300 requests already in flight on Server C when it gets pulled out for a deploy? Should &lt;code&gt;/payments&lt;/code&gt; and &lt;code&gt;/images&lt;/code&gt; even go to the same pool of servers? What happens when one request takes 20 milliseconds and the next takes 20 seconds? Where does TLS actually terminate? What happens to a user's session when the server holding it in memory isn't the one that answers their next request?&lt;/p&gt;

&lt;p&gt;Once those questions show up, "distribute requests across servers" stops being an adequate description of what's happening. I put together a visual walkthrough of the request flow, the algorithms, and the production failure modes on &lt;a href="https://seeitflow.com/networking/traffic-routing/loadbalancer/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to see it rather than read through it. Here's the written version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this exists in the first place
&lt;/h2&gt;

&lt;p&gt;Start with an application running on a single server. Eventually it hits a ceiling, CPU, memory, open connections, network bandwidth, doesn't matter which, you can only buy a bigger machine for so long before vertical scaling stops being an option. So you add more servers. And the moment you have more than one, a new question appears that didn't exist before: who decides which server gets each request? That's the first job a load balancer does.&lt;/p&gt;

&lt;p&gt;But horizontal scaling isn't the only reason it exists. Say Server B crashes outright. Without anything watching for that, some clients keep sending traffic to a server that's already dead. With health-aware load balancing, B simply stops receiving requests the moment it's detected as unhealthy, and traffic keeps flowing to A and C without anyone noticing B was ever a problem. So a load balancer is really solving two separate problems at once: throughput, spreading load across more capacity than one machine has, and availability, making sure a dead or struggling server doesn't keep receiving traffic just because nobody told it to stop.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much should the load balancer actually understand?
&lt;/h2&gt;

&lt;p&gt;One of the more consequential decisions in load balancer design is where it sits in the network stack, and how much of the traffic it's actually allowed to see.&lt;/p&gt;

&lt;p&gt;A Layer 4 load balancer works purely with connections, source IP, destination IP, source port, destination port, TCP or UDP. It has no idea what HTTP is, and it doesn't need to, which is exactly what makes it fast and protocol-agnostic. A Layer 7 load balancer, by contrast, actually understands the application protocol riding on top, so for HTTP traffic it can look inside the request itself, the path, the host header, custom headers, even cookies, and route based on what it finds there:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/images/*    → image servers
/api/*       → API servers
/admin/*     → admin service
X-Version:v2 → new backend
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the real trade-off underneath L4 versus L7: L4 sees less and costs less to process, L7 sees far more and can make far richer routing decisions because of it. Most real architectures don't pick one and stick with it exclusively, they layer both, an L4 balancer handling raw connection distribution with an L7 layer making smarter routing decisions on top.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round robin works right up until it doesn't
&lt;/h2&gt;

&lt;p&gt;The simplest algorithm is round robin, request one goes to A, request two to B, request three to C, and back around. That works fine when every server has roughly the same capacity and every request costs roughly the same amount of work. It falls apart the moment that stops being true. Picture Server A stuck processing a 20-second request while B and C sit idle, and the next request arrives. Pure round robin doesn't know or care that A is busy, it might send the new request straight to A anyway, because "who's next in rotation" was never the same question as "who actually has capacity right now."&lt;/p&gt;

&lt;p&gt;A few other algorithms exist specifically to fix this. Least connections sends each new request to whichever server currently has the fewest active connections, which matters a lot once request durations start varying widely. Weighted round robin accounts for servers that genuinely aren't equal, if A has twice the CPU capacity of B, giving A a weight of 2 against B's weight of 1 means A gets proportionally more traffic instead of being treated as identical. IP hash routes based on a hash of the client's IP, so the same client tends to land on the same server, useful for affinity, though it can produce uneven distribution when a lot of users happen to be sitting behind the same corporate NAT and therefore hash to the same bucket. There's no universally correct choice among these, the right one depends entirely on what's actually varying in your traffic, request duration, server capacity, or the need for a client to consistently land in the same place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Health checks are easy to get subtly wrong
&lt;/h2&gt;

&lt;p&gt;The basic idea is straightforward: the load balancer periodically checks something like &lt;code&gt;GET /healthz&lt;/code&gt; on each server, and after enough consecutive failures, it pulls that server out of rotation, then lets it back in once it starts passing checks again. Simple in concept, and surprisingly easy to implement badly in either direction.&lt;/p&gt;

&lt;p&gt;Make the check too shallow, "is the process running", and a server can report healthy while being completely unable to actually serve a request, maybe its database connection pool is exhausted, maybe a critical dependency is down, the process itself never noticed. Make it too deep instead, checking the database, checking Redis, checking two downstream APIs, and now one unrelated dependency having a brief blip can cause every single instance in the fleet to simultaneously report itself unhealthy. At that point the load balancer, doing exactly what it was told, pulls the entire fleet out of rotation at once, and the health check itself just caused the outage it was supposed to prevent.&lt;/p&gt;

&lt;p&gt;This is why modern deployments generally separate three distinct questions that sound similar but aren't: startup, has the application finished initializing, liveness, should this specific process be killed and restarted, and readiness, should traffic be sent here right this moment. For the load balancer's actual routing decision, readiness is almost always the one question it should actually be asking, since a process can be alive and still not be in any condition to usefully handle a request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sticky sessions fix one problem by creating another
&lt;/h2&gt;

&lt;p&gt;Say a user's login state lives in Server A's memory. Their first request lands on A and gets recorded there. Their next request happens to land on Server B instead, which has never heard of that session, and from the user's perspective they just got randomly logged out for no reason they can see.&lt;/p&gt;

&lt;p&gt;Sticky sessions solve this directly, by pinning a given client to the same server every time. It works, but it costs the load balancer some of its freedom to make good decisions elsewhere. If the server a user is pinned to gets overloaded, that user keeps going there anyway, because pinning doesn't know or care about current load. If that server needs to come out of rotation for a deploy, every user pinned to it has to move somewhere else all at once. And if it crashes outright, whatever session state was only living in its memory disappears with it.&lt;/p&gt;

&lt;p&gt;The cleaner fix, and the one most systems eventually converge on, is to stop keeping session state on individual servers at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             ┌── Server A ──┐
User → LB ───┼── Server B ──┼── Redis / DB
             └── Server C ──┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Move the session into something shared, Redis, a database, wherever, and every application server becomes genuinely stateless. Any server can now handle any request from any user, which is the same statelessness payoff that shows up everywhere else in distributed system design, once state lives externally instead of in a particular process's memory, the load balancer gets its full freedom back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The deployment detail people forget: connection draining
&lt;/h2&gt;

&lt;p&gt;Say Server B is being taken out of rotation for a deploy, and it currently has 300 requests actively in progress. Kill it outright and all 300 of those requests fail immediately, from the user's perspective for no visible reason, even though the rest of the infrastructure is completely healthy.&lt;/p&gt;

&lt;p&gt;The fix is to stop sending B &lt;em&gt;new&lt;/em&gt; traffic first, while letting the requests already running on it finish naturally, and only shut the instance down once those have actually completed. That's connection draining, and it's a small enough detail that it's easy to skip when you're first setting up a load balancer, right up until the first deploy that takes down 300 in-flight requests teaches everyone why it matters. It's genuinely one of the things separating "we have a load balancer" from "we can safely operate a load-balanced system" in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployments the load balancer makes possible
&lt;/h2&gt;

&lt;p&gt;Once traffic routing is centralized in one place, that same control point turns into a deployment tool, not just a distribution mechanism.&lt;/p&gt;

&lt;p&gt;Blue/green deployment keeps two full versions running side by side, the current version, Blue, taking 100% of traffic while the new version, Green, sits at 0% and gets verified. Once it's confirmed healthy, traffic switches, Blue to 0%, Green to 100%, and if something's wrong, switching back is just as immediate.&lt;/p&gt;

&lt;p&gt;Canary deployment is more gradual: instead of an all-at-once switch, the new version gets a small slice, say 5%, while the stable version keeps the rest. Latency, error rate, and other metrics get watched closely on that 5%, and if it looks healthy, the percentage climbs, 5% to 25% to 50% to 100%. If anything looks wrong at any step, traffic routes back to the stable version before the bad version ever sees full load. At this point the load balancer has stopped being a traffic distributor and become an actual part of how deployments happen safely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who load-balances the load balancer?
&lt;/h2&gt;

&lt;p&gt;There's an obvious uncomfortable question once every request is passing through this one component: doesn't that make the load balancer itself a single point of failure? Yes, if you only deploy one of them. Production systems make the load-balancing layer redundant too, and in most cloud environments this is handled for you, hidden behind a managed service that's already spread across multiple failure zones.&lt;/p&gt;

&lt;p&gt;At larger scale, another layer shows up entirely before the regional load balancer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                     ┌── US Region → Regional LB → Servers
Users → Global LB ───┤
                     └── EU Region → Regional LB → Servers
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GeoDNS or anycast routes a user toward whichever region is closest and healthy, and load balancing becomes genuinely hierarchical, a global layer deciding which region, a regional layer deciding which server within that region.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this actually looks like assembled together
&lt;/h2&gt;

&lt;p&gt;Put all of the above in one diagram and the trivial three-box picture from the start has grown considerably:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                        Internet
                           │
                           ▼
                 Global Routing / DNS
                           │
                           ▼
                  Regional Load Balancer
                     TLS termination
                     health checking
                     request routing
                           │
              ┌────────────┼────────────┐
              ▼            ▼            ▼
           Server A     Server B     Server C
              │            │            │
              └────────────┼────────────┘
                           ▼
                    Shared Redis / DB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Servers get added and removed freely. Unhealthy instances stop receiving traffic without anyone having to notice and intervene manually. Deployments drain old instances instead of killing in-flight requests. Canaries take a small, controlled slice of traffic before a rollout goes wide. TLS gets managed in one central place instead of on every individual server. And because the application servers themselves hold no state, any healthy instance really can serve any request from any user.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model worth keeping
&lt;/h2&gt;

&lt;p&gt;A load balancer isn't well described as "something that sends requests to different servers," that framing undersells almost everything it actually ends up doing. A better way to think about it: it's the control point sitting between clients and a pool of backend capacity that's constantly changing, servers coming up, going down, being deployed, being drained. Once something sits in that position, health checking, TLS termination, protocol-aware routing, connection draining, deployment control, regional failover, and observability all naturally accumulate around it, not because anyone planned for the load balancer to do all of that from day one, but because it's the one place in the architecture that already sees every request before it goes anywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;This post covers the core mechanics and production behavior of load balancing. The full visual breakdown on &lt;a href="https://seeitflow.com/networking/traffic-routing/loadbalancer/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; covers the fundamentals in more depth, there's a dedicated &lt;a href="https://seeitflow.com/networking/traffic-routing/loadbalancer/visual-learning" rel="noopener noreferrer"&gt;visual learning walkthrough&lt;/a&gt; that shows requests actually moving through the load balancer, health checks failing and recovering, and deployments draining in real time, and an &lt;a href="https://seeitflow.com/networking/traffic-routing/loadbalancer/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on the trade-offs behind algorithm choice, sticky sessions, and canary rollouts.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>systemdesign</category>
      <category>webdev</category>
      <category>backend</category>
    </item>
    <item>
      <title>TCP Isn't Reliable Because It's Magic. It's Reliable Because It Never Trusts the Network.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Sat, 26 Sep 2026 14:11:28 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/tcp-isnt-reliable-because-its-magic-its-reliable-because-it-never-trusts-the-network-b5d</link>
      <guid>https://dev.to/mangeshmandlik/tcp-isnt-reliable-because-its-magic-its-reliable-because-it-never-trusts-the-network-b5d</guid>
      <description>&lt;p&gt;"Reliable, ordered delivery over IP" is a fine one-line summary of TCP, and it's also the kind of definition that tells you almost nothing about why TCP is built the way it is. Why does every connection start with three messages instead of one? What's actually making delivery reliable under the hood? Why are flow control and congestion control treated as two separate mechanisms instead of one? And why can a packet loss rate of a fraction of a percent tank throughput on an otherwise fast connection?&lt;/p&gt;

&lt;p&gt;Those questions all have the same root: TCP is built entirely on top of a network layer, IP, that makes no promises at all. IP doesn't guarantee a packet arrives, arrives exactly once, or arrives in the order it was sent. Everything TCP does is machinery built specifically to hide that fact from the application sitting on top of it. I put together an interactive walkthrough of that machinery, handshakes, sliding windows, congestion control, on &lt;a href="https://seeitflow.com/networking/foundations/tcp/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to watch it rather than read through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A reliable stream, built on a network that promises nothing
&lt;/h2&gt;

&lt;p&gt;TCP's job is to take an unreliable, unordered, best-effort packet network and present something completely different to the application above it: a reliable, ordered byte stream. Sequence numbers identify exactly where each byte belongs in that stream. Acknowledgements tell the sender what actually made it across. Retransmissions recover whatever didn't. Reordering logic handles packets that arrive out of sequence, which happens constantly on a real network. Flow control keeps a fast sender from overwhelming a slow receiver. Congestion control keeps a fast sender from overwhelming the network itself. None of these individually is complicated, but together they're doing the work of turning "packets that might not show up" into "a stream you can trust."&lt;/p&gt;

&lt;p&gt;One detail here trips people up constantly: TCP is a byte-stream protocol, not a message protocol. If your application makes two separate &lt;code&gt;write()&lt;/code&gt; calls, there's no guarantee the other end sees two corresponding &lt;code&gt;read()&lt;/code&gt; calls. TCP doesn't preserve your message boundaries, it just guarantees the bytes arrive in order. Any protocol built on top of TCP, HTTP included, has to invent its own way of marking where one message ends and the next begins, because TCP itself doesn't know or care.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the handshake needs three messages, not two
&lt;/h2&gt;

&lt;p&gt;Before any application data moves, TCP establishes state on both ends of the connection:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client                         Server

   -------- SYN ------------&amp;gt;
   &amp;lt;----- SYN + ACK ----------
   -------- ACK ------------&amp;gt;

          ESTABLISHED
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both sides pick an initial sequence number for their side of the stream and confirm they've received the other's. The reason it takes three messages rather than two comes down to needing both directions verified independently: the client needs confirmation the server actually got its SYN, and separately, the server needs confirmation the client received the server's response. Two messages would leave one direction unconfirmed.&lt;/p&gt;

&lt;p&gt;That reliability isn't free, though. A brand-new connection pays a full round trip before any actual application data starts flowing, and on a high-latency path that round trip is real, measurable time spent doing nothing but confirming both sides are listening. This is exactly why connection reuse and connection pooling matter as much as they do in real systems, every new connection you avoid opening is a round trip you don't have to pay for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sequence numbers are what make "reliable" actually mean something
&lt;/h2&gt;

&lt;p&gt;TCP's sequence numbers track bytes, not packets. If a sender transmits bytes 1000 through 1999, the receiver's reply of &lt;code&gt;ACK 2000&lt;/code&gt; is saying something specific: everything before byte 2000 has arrived, send byte 2000 next. These acknowledgements are cumulative, so one ACK can implicitly confirm several earlier segments at once, the receiver doesn't need to acknowledge every single segment individually.&lt;/p&gt;

&lt;p&gt;This one mechanism is doing a surprising amount of work. It's what lets TCP detect a gap in the stream, discard a duplicate that arrived twice, reconstruct data that showed up out of order, and figure out precisely what needs to be resent when something goes missing. Reliability isn't a separate feature bolted onto sequence numbers, it emerges directly from tracking position in the byte stream this precisely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sliding windows: sending more than one thing before waiting
&lt;/h2&gt;

&lt;p&gt;A naively "reliable" protocol could send one segment, wait for its acknowledgement, then send the next. On a high-latency connection that would be painfully slow, most of the time would be spent waiting rather than transmitting. TCP instead allows multiple segments to be in flight, unacknowledged, at once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Sender                         Receiver

Segment 1  -------------------&amp;gt;
Segment 2  -------------------&amp;gt;
Segment 3  -------------------&amp;gt;
           &amp;lt;------------------- ACK

       window moves forward
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the sliding window, and how much data TCP is willing to keep in flight at once is governed by two separate mechanisms that get confused with each other constantly, because they sound like they might be the same thing and they're not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two different windows, protecting two different things
&lt;/h2&gt;

&lt;p&gt;The receiver advertises a receive window, &lt;code&gt;rwnd&lt;/code&gt;, representing how much buffer space it currently has available. If the application reading from the socket is slow to keep up, that buffer fills, &lt;code&gt;rwnd&lt;/code&gt; shrinks, and the sender is told, explicitly, to slow down. This exists purely to protect the receiver from being overwhelmed by data it can't process fast enough.&lt;/p&gt;

&lt;p&gt;Separately, the sender maintains its own congestion window, &lt;code&gt;cwnd&lt;/code&gt;, representing how much traffic it believes the network path can currently handle without breaking down. Packet loss and other congestion signals shrink &lt;code&gt;cwnd&lt;/code&gt;, independent of anything the receiver is doing. This exists purely to protect the network itself.&lt;/p&gt;

&lt;p&gt;The actual amount of data allowed in flight at any moment is roughly &lt;code&gt;min(rwnd, cwnd)&lt;/code&gt;, whichever constraint is tighter wins. The distinction worth keeping straight: flow control protects the receiver, congestion control protects the network. They can constrain a connection independently and for completely different reasons, and conflating them makes debugging a slow connection much harder than it needs to be.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a little bit of packet loss hurts a lot
&lt;/h2&gt;

&lt;p&gt;TCP treats packet loss as a signal, not just a nuisance to route around. When loss happens, TCP doesn't simply retransmit the missing bytes and move on, its congestion-control algorithm typically also shrinks &lt;code&gt;cwnd&lt;/code&gt;, on the assumption that loss might mean the network is congested and sending less aggressively is the safer bet. That means a relatively small loss rate can meaningfully reduce throughput, especially on a high-bandwidth, high-latency path where the window needs to stay large to keep the pipe full, and every loss event knocks it back down.&lt;/p&gt;

&lt;p&gt;TCP has two distinct ways of recovering from loss. Fast retransmit kicks in when enough duplicate ACKs arrive to signal a specific segment went missing, and TCP resends it immediately without waiting on any timer. Retransmission timeout is the fallback: if no acknowledgement shows up at all, TCP eventually resends after its retransmission timeout, RTO, expires. That RTO isn't a fixed number, it's calculated dynamically from the observed round-trip time and how much that RTT has been varying. This matters more than it sounds like it should for anyone setting application-level timeouts, an application timeout set too aggressively can give up on a request and abandon it before TCP itself has even had a chance to recover from what might have been a single, transient lost packet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bandwidth alone doesn't tell you what a connection can actually do
&lt;/h2&gt;

&lt;p&gt;Picture a very fast link connecting two distant regions. High bandwidth on that link doesn't automatically mean a single TCP connection can use all of it. The concept that actually determines this is the bandwidth-delay product:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;BDP = bandwidth × RTT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This approximates how much data needs to be in flight at once to fully saturate the link. If the TCP window is meaningfully smaller than the BDP, the connection becomes window-limited, it's sitting there with unused network capacity available, but it isn't allowed to have enough data in flight to actually use it. Modern operating systems generally auto-tune TCP buffer sizes to account for this, so manually cranking window settings shouldn't be the first move if a connection seems underutilized, measuring the actual RTT and throughput first tells you whether that's even the real bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep-alive means two different things depending on which layer you're at
&lt;/h2&gt;

&lt;p&gt;The names collide unfortunately here. TCP keep-alive is the transport layer probing an idle connection just to check whether the other end is still reachable at all. HTTP keep-alive, also called persistent connections, is an entirely different, application-layer idea: reusing one already-open TCP connection to carry multiple HTTP requests instead of opening a new one for each. Same phrase, two unrelated mechanisms at two different layers.&lt;/p&gt;

&lt;p&gt;Cloud environments add another wrinkle on top of both: load balancers, NAT gateways, proxies, and firewalls all frequently terminate idle connections according to their own independent timeout policies, ones you often don't control and might not even know about until a connection you assumed was still open turns out to have been silently dropped somewhere in the middle. This is exactly why long-lived connections in production often need a combination of application-level connection reuse, TCP keep-alive, and application-level heartbeats working together, no single one of the three covers every layer where a connection can quietly die.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three things that show up in real production systems
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Connection churn and TIME_WAIT.&lt;/strong&gt; Opening a fresh TCP connection for every small operation means paying a full handshake every single time, and at high enough request rates it also leaves large numbers of sockets sitting in &lt;code&gt;TIME_WAIT&lt;/code&gt;, which contributes to ephemeral port pressure and other resource exhaustion. The fix is almost always the boring one: reuse connections instead of constantly opening new ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Head-of-line blocking.&lt;/strong&gt; TCP's guarantee of ordered delivery cuts both ways. If one segment goes missing, none of the bytes that arrived after it can be handed to the application until that gap gets filled in, even if those later bytes have nothing to do with whatever was lost. This exact limitation, blocking at the transport layer regardless of what the application actually needs, is a big part of why HTTP/3 and QUIC were built to give independent streams their own delivery guarantees instead of sharing one ordered stream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nagle's algorithm interacting badly with delayed ACKs.&lt;/strong&gt; Nagle's algorithm batches small writes together to avoid flooding the network with tiny packets, which is genuinely useful for efficiency. But for latency-sensitive applications, that batching behavior can interact with the receiver's delayed-ACK behavior in a way that introduces noticeable, avoidable delay. This is why applications sending frequent small, latency-sensitive messages sometimes explicitly set &lt;code&gt;TCP_NODELAY&lt;/code&gt;, but that should be a deliberate, understood trade-off, not a default tuning knob flipped out of habit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually look at when a request is slow for no obvious reason
&lt;/h2&gt;

&lt;p&gt;Application logs are usually good at telling you a request was slow. They're rarely any good at telling you why, because the actual cause frequently isn't in the application code at all, it's sitting in the transport layer underneath it. The signals worth checking include the RTT distribution, the retransmission rate, connection counts, socket states like &lt;code&gt;TIME_WAIT&lt;/code&gt;, and the actual values of &lt;code&gt;rwnd&lt;/code&gt; and &lt;code&gt;cwnd&lt;/code&gt; at the time. On Linux, tools like &lt;code&gt;ss&lt;/code&gt;, &lt;code&gt;tcpdump&lt;/code&gt;, Wireshark, and eBPF-based utilities can surface all of this directly.&lt;/p&gt;

&lt;p&gt;One distinction here saves a lot of guesswork: a small &lt;code&gt;rwnd&lt;/code&gt; points at the receiver or the application reading from the socket not keeping up, while a small &lt;code&gt;cwnd&lt;/code&gt; combined with visible retransmissions points at an actual network or congestion problem. Those are different problems with different fixes, and conflating them tends to send people tuning the wrong layer entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  TCP vs UDP isn't reliable vs broken
&lt;/h2&gt;

&lt;p&gt;UDP deliberately skips TCP's reliability and ordering guarantees, and that's a design choice, not a shortcoming. For a file transfer, losing bytes silently is unacceptable, so TCP's retransmission behavior is exactly what you want. For live voice, video, or a multiplayer game, a packet that finally arrives after the moment it was relevant for has already passed is often worse than useless, waiting for it to be retransmitted just delays everything behind it for data nobody needs anymore. In systems like that, skipping the lost packet and moving on is frequently the better trade-off, one TCP won't make for you on its own.&lt;/p&gt;

&lt;p&gt;So UDP isn't a stripped-down, worse version of TCP. It's a protocol that hands the reliability decision back to the application instead of making it unconditionally. QUIC is a good example of what that flexibility enables: it runs on top of UDP but implements its own sophisticated reliability and congestion control, with independent streams specifically designed to avoid the cross-stream head-of-line blocking that's baked into how TCP works.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model worth keeping
&lt;/h2&gt;

&lt;p&gt;TCP takes IP's unreliable, best-effort packet delivery and turns it into a reliable, ordered byte stream, and it does that entirely through sequence numbers and acknowledgements working together. Sliding windows are what make that reliable stream fast rather than painfully slow, flow control protects the receiver from being overwhelmed, and congestion control protects the network from the same fate, for different reasons and via different mechanisms. Connection setup, retransmissions, packet loss, and connection churn all carry real, measurable latency and capacity costs once you're running this in production, they aren't just theoretical concerns from a networking course. And maybe the most useful habit to take from all of this: when a backend request is mysteriously slow, the application code isn't always where the answer lives. Sometimes the transport layer already told you exactly what went wrong, if you know which signal to check.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;This post covers TCP's core reliability mechanisms. The full walkthrough on &lt;a href="https://seeitflow.com/networking/foundations/tcp/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; covers the handshake, sliding windows, and congestion control visually, step by step. There's also a dedicated &lt;a href="https://seeitflow.com/networking/foundations/tcp/production-engineering" rel="noopener noreferrer"&gt;production engineering guide&lt;/a&gt; covering connection churn, keep-alive behavior, and Nagle's algorithm in more depth, and an &lt;a href="https://seeitflow.com/networking/foundations/tcp/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on debugging real transport-layer problems and the trade-offs behind TCP's design decisions.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>tcp</category>
      <category>backend</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>DNS Isn't a Phonebook. It's a Delegation Tree With a Cache Bolted On.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Sat, 26 Sep 2026 13:47:30 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/dns-isnt-a-phonebook-its-a-delegation-tree-with-a-cache-bolted-on-28cj</link>
      <guid>https://dev.to/mangeshmandlik/dns-isnt-a-phonebook-its-a-delegation-tree-with-a-cache-bolted-on-28cj</guid>
      <description>&lt;p&gt;You type &lt;code&gt;example.com&lt;/code&gt; into a browser, and a moment later it's talking to a server somewhere on the internet. That transition feels almost too smooth to be interesting, but there's a real system doing work in between, and it's worth understanding on its own terms rather than as the thing you only think about when it breaks.&lt;/p&gt;

&lt;p&gt;DNS gets called "the phonebook of the internet" often enough that the phrase has become the default mental model, and it's a reasonable starting point, but it hides most of what makes DNS actually work. DNS is a distributed, hierarchical, heavily cached naming system, and that design, not any single clever trick, is why billions of clients can resolve domain names without all of them depending on one enormous global database. I put together an interactive walkthrough of the whole resolution path on &lt;a href="https://seeitflow.com/networking/foundations/dns/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to see it laid out visually rather than read it top to bottom.&lt;/p&gt;

&lt;h2&gt;
  
  
  There's no single server that knows everything
&lt;/h2&gt;

&lt;p&gt;At the simplest level, DNS turns a name a human can read into information a machine can use: &lt;code&gt;example.com&lt;/code&gt; goes in, &lt;code&gt;93.184.216.34&lt;/code&gt; comes out. But there's no server anywhere holding every domain on the internet in one table. Responsibility is split across a hierarchy instead, and a lookup walks down through it: browser, to recursive resolver, to a root server, to the TLD server for &lt;code&gt;.com&lt;/code&gt;, to the authoritative nameserver for &lt;code&gt;example.com&lt;/code&gt;, which finally hands back the actual IP.&lt;/p&gt;

&lt;p&gt;The reason this scales is that each layer only needs to know a little. The root doesn't need example.com's IP address, it just needs to know who's responsible for &lt;code&gt;.com&lt;/code&gt;. The &lt;code&gt;.com&lt;/code&gt; servers don't store every record for every &lt;code&gt;.com&lt;/code&gt; domain, they just know which nameservers are authoritative for &lt;code&gt;example.com&lt;/code&gt;. The authoritative nameserver is the only one that actually holds the real records. Responsibility gets delegated one level at a time, and that delegation is what lets DNS scale without any single organization maintaining global knowledge of the entire internet's names.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recursive and iterative aren't two competing systems
&lt;/h2&gt;

&lt;p&gt;These two words get mixed up constantly because they both describe the same lookup, just from different vantage points. When your machine asks a recursive resolver for &lt;code&gt;example.com&lt;/code&gt;, it's essentially saying "give me the final answer, I don't want to deal with the hierarchy myself." That's the recursive part, from the client's perspective, one question goes out and one answer comes back.&lt;/p&gt;

&lt;p&gt;What the resolver does next, though, is iterative. It asks the root server, gets pointed toward &lt;code&gt;.com&lt;/code&gt;. It asks the &lt;code&gt;.com&lt;/code&gt; server, gets pointed toward &lt;code&gt;example.com&lt;/code&gt;'s authoritative nameserver. It asks that nameserver and finally gets the actual IP. Each of those exchanges is the resolver getting the best answer a given server has, which is often just a referral to the next one down. The useful way to hold both ideas at once: client to resolver is recursive, the client hands off the whole problem, resolver to the rest of the hierarchy is iterative, the resolver does the legwork itself, one hop at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Most lookups never reach the authoritative server at all
&lt;/h2&gt;

&lt;p&gt;Walking root, then TLD, then authoritative nameserver for every single request would be enormously wasteful, and DNS avoids that through caching at nearly every layer it can. A lookup might get answered from the browser's own cache, or the operating system's cache, or the recursive resolver's cache, and only fall through to the actual DNS hierarchy if all of those miss. In practice, the overwhelming majority of lookups get served from a cache somewhere before they ever reach the domain's real authoritative server.&lt;/p&gt;

&lt;p&gt;Every DNS record carries a TTL, time to live, that tells every cache along the way how long it's allowed to keep reusing the answer without asking again. A record with a TTL of 3600 seconds means a resolver can keep answering with that same IP for up to an hour without checking back in. That single number is behind one of DNS's more important production trade-offs: a long TTL means fewer queries and more cache hits, which is good for load and latency, but it also means changes take longer to actually reach everyone. A short TTL makes migrations and failover react faster, at the cost of more queries constantly landing on your DNS infrastructure. This is exactly why the standard advice is to lower a TTL well before a planned migration, not after, lowering it after you've already changed the IP does nothing for the copies of the old answer that resolvers cached while the TTL was still long.&lt;/p&gt;

&lt;h2&gt;
  
  
  "DNS propagation" is really just cache expiration, spread unevenly
&lt;/h2&gt;

&lt;p&gt;Say you update &lt;code&gt;api.example.com&lt;/code&gt; from &lt;code&gt;10.0.0.10&lt;/code&gt; to &lt;code&gt;10.0.0.20&lt;/code&gt;. The new value can be sitting on the authoritative nameserver immediately, correct and ready to be served. And yet some users will keep reaching the old IP for a while, sometimes minutes, sometimes hours, depending entirely on what their particular resolver happened to have cached and when that cache entry's TTL runs out.&lt;/p&gt;

&lt;p&gt;That's really the whole story behind what people casually call DNS propagation: the record changes at the authoritative server right away, some resolvers out there are still holding the old cached value, their TTL eventually expires, they ask again, and only then does the new value show up for whoever's using that resolver. There's no single moment when "the change propagates", there's a scattered set of moments, one per resolver, each governed by whenever its particular cached copy happens to expire. That's also exactly why the same DNS change can look instant to one person and take an hour to reach someone else, they're just hitting different resolvers with different cache states.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why DNS mostly uses UDP, but not always
&lt;/h2&gt;

&lt;p&gt;Most ordinary DNS queries are small, a hostname in, an IP out, so UDP is a good fit: send the query, get the response back, with very little transport overhead in either direction. And because a DNS query is safe to simply repeat if it goes missing, losing a UDP packet here and there isn't a serious problem, the client or resolver just asks again.&lt;/p&gt;

&lt;p&gt;But "DNS uses UDP" isn't the complete picture. DNS falls back to TCP when a response is too large to fit in the space UDP allows, and TCP is also what's used for things like zone transfers, where an entire zone's records move between servers at once. The more accurate framing is that DNS commonly uses UDP for ordinary queries, but TCP is very much still part of the protocol whenever the situation calls for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  A record isn't just a name pointing at an IP
&lt;/h2&gt;

&lt;p&gt;DNS stores more kinds of information than most people realize. An &lt;code&gt;A&lt;/code&gt; record maps a hostname to an IPv4 address, and &lt;code&gt;AAAA&lt;/code&gt; does the same for IPv6. A &lt;code&gt;CNAME&lt;/code&gt; points one hostname at another hostname rather than at an IP directly. &lt;code&gt;MX&lt;/code&gt; records tell the world which mail servers handle a domain's email. &lt;code&gt;TXT&lt;/code&gt; records hold arbitrary text, often used for domain verification or email security policies. &lt;code&gt;NS&lt;/code&gt; records say which nameservers are authoritative for a zone, and &lt;code&gt;PTR&lt;/code&gt; records do the reverse lookup, IP address back to hostname.&lt;/p&gt;

&lt;p&gt;That range of record types is why "distributed database" is a more accurate mental model for DNS than "phonebook." A phonebook only ever maps a name to a number. DNS is storing several different kinds of structured, typed data, and different applications query different record types for entirely different reasons.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two roles that get confused constantly
&lt;/h2&gt;

&lt;p&gt;A recursive resolver does the actual lookup work on a client's behalf and caches whatever it finds along the way. An authoritative nameserver is the one that actually stores the real records for a zone, it's the source of truth, not a caching layer. If someone asks "which server actually stores my domain's records," the answer is always the authoritative nameserver, never the resolver. The resolver is just the caching middleman that goes and finds those records for you, and remembers the answer for a while afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  DNS is also a routing mechanism, not just a lookup
&lt;/h2&gt;

&lt;p&gt;Once a system spans multiple regions, DNS starts doing more interesting work than simple name resolution. The same domain can resolve to different endpoints depending on where the request is coming from, a European visitor gets routed to an EU endpoint, someone in India gets an India endpoint, and DNS providers commonly support geo-based, latency-based, weighted, and failover routing to make that happen. Anycast takes a related idea further: the same IP address gets announced from many physical locations at once, and ordinary network routing carries a given request toward whichever announcing location is actually closest to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  DNS round robin isn't a load balancer, even though it looks like one
&lt;/h2&gt;

&lt;p&gt;A single hostname can resolve to several different IP addresses, &lt;code&gt;api.example.com&lt;/code&gt; returning &lt;code&gt;10.0.0.1&lt;/code&gt;, &lt;code&gt;10.0.0.2&lt;/code&gt;, and &lt;code&gt;10.0.0.3&lt;/code&gt; in rotation, and that does spread traffic across multiple servers in a rough sense. But DNS isn't sitting in the actual request path watching what happens to each request the way a real load balancer is. If one of those three servers becomes unhealthy, DNS has no immediate way to know or react, clients that already cached that IP will keep sending traffic to a dead server until their TTL expires, regardless of what's actually happening on the other end. This is exactly why production systems combine DNS-level routing with real load balancers and health checks, rather than treating round robin DNS as a full substitute for either.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this actually goes wrong
&lt;/h2&gt;

&lt;p&gt;The protocol itself is rarely the hard part, it's the operational assumptions built on top of it that cause real incidents. The classic one: a team runs an infrastructure migration, changes an IP, and expects traffic to move over right away, only to discover the old record had a long TTL, so resolvers everywhere keep sending users to the old destination until those cached answers finally expire on their own schedule, sometimes hours after the "migration" was supposed to be complete.&lt;/p&gt;

&lt;p&gt;A more dangerous version of the same underlying issue is a dangling DNS record, a &lt;code&gt;CNAME&lt;/code&gt; still pointing at some third-party resource that's since been deleted or deprovisioned. If that external resource isn't yours anymore but your DNS record still points there, someone else can potentially claim that resource and effectively take control of your subdomain, a real subdomain-takeover vulnerability that starts entirely from an unmonitored, stale DNS entry.&lt;/p&gt;

&lt;p&gt;DNS is also worth monitoring in its own right, not just assumed to be working silently in the background. Resolution latency, &lt;code&gt;NXDOMAIN&lt;/code&gt; and &lt;code&gt;SERVFAIL&lt;/code&gt; rates, unexpected record drift, domain expiry dates, and DNSSEC validation failures can all quietly degrade an application that looks completely healthy from every other angle.&lt;/p&gt;

&lt;h2&gt;
  
  
  DNSSEC and encrypted DNS solve different problems
&lt;/h2&gt;

&lt;p&gt;Classic DNS has no built-in way to prove a response is authentic, nothing stops a response from being tampered with in transit. DNSSEC adds cryptographic signatures so a validating resolver can verify that the DNS data it received hasn't been altered along the way. What DNSSEC doesn't do is encrypt the query itself, someone watching the network can still see which domain you looked up, they just can't tamper with the answer undetected. Encrypted transport, DoH or DoT, solves that separate problem, hiding the query from anyone observing the network. DNSSEC is about authenticity and integrity. DoH and DoT are about privacy. They're not substitutes for each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  A debugging habit worth having
&lt;/h2&gt;

&lt;p&gt;When a DNS change looks like it's "not propagating," the first useful move is to ask the authoritative nameserver directly rather than guessing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;dig example.com @ns1.example.com
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the authoritative server is already returning the new value, your configuration is correct and what you're actually looking at is a stale answer cached somewhere else in the chain, not a broken change. To see the entire resolution path as it actually happens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;dig +trace example.com
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And to check a specific record type against a specific resolver:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="go"&gt;dig MX example.com @8.8.8.8
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These three commands cover a surprising share of real DNS confusion, most of it turns out to be a cache holding an old answer rather than anything actually misconfigured.&lt;/p&gt;

&lt;h2&gt;
  
  
  The model worth keeping
&lt;/h2&gt;

&lt;p&gt;DNS isn't just "domain in, IP out." It's a distributed hierarchy, with delegation splitting responsibility across layers, typed records carrying more than just addresses, aggressive caching at nearly every layer a request passes through, TTLs deciding how quickly changes actually take effect everywhere, and a routing layer capable of sending different users to different places entirely. Once DNS looks like that instead of a simple lookup table, things like propagation delays, resolver caching behavior, CDN routing, and failover all stop being mysterious and start being the predictable consequence of a system built this way on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;This post covers the resolution path and the core ideas behind it. The full walkthrough on &lt;a href="https://seeitflow.com/networking/foundations/dns/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; goes further into the resolution flow, caching layers, TTL behavior, and record types visually. There's also a dedicated &lt;a href="https://seeitflow.com/networking/foundations/dns/production-engineering" rel="noopener noreferrer"&gt;production engineering guide&lt;/a&gt; covering migrations, dangling records, and monitoring in more depth, and an &lt;a href="https://seeitflow.com/networking/foundations/dns/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on the trade-offs behind TTL choices, routing strategies, and DNSSEC.&lt;/p&gt;

</description>
      <category>dns</category>
      <category>webdev</category>
      <category>networking</category>
      <category>beginners</category>
    </item>
    <item>
      <title>HTTP Looks Like One Arrow. It's Actually Seven Decisions.</title>
      <dc:creator>MANGESH MANDLIK</dc:creator>
      <pubDate>Fri, 25 Sep 2026 10:18:17 +0000</pubDate>
      <link>https://dev.to/mangeshmandlik/http-looks-like-one-arrow-its-actually-seven-decisions-21j5</link>
      <guid>https://dev.to/mangeshmandlik/http-looks-like-one-arrow-its-actually-seven-decisions-21j5</guid>
      <description>&lt;p&gt;Draw HTTP the way most people first learn it and you get one arrow going out and one coming back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client → Request → Server
Client ← Response ← Server
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's not wrong, exactly, it's just incomplete enough to be misleading. A real page load involves resolving DNS, establishing a connection, negotiating TLS, sending the request, receiving the response, checking whether any of it could've been served from cache, and usually fetching several more resources the same way. Understanding HTTP well has less to do with memorizing status codes and more to do with understanding why the protocol is shaped the way it is, because most of those shapes are deliberate answers to "how do we make this work at the scale of the entire web."&lt;/p&gt;

&lt;p&gt;I put together an interactive breakdown of the whole thing on &lt;a href="https://seeitflow.com/networking/foundations/http/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt;, if you'd like to see it laid out visually. Here's the written version.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual contract underneath everything
&lt;/h2&gt;

&lt;p&gt;Strip away every feature built on top of it, and HTTP is one exchange. The client sends a request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="nf"&gt;GET&lt;/span&gt; &lt;span class="nn"&gt;/products/42&lt;/span&gt; &lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt;
&lt;span class="na"&gt;Host&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;example.com&lt;/span&gt;
&lt;span class="na"&gt;Accept&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;application/json&lt;/span&gt;
&lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Bearer ...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and the server sends back a response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="k"&gt;HTTP&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="m"&gt;1.1&lt;/span&gt; &lt;span class="m"&gt;200&lt;/span&gt; &lt;span class="ne"&gt;OK&lt;/span&gt;
&lt;span class="na"&gt;Content-Type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;application/json&lt;/span&gt;
&lt;span class="na"&gt;Cache-Control&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s"&gt;max-age=300&lt;/span&gt;

&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Keyboard"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A method, a target, headers, optionally a body, going one direction. A status code, headers, optionally a body, coming back. That's genuinely the whole protocol at its foundation. Cookies, authentication, caching, compression, content negotiation, all of it is built on top of this one request/response shape, not bolted on as something separate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why statelessness is a feature, not a gap
&lt;/h2&gt;

&lt;p&gt;One property of HTTP surprises people the first time they think about it carefully: the server doesn't automatically know that this request came from the same person as the last one. Every request stands alone. If an application wants to recognize a returning user, it has to build that itself, a cookie, a session ID, a JWT, a record in a database or a Redis store keyed on some token.&lt;/p&gt;

&lt;p&gt;That sounds like something HTTP is missing. It's actually a large part of why HTTP-based systems scale as well as they do. Picture three application servers behind a load balancer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                ┌── Server A
Client → LB ────┼── Server B
                └── Server C
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If each server kept a user's session in its own local memory, the load balancer would need to remember which server that user landed on and keep routing them back there, sticky sessions. Lose that server and you lose the session with it. Keep the state external instead, in a cookie the client carries or a store every server can read, and any healthy server can answer the next request. That's what makes horizontal scaling and failover straightforward instead of a coordination problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The method is a promise the rest of the internet relies on
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;GET&lt;/code&gt;, &lt;code&gt;POST&lt;/code&gt;, &lt;code&gt;PUT&lt;/code&gt;, &lt;code&gt;PATCH&lt;/code&gt;, &lt;code&gt;DELETE&lt;/code&gt; aren't just different names for "do something," picked by convention. Each one tells every piece of infrastructure between the client and the server what kind of operation this is, and that infrastructure acts on it without checking with you first.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;GET&lt;/code&gt; means read, and more specifically it means safe, nothing changes as a result. That's exactly why browsers prefetch links, why crawlers follow them automatically, why caches and proxies feel entitled to store the response and serve it to someone else. An endpoint like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /delete-account
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;breaks every one of those assumptions at once. A crawler indexing the site can trigger a deletion just by following a link, because it was told &lt;code&gt;GET&lt;/code&gt; requests are safe to follow, and this one wasn't. That's not a bug in the crawler, it's the endpoint lying about what kind of request it is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Status codes are instructions, not just labels
&lt;/h2&gt;

&lt;p&gt;The useful grouping is &lt;code&gt;2xx&lt;/code&gt; for success, &lt;code&gt;3xx&lt;/code&gt; for redirects and cache validation, &lt;code&gt;4xx&lt;/code&gt; for something wrong with the request itself, &lt;code&gt;5xx&lt;/code&gt; for something that broke on the server's or an upstream's side while handling it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;200 OK              301 Moved Permanently     400 Bad Request        500 Internal Server Error
201 Created         304 Not Modified          401 Unauthorized       502 Bad Gateway
204 No Content                                403 Forbidden          503 Service Unavailable
                                               404 Not Found          504 Gateway Timeout
                                               429 Too Many Requests
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That grouping matters most when deciding what's worth retrying. A &lt;code&gt;400&lt;/code&gt; or &lt;code&gt;401&lt;/code&gt; will come back identical no matter how many times the exact same request is resent, so retrying just wastes a round trip confirming what you already knew. A &lt;code&gt;503&lt;/code&gt; or &lt;code&gt;502&lt;/code&gt; often represents something genuinely temporary, a struggling dependency, a proxy that couldn't reach its upstream in time, and retrying with backoff can actually succeed. Treating these two categories the same is how naive retry logic turns a brief hiccup into wasted load on an endpoint that was never going to say yes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching: the fastest request is the one you skip
&lt;/h2&gt;

&lt;p&gt;A response header like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;Cache-Control: max-age=3600
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;tells the browser it can reuse this exact response for an hour without asking again. When freshness needs checking without re-downloading the whole payload, ETags handle that cheaply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;If-None-Match: "abc123"
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and if nothing changed, the server replies &lt;code&gt;304 Not Modified&lt;/code&gt; with no body at all, just confirmation that what's already cached is still good.&lt;/p&gt;

&lt;p&gt;Production systems take this further with content-hashed filenames, &lt;code&gt;app.a1b2c3.js&lt;/code&gt;, &lt;code&gt;styles.8f91de.css&lt;/code&gt;, so those assets can be cached for close to forever, a new deployment produces a new filename rather than overwriting the old content under the same one. The HTML entry point is the one piece kept revalidated, so a new deploy takes effect immediately: fresh HTML references the new hashed filenames, and everything downstream of that can be cached as aggressively as you like.&lt;/p&gt;

&lt;h2&gt;
  
  
  HTTPS is HTTP, not something separate from it
&lt;/h2&gt;

&lt;p&gt;HTTPS isn't a competing protocol, it's HTTP carried inside a TLS-encrypted connection. TLS adds three properties HTTP alone doesn't have on its own: confidentiality, so nobody watching the network can read the traffic, integrity, so nobody can modify it in transit undetected, and authentication, so the client has a real way to verify which server it's actually talking to.&lt;/p&gt;

&lt;p&gt;That gives you an ordering: DNS resolves, a connection gets established, TLS negotiates, and only then does the HTTP request actually go out.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DNS → Connection → TLS handshake → HTTP request → HTTP response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every step there is a round trip happening before your application code ever sees the request, which is exactly why connection setup and handshake latency matter as much as they do for how fast a page feels.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why HTTP/2 and HTTP/3 exist
&lt;/h2&gt;

&lt;p&gt;HTTP/1.1 worked well for a long time, but a modern page pulls in dozens of resources, HTML, CSS, several JS bundles, API calls, images, and browsers historically worked around HTTP/1.1's limits by opening several parallel connections just to fetch more of them at once.&lt;/p&gt;

&lt;p&gt;HTTP/2 addressed the actual bottleneck with multiplexing: many requests share a single TCP connection.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;One TCP connection
├── Stream 1 → HTML
├── Stream 3 → CSS
├── Stream 5 → JavaScript
└── Stream 7 → API response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The catch is that HTTP/2 still rides on TCP, and TCP guarantees ordered delivery, so if one packet gets lost, every stream sharing that connection waits for it to be resent, even streams that had nothing to do with the lost packet.&lt;/p&gt;

&lt;p&gt;HTTP/3 solves that specific issue by moving off TCP and onto QUIC, which runs over UDP and gives each stream independent delivery. Lose a packet on one stream and only that stream stalls, the rest keep moving. The short version: HTTP/2 multiplexes requests over one connection, HTTP/3 keeps that multiplexing while getting rid of the head-of-line blocking TCP still imposes underneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually sits between a browser and your application
&lt;/h2&gt;

&lt;p&gt;A request rarely travels straight from a browser to an application server. A more realistic path looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
  ↓
CDN / Edge
  ↓
Reverse Proxy / Load Balancer
  ↓
Application Servers
  ↓
Cache / Database / Other Services
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CDN serves cached content from somewhere physically close to the user and often terminates TLS right there at the edge. A reverse proxy handles routing, TLS termination further in, buffering, compression, rate limiting. A load balancer spreads requests across whichever application instances are healthy right now. And because the application servers stay stateless, adding or removing instances doesn't require coordinating anyone's in-flight session, the statelessness property from earlier showing up again, one layer further out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this connects to everything else
&lt;/h2&gt;

&lt;p&gt;Once HTTP is running at real scale, most of the interesting problems aren't about HTTP syntax at all, they're things like connection pool exhaustion, cache stampedes, retry storms, a slow upstream dragging down an otherwise healthy service, or non-idempotent operations getting retried into duplicate side effects. A payment request can succeed on the server while the response itself gets lost on the way back, the client sees a timeout, retries, and if that operation isn't idempotent, the retry charges the customer a second time for something that already worked. That single failure mode is why production HTTP design ends up connected directly to idempotency keys, exponential backoff, jitter, rate limiting, caching, and having enough observability to tell these situations apart.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;HTTP looking simple from the outside is exactly why it's worth understanding what's underneath: statelessness that enables scaling, methods that are promises other systems rely on, status codes that carry instructions rather than just outcomes, caching that eliminates requests before they happen, and a whole layer of proxies and CDNs doing real work before a request ever reaches your code. None of that shows up in the one-arrow diagram, and all of it is what actually makes the web work at its current scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;p&gt;I kept this one focused on the request path itself. The fuller guide on &lt;a href="https://seeitflow.com/networking/foundations/http/fundamentals" rel="noopener noreferrer"&gt;SeeItFlow&lt;/a&gt; goes further into cookies and sessions, browser caching in more depth, and HTTPS/TLS handshakes step by step. There's also a dedicated &lt;a href="https://seeitflow.com/networking/foundations/http/production-engineering" rel="noopener noreferrer"&gt;production engineering walkthrough&lt;/a&gt; covering CDNs, reverse proxies, and load balancing, and an &lt;a href="https://seeitflow.com/networking/foundations/http/engineering-insights" rel="noopener noreferrer"&gt;engineering insights guide&lt;/a&gt; focused on production failure patterns.&lt;/p&gt;

</description>
      <category>networking</category>
      <category>http</category>
      <category>webdev</category>
      <category>systemdesign</category>
    </item>
  </channel>
</rss>
