Picture a user in Mumbai requesting content from a server sitting in Virginia. No matter how fast that server is, no matter how optimized the application code running on it happens to be, the request still has to physically travel there and back, and that round trip has a hard floor set by the speed of light over fiber, measured in hundreds of milliseconds. You cannot optimize your way out of geography with a faster CPU.
That's the first reason a CDN exists, and it's the one most people already know. The second reason is less obvious and arguably more important: a single origin server is a single point of load. If a piece of content suddenly goes viral, every one of those requests would otherwise land on that one server, all at once, regardless of whether it was ever provisioned to handle that kind of traffic. A CDN turns "thousands of requests hitting the origin" into "one request hits the origin, thousands get served from a cache somewhere else entirely." Latency and origin protection are both the real answer here, not just one or the other. I put together a visual walkthrough of how this actually works, edge locations, cache hits and misses, invalidation, on SeeItFlow, if you'd like to see it rather than read through it.
What a CDN actually is
A CDN is a globally distributed network of servers that caches content physically close to users, so a request gets served from somewhere nearby instead of making the full trip to a single distant origin every time. From a user's perspective, it's completely invisible, they request a URL and get a fast response back, with no idea that a network of edge locations, also called points of presence or PoPs, is sitting between them and the server that actually owns the content.
The CDN's job is to answer as many requests as it possibly can directly from a cache at the edge, and only fall back to the origin when it genuinely has no other choice. Because it sits on every single request at the network's edge, it also ends up being a natural place to handle DDoS protection, TLS termination, and request routing, and increasingly even small amounts of actual compute. Almost every production CDN is wearing several of these hats simultaneously, not just the caching one.
Hit or miss, and why the ratio is the metric that matters
Every request reaching an edge location resolves to one of two outcomes. A cache hit means the requested object already exists at that edge location and hasn't expired yet, so the response gets served directly, no trip to the origin, minimal latency. A cache miss means the object isn't there, or it's expired, so the edge has to fetch it from the origin first, store a copy locally, and only then return it to the user. The very first request for anything new is always a miss by definition. Every request after that, until the cached copy expires, is a hit.
Cache-hit ratio, the percentage of requests served from cache versus the total, is the single most-watched CDN health metric that exists, and for good reason. A dropping hit ratio is usually one of the earliest signals that something's gone wrong, a misconfigured cache key, a TTL set too short, a sudden spike in genuinely unique, uncacheable requests, and it tends to show up on a dashboard well before origin load becomes visibly dangerous.
TTL: the number that decides how fresh is fresh enough
Time-to-live determines how long a cached object is considered fresh before the edge needs to re-check with the origin. It's usually set through HTTP response headers, most commonly Cache-Control: max-age=<seconds>, or the older Expires header. Once that TTL runs out, the object is considered stale, and the next request for it triggers either a fresh fetch from the origin or a conditional request using ETag/If-None-Match to check whether it actually changed at all.
There's no universally correct TTL, it's a genuine trade-off made per asset. A short TTL means fresher content but more load reaching the origin. A long TTL means far less origin load but more risk of serving something outdated if it changes unexpectedly. A hashed, versioned JS bundle can reasonably cache for a year, since its filename changes the moment its content does. A homepage's HTML might only be safe to cache for a few seconds, or not at all. Two patterns soften this trade-off further: stale-while-revalidate serves the stale cached copy immediately while refreshing it in the background, so no single user ever waits on the refresh, and stale-if-error keeps serving the last known-good cached copy if the origin becomes unreachable or starts erroring, rather than failing the request outright just because the origin is having a bad moment.
When you can't wait for the TTL to run out
Sometimes a TTL expiring naturally isn't fast enough. Cache invalidation, or purging, forcibly removes a stale object from edge caches immediately, typically triggered right after a deployment pipeline publishes new content, "purge /app.js and /styles.css", so users don't have to sit through a long TTL window just to see an update that already shipped.
The mistake worth avoiding here is assuming a purge is instantaneous everywhere at once. Propagating a purge across hundreds of globally distributed PoPs takes measurable time, seconds at minimum, sometimes longer for wildcard purges, so there's genuinely a short window where some regions are still serving the old version while others have already updated. A common alternative sidesteps the whole problem: bake a content hash directly into the filename, app.a1b2c3.js, so a new deployment automatically produces a new cache key with no purge step required at all, and the old version can safely be cached forever since nothing ever references it again.
Not everything deserves to be cached
Static content, images, CSS, JS bundles, videos, fonts, is identical for every single user and only changes on deployment, which makes it ideal for long TTLs and aggressive edge caching. Dynamic content, personalized API responses, authenticated pages, real-time data, differs per user or per request, and usually bypasses the cache entirely via Cache-Control: no-store, getting forwarded straight through to the origin instead.
But the line between those two categories is softer than it first looks. Some dynamic-looking content is genuinely cacheable for short windows, a product listing that updates every few minutes can still benefit meaningfully from a 30-second edge cache, cutting a large share of origin load without serving anything a user would notice as stale. The more useful question for any given endpoint isn't "is this static or dynamic," it's "how stale can this be before someone actually notices or cares." That answer, not the content type on paper, is what should actually drive the TTL.
The cache key decides what counts as "the same request"
By default, most CDNs build a cache key from the request path plus, often, the full query string. Two requests differing only by ?color=red versus ?color=blue get treated as entirely separate cached objects unless you configure it otherwise. CDNs let you customize this, stripping out tracking parameters like utm_source so semantically identical requests share one cache entry instead of fragmenting into thousands of near-duplicate ones, or including specific headers via the Vary header when a response genuinely differs based on something like Accept-Language.
This is also where one of the more serious CDN mistakes lives. Caching a response that varies by Authorization or a session cookie, without including that in the cache key or declaring it in Vary, can leak one user's personalized or authenticated response straight to a completely different user who happens to request the same URL next. It's the single most damaging category of CDN misconfiguration in production, not because it's exotic, but because it's extremely easy to miss during a routine setup and invisible until someone notices they're looking at a stranger's data.
Keeping the origin from getting hit all at once
Without any special handling, every regional edge location independently experiences a cache miss the first time a newly popular object shows up, and all of them hit the origin simultaneously, a thundering herd from the origin's point of view even though no single PoP did anything wrong. Origin shielding designates one edge location as the only one allowed to talk directly to the origin. Every other PoP routes its misses through that shield location instead, which fetches and caches the object exactly once and then serves every other region from there. It's one of the lowest-effort, highest-impact changes available for any origin that isn't trivially horizontally scalable, especially during a cold-cache event like a fresh deploy or a sudden viral spike.
Security that happens before traffic ever reaches you
Sitting on every single request makes a CDN a natural security enforcement point, not just a performance layer bolted on top. TLS termination happens right at the edge, closest to the user, cutting handshake latency compared to a distant origin doing the same negotiation itself. A Web Application Firewall can filter known malicious patterns, SQL injection attempts, XSS payloads, known bad bot signatures, before any of it reaches an actual application server. Signed URLs restrict access to protected content, paid media, private downloads, using time-limited tokens validated right at the edge, with no need to hit the origin just to authorize each individual request.
Because a CDN spreads traffic across every PoP globally via anycast, a volumetric DDoS attack gets naturally distributed across the whole network rather than concentrating entirely on one target. Rate limiting at the edge blocks abusive clients before they ever reach the origin, and challenge pages or JS-based bot mitigation can filter automated attacks while letting legitimate users straight through. None of this matters much, though, if the origin's real IP address is still publicly reachable, an attacker who can hit the origin directly bypasses every one of these protections entirely, which is why locking the origin down to only accept traffic from the CDN's own IP ranges is such a disproportionately valuable, low-effort step.
A CDN isn't a load balancer, even though they overlap
These two get confused constantly because they both sit somewhere in the request path distributing traffic, but they're solving genuinely different problems. A CDN sits between the user and the entire origin infrastructure, deciding whether a request even needs to reach the origin at all. A load balancer sits in front of a specific pool of backend servers, typically at or near the origin itself, deciding which one of those servers should actually handle a request that does make it through.
They compose rather than compete. A fairly typical architecture looks like user, to CDN edge, to load balancer, to application servers, the CDN absorbs everything cacheable before it ever becomes the load balancer's problem, and the load balancer distributes whatever's left, the genuinely dynamic, uncacheable remainder, across the application fleet. A CDN can't replace a load balancer here, it has no way to meaningfully distribute per-request, uncacheable dynamic traffic across a server pool, that's precisely the job a load balancer exists to do.
A CDN is really a reverse proxy, scaled globally
The cleanest way to understand a CDN's actual mechanism is as a globally distributed, managed reverse proxy. A reverse proxy like Nginx or Varnish sits in front of an origin, forwarding and optionally caching requests, usually from one location, close to or inside the origin's own infrastructure. A CDN takes that exact same forwarding-and-caching idea and spreads it across hundreds of geographically dispersed edge locations instead of one, adding global routing, DDoS absorption, and origin protection at a scale a single reverse proxy was never meant to operate at. Many CDN vendors are, quite literally, running reverse-proxy software, or a heavily modified fork of one, at every single PoP. The products differ in scale and geographic distribution, not in the underlying mechanism.
Where this goes wrong in production
A handful of mistakes account for most of the real incidents once a CDN is actually running in front of production traffic. Treating it as fire-and-forget, enabling it with default settings and never revisiting cache-control headers, TTLs, or cache keys as the application evolves, is probably the most common one, caching policy that lives entirely in a dashboard rather than version-controlled code is invisible during code review and easy to silently break. Caching a personalized response by accident, forgetting to mark an authenticated page as private or no-store, causes exactly the user-data-leak scenario described above. Not locking down the origin leaves its real IP reachable and every CDN-level protection trivially bypassable. And ignoring cache-hit ratio until an actual incident forces attention onto it means a caching regression only gets discovered during a traffic spike, instead of being caught early through routine monitoring that was watching for it all along.
One more worth calling out specifically because the naming is actively misleading: no-cache does not mean "don't cache." It means "cache this, but always revalidate with the origin before serving it." no-store is the directive that actually forbids caching outright. Mixing these two up is a genuinely common production gotcha, not just an interview trick question.
The question worth asking about any caching strategy
A useful test for whether a caching setup is actually protecting the origin, rather than just happening to look fast most of the time: what happens to the origin if the CDN's cache were completely cold right now, this second? If the honest answer is "it falls over," the caching strategy is under-protecting the origin no matter how good the measured latency numbers currently look. Treating origin protection as an explicit design goal, not an incidental side effect of caching for speed, changes real decisions, it's what justifies micro-caching a dynamic endpoint for even a few seconds, or paying for origin shielding, even at the cost of a little extra staleness, because the alternative is an origin that can't survive its own traffic the moment the cache goes cold.
Where this is heading
The newest layer on top of all of this is edge computing, running actual application logic at the same distributed locations that used to only cache static files. Edge functions can handle request rewriting, A/B test bucketing, authentication checks, even full server-side rendering, physically close to the user, cutting out round trips to the origin for logic that never needed the full backend stack in the first place. This is steadily blurring the line between "CDN" and "application platform," and it's worth knowing the trend exists even if a given system isn't using it yet, since it comes up constantly in forward-looking system design conversations.
The mental model worth keeping
A CDN caches content at globally distributed edge locations for two reasons at once, cutting latency by physically moving content closer to users, and shielding the origin from load it was never provisioned to handle directly. Every request resolves to a hit or a miss, TTL controls how long something stays fresh before that distinction gets re-evaluated, and invalidation exists for the moments when waiting out a TTL isn't fast enough. Not everything belongs in a cache, and the real question for any given piece of content is how stale it can get away with being, not whether it's technically static or dynamic. And the cache key, quietly, is deciding what counts as "the same request" the entire time, get it wrong and you're either wasting cache space or handing one user's data to another.
References
This post covers the core caching mechanics and production trade-offs of a CDN. The full guide on SeeItFlow covers the fundamentals in more depth across edge locations, TTL, and cache keys, there's a dedicated production engineering guide covering cache strategy, origin shielding, DDoS protection, and a full debugging checklist, and an engineering insights guide focused on CDN versus load balancer, CDN versus reverse proxy, and the real trade-offs behind invalidation and origin protection.
Top comments (0)