DEV Community

Cover image for A Reverse Proxy Isn't One Feature. It's Five Cross-Cutting Concerns Living in One Place.
MANGESH MANDLIK
MANGESH MANDLIK

Posted on

A Reverse Proxy Isn't One Feature. It's Five Cross-Cutting Concerns Living in One Place.

Start with the simplest possible setup: a client talking directly to a single application server. No proxy, no extra hop, just a browser hitting one address and getting an answer back. That server ends up doing a surprising amount on its own once you look closely, it's terminating its own TLS certificate, serving its own static files like images and stylesheets alongside its actual application logic, and it's directly exposed to the internet with no layer in front of it filtering anything out.

Add a second server and every one of those responsibilities gets duplicated. Two TLS certificates to renew instead of one, two copies of the same static-file-serving logic, two separate attack surfaces facing the internet. Add an admin panel, a public API, and a static marketing site, and now routing logic starts getting tangled into the application itself, or worse, hardcoded into the clients that call it. None of these problems are hard individually. They're expensive specifically because they're the same problem, repeated once per server, forever. I put together a visual walkthrough of exactly this evolution, from direct access to a full proxy tier, on SeeItFlow, building it up the way I wish someone had explained it to me, one problem at a time.

What a reverse proxy actually is

A reverse proxy is a server that sits in front of your application servers, receives every client request on their behalf, and forwards it to whichever backend should actually handle it, then relays the response back. From the client's point of view, the reverse proxy is the application. The browser connects to one address, sends a request, gets a response back, and has no idea whether there's one app server behind that address or fifty, whether the response came from a cache, or whether a file was served straight off the proxy's own disk without an application ever getting involved. That indirection is the entire point.

Because the proxy sits on every single request without exception, it becomes the natural home for exactly the concerns that shouldn't live in individual applications: terminating HTTPS, serving static files, routing by URL, caching responses, and filtering out hostile traffic. Pull those five things out of the app and into one shared edge layer, and every backend server gets simpler for it.

Two proxies, opposite directions

The word "proxy" covers two roles that face in opposite directions, and mixing them up is an easy mistake to make once, since the underlying mechanics look similar.

A forward proxy sits in front of a group of clients. When those clients want to reach the internet, their requests go through the forward proxy first, which can filter, log, cache, or anonymize that outbound traffic. A corporate web filter or a VPN's egress point is a forward proxy, it works on behalf of the client, and the destination server on the other end usually has no idea the proxy was ever involved.

A reverse proxy sits in front of a group of servers instead. When clients out on the internet want to reach your application, they hit the reverse proxy first, which forwards the request to the right backend. It works on behalf of the server, and the client usually has no idea it's there. Same basic machinery, pointed in the opposite direction, serving the opposite side's interests.

It overlaps with a load balancer without being one

A reverse proxy can spread traffic across backends, and that shared capability is exactly why people conflate it with a dedicated load balancer. But load balancing is one job among several for a reverse proxy, and it's the only job for a dedicated load balancer.

A dedicated load balancer is deliberately narrow: distribute requests across identical replicas of a service, health-check them, do all of that at very high throughput, typically operating at Layer 4, connections, or a fairly basic Layer 7, with almost no understanding of the application itself. Its narrowness is what lets it be extremely fast and extremely reliable. A reverse proxy load-balancing across replicas is doing that same distribution, but as one feature bundled alongside HTTPS termination, response caching, static file serving, and path-based routing, all running on the same box. In most real production stacks, you actually run both, a dedicated L4 load balancer in front of a cluster of L7 reverse proxies, letting each layer stay narrowly focused and scale independently of the other.

It's also not the same thing as an API gateway

An API gateway can be thought of as a reverse proxy that specialized specifically for APIs and cross-cutting policy. A reverse proxy routes requests to backends, often identical replicas of the same application, and handles edge concerns like SSL, caching, and static files. It's content-aware enough to route by path, but it doesn't inherently understand your API's identity or enforce business-level rules.

An API gateway routes fundamentally different requests to different services, and layers on API-specific policy on top of that routing: authentication, authorization, per-key rate limiting, request aggregation, versioning. Under the hood a gateway is still built on reverse-proxy technology, but its center of gravity is policy and coordinating many distinct services, not edge plumbing for one application. Nginx, for what it's worth, can genuinely wear either hat, the question in any given deployment is which one it's actually configured to be.

What happens to one request, in order

A single request through a reverse proxy is really two independent connections and four distinct hops, with the proxy sitting as the pivot in the middle. The client opens a connection to the proxy and sends its request, that's hop one. The proxy then opens its own, separate connection to a backend and forwards the request along, hop two. The backend does the actual work and replies to the proxy, hop three. The proxy relays that response back down the original client-side connection, hop four.

That two-connection design is what unlocks essentially everything else a reverse proxy does. Because the client-side and backend-side connections are genuinely independent of each other, the proxy can terminate TLS on the client side while speaking plain HTTP to the backend, pool and reuse backend connections across many different clients, buffer a slow client without tying up a backend worker waiting on them, and even swap which backend it's talking to entirely without the client ever needing to reconnect. It's a full intermediary sitting on both halves of the exchange, not a redirect that steps out of the way after the first hop.

Terminating TLS once instead of everywhere

HTTPS has to be decrypted somewhere, and doing that once at the proxy instead of separately on every backend server is one of the largest operational wins a reverse proxy provides. With SSL termination, also called TLS offload, the client negotiates HTTPS with the proxy and only the proxy. The certificate and private key live in exactly one place. From the proxy inward, on your trusted private network, traffic runs as plain HTTP, so backend servers never touch a certificate and never pay the CPU cost of encryption themselves.

The operational payoff compounds quickly: one certificate to install, monitor, and rotate instead of one per server, a new backend inheriting HTTPS automatically the moment it's added behind the proxy, and upgrading cipher suites or TLS versions becoming a single config change instead of a fleet-wide rollout. The trade-off worth knowing: the leg from proxy to backend really is unencrypted plain HTTP. That's a completely reasonable choice on a trusted private network, but in zero-trust or regulated environments, that hop typically gets re-encrypted too, sometimes called TLS re-encryption or end-to-end TLS, via mTLS between the proxy and the app. You keep the centralized-certificate win on the public side while still encrypting the internal hop.

server {
  listen 443 ssl;
  server_name example.com;

  location / {
    proxy_pass http://app_servers;
    proxy_set_header Host $host;
    proxy_set_header X-Real-IP $remote_addr;
  }
}
Enter fullscreen mode Exit fullscreen mode

Serving static files without waking up the application

A large fraction of web traffic is static files that never change per request, a logo, a compiled JS bundle, a stylesheet. Making a full application runtime generate those on every request is genuinely wasteful. When a request for /logo.png arrives, the proxy can match it against its static rules, find the file directly on its own disk, and return it without ever contacting the application server at all.

A compiled proxy like Nginx serving a file straight from disk, often using sendfile so the kernel copies bytes from the file cache to the socket with zero application involvement, is dramatically faster and cheaper than a language runtime doing the equivalent work. Splitting static traffic off from dynamic traffic at the proxy is frequently the single biggest reason teams put something like Nginx in front of an application in the first place, all that repetitive traffic gets absorbed at the edge, leaving the application's CPU free for the dynamic work only it can actually compute. Fingerprinted filenames, app.4f2c.js rather than app.js, let you set very long cache TTLs without any risk of serving stale content, since a new deploy simply produces a new filename.

location /static/ {
  root /var/www;
  expires 30d;
}

location / {
  proxy_pass http://app_servers;
}
Enter fullscreen mode Exit fullscreen mode

Routing one domain to several independent services

A real site is rarely a single application. The proxy can front several independent services behind one domain and route each request based on its URL path, /api/* to the API service, /admin to the admin panel, /static/* to the asset service. To the outside world it looks like one coherent site. Behind the proxy, it's several genuinely independent services that different teams can deploy, scale, and own separately.

This is exactly the point where a reverse proxy starts to resemble an API gateway, the difference is that the proxy is routing purely by host and path, without layering on authentication, rate limiting, or aggregation policy on top. Nginx resolves each request to exactly one location block, matched by a defined precedence, exact matches first, then longest-prefix, then regex in file order. Get that precedence wrong and a broad location / placed above a specific /api rule will silently swallow API traffic that was supposed to go somewhere else, a routing bug that stays completely invisible until a user ends up on the wrong backend.

location /api/   { proxy_pass http://api_service; }
location /admin  { proxy_pass http://admin_service; }
location /static { proxy_pass http://asset_service; }
Enter fullscreen mode Exit fullscreen mode

Answering a request without waking the backend at all

If ten thousand people ask for the same page within a minute, regenerating it ten thousand times is pure waste. On a cache miss, the proxy has nothing stored yet, so it forwards to the backend and then keeps a copy of the response, tagged with a time-to-live, before passing it along. On a subsequent request for the same thing, while that copy is still fresh, the proxy can answer immediately and the backend never gets touched.

Hit ratio is the entire game here. At a 90% hit ratio, your backend sees roughly a tenth of the traffic it otherwise would, and users get responses back almost instantly. Request collapsing matters too: when a cached item finally does expire, a thundering herd of simultaneous requests for it should trigger exactly one backend fetch to repopulate the cache, not thousands of simultaneous ones. The classic mistake here is caching something that was never safe to share in the first place, per-user or rapidly changing content, under a key that doesn't account for who's actually asking. Caching a logged-in user's page under a shared key doesn't just fail to help, it leaks that user's data to the next visitor who happens to request the same URL.

The perimeter your backends never see

Because the proxy is the only thing actually exposed to the internet, it's the natural place to defend the entire system at once. Backends sit on a private network with no public address, so nobody outside can even attempt to reach them directly, only the proxy can. Hiding the internal topology genuinely matters here, an attacker can't target a box they can't see or address in the first place.

Because every request flows through this one choke point, it's also the natural place to inspect and reject bad traffic before it ever reaches an application that would otherwise have to parse it. TLS with modern cipher suites, security headers like HSTS and CSP, IP allow or deny lists, connection and request rate limiting, request size and timeout limits, and optionally a WAF module for known exploit patterns, all of it layers up at this one point. Blocking at the edge shrinks your actual blast radius meaningfully: the application never even parses hostile input, because the request already died at the proxy before reaching it.

The proxy itself can't be the single point of failure

A single proxy sitting on every request makes that one process a single point of failure for the entire site. If it dies, SSL, routing, and caching all go down with it, total outage, not a partial degradation. The fix is running several identical proxy instances as a stateless cluster and placing a dedicated load balancer in front of them, health-checking each instance individually. If one proxy fails, the load balancer simply stops routing to it, and users never notice anything happened.

Each layer in this stack scales independently and horizontally. Under load, you add proxy instances for more edge capacity; backends scale on their own separate axis for more compute. Because proxies absorb static, cached, and blocked traffic before it ever reaches an application, the backend tier that actually needs scaling is often just the genuinely dynamic remainder of the traffic, which tends to be a lot smaller than the total request volume hitting the edge.

What this looks like assembled together

Put every piece in one place and you get the shape behind most large-scale web systems: traffic from the internet hits a dedicated load balancer, which spreads it across a resilient tier of reverse proxies. That proxy tier terminates SSL, serves static files, caches responses, routes by path, and filters hostile traffic, then forwards whatever dynamic requests actually survive all of that to the backend services behind it. Every layer scales independently, and no single layer is a point of failure for the whole system.

Internet → Load Balancer → Reverse Proxy Cluster → Backend Services
Enter fullscreen mode Exit fullscreen mode

The handful of mistakes that take sites down repeatedly

A small set of reverse-proxy mistakes account for a disproportionate share of real outages. Running a single proxy instance with no redundancy behind it is the most basic one, one failure away from a total outage, always run a cluster of at least two behind a load balancer. A catch-all location / placed above more specific routes will silently swallow traffic meant for those specific routes, order rules from specific to general and prefer exact or longest-prefix matches for anything critical. Caching an authenticated response under a shared cache key leaks one user's private data to whoever else requests that same URL next, vary the cache key on session, or simply mark authenticated responses as non-cacheable entirely. Forgetting certificate renewal turns into a hard, total HTTPS outage the moment the cert expires, browsers refuse the handshake outright, so automate renewal and alert well before expiry rather than relying on someone remembering. Dropping the client's real IP is a quieter mistake, backends end up seeing the proxy's IP address for every single request unless X-Forwarded-For and X-Real-IP are explicitly set and, just as importantly, trusted carefully so they can't be spoofed by a malicious client. And reloading a proxy's configuration without validating it first means a single bad config gets pushed to every instance in the cluster simultaneously, nginx -t before every reload and canarying config changes catches this before it becomes an outage rather than after.

The mental model worth keeping

Reverse proxy, load balancer, and API gateway all sit in a similar physical position in the architecture, and it's tempting to treat them as interchangeable because of that. They're not answering the same question. A load balancer asks which replica should handle this. An API gateway asks which service should handle this, and whether the request is even allowed to happen. A reverse proxy is broader than either of those framings alone, it's the shared edge layer absorbing SSL termination, static serving, path-based routing, caching, and security, so that none of the backend services sitting behind it have to reimplement any of those five concerns on their own.

References

This post covers the core concepts and request lifecycle of a reverse proxy. The full guide on SeeItFlow covers the fundamentals in more depth across 13 chapters, there's a dedicated production engineering guide covering forwarding internals, TLS termination flow, and a full failure-and-scale playbook, and an engineering insights guide focused on where to cache, where to terminate TLS, and the comparisons against load balancers and API gateways in more detail.

Top comments (0)