DEV Community

Alex Georgiev
Alex Georgiev

Posted on AI-assisted

nginx 1.29.6 moves cookie-based session affinity out of nginx Plus

For as long as I've configured nginx upstream blocks, cookie-based session affinity has been something you paid for. The free build had ip_hash and the hash directive, both of which pin a client to a backend using properties of the connection rather than anything the application controls. Proper sticky sessions, the kind that follow a browser cookie, lived behind an nginx Plus subscription.

nginx 1.29.6, released in March 2026, moves the sticky directive into the open-source build. I pulled it, broke my usual ip_hash setup on purpose, and measured what changes when session affinity stops depending on where a request comes from and starts depending on a cookie.

The setup

Three backend containers, each a plain Python HTTP server that returns its own container hostname so I could tell which one answered. One nginx container in front, reconfigured between runs: default round robin, sticky cookie, and ip_hash, all pointed at the same three backends over a Docker bridge network. Every backend was reachable and healthy throughout; nothing here depends on failure.

The headline number

One client, 30 requests, three configurations:

upstream config backend 1 backend 2 backend 3
default (round robin) 10 10 10
sticky cookie srv_id 30 0 0

With the cookie set, every one of 30 requests from the same client landed on the backend nginx picked on request one. Without it, the same client's requests rotated evenly across all three. That's the entire pitch of the feature, and it held up exactly as advertised on the first test.

What ip_hash actually does to five different users

The reason I'd been using ip_hash for years is that it needed no application cooperation: no cookie, no header, just the client's address. The problem only shows up when several distinct users share one address, which is normal behind NAT, a corporate proxy, or a mobile carrier. I simulated that by running five separate clients, each with its own cookie jar, all issuing requests from the same source container, so nginx saw one IP address for all of them.

upstream config client 1 client 2 client 3 client 4 client 5
ip_hash backend 1 backend 1 backend 1 backend 1 backend 1
sticky cookie backend 2 backend 3 backend 1 backend 2 backend 3

Under ip_hash, all five clients, despite being genuinely different sessions, landed on the identical backend, every single request, 30 hits out of 30. Under sticky cookie, the same five clients, from the same shared address, spread themselves across all three backends and each one stayed put once assigned. This is the actual argument for the feature. It isn't really about stickiness at all; ip_hash was already sticky. It's about stickiness that doesn't quietly merge unrelated users the moment they share a network path.

I re-ran this twice more to be sure it wasn't an artefact of request order, and both numbers reproduced exactly: five-for-five collapse under ip_hash, correct separation under sticky cookie.

It combines with least_conn, which surprised me

I assumed sticky would be mutually exclusive with another balancing method, the way ip_hash and least_conn can't be used together (nginx rejects that combination outright). I was wrong. sticky cookie layered on top of least_conn passed nginx -t cleanly and behaved correctly at runtime: new sessions were assigned using least_conn's logic, and once assigned, every client stayed on its backend for the rest of the test, five clients, ten requests each, zero migration. The two aren't the same kind of directive. ip_hash and least_conn are both load-balancing methods, and nginx only allows one. sticky is a binding layer that sits on top of whichever method you choose for new sessions. Worth knowing if, like me, you assumed otherwise from the ip_hash precedent.

What's still gated behind a subscription

The open-sourcing isn't complete. sticky learn, which watches upstream responses for an application-issued session cookie rather than inventing its own, now works in the free build:

$ nginx -t
nginx: configuration file /etc/nginx/nginx.conf test is successful
Enter fullscreen mode Exit fullscreen mode

But its sync parameter, which replicates the learned-session table across a cluster of nginx instances, is still commercial-only, and says so plainly:

$ nginx -t
nginx: [emerg] unknown parameter "sync" in /etc/nginx/nginx.conf:7
nginx: configuration file /etc/nginx/nginx.conf test failed
Enter fullscreen mode Exit fullscreen mode

So the boundary is specific: cookie, route, plain learn, and the new drain/route server parameters are free as of 1.29.6. Cross-instance session replication is not. If your actual requirement is a cluster of load balancers that agree on where a session lives, this release doesn't get you there by itself.

And none of it exists before 1.29.6 at all. Running the identical config against 1.29.5:

$ nginx -t
nginx: [emerg] unknown directive "sticky" in /etc/nginx/nginx.conf:7
Enter fullscreen mode Exit fullscreen mode

Not a deprecated-but-working directive, not a warning. The whole block fails to parse.

Draining a backend without losing bound sessions

The drain server parameter is the other half of this release, also previously commercial. The documented behaviour is that a draining server keeps serving clients already bound to it via sticky, while refusing any brand-new session. I tested both halves. Across 40 fresh, cookie-less requests against two separately configured draining backends, zero landed on the drained server. Against a session cookie established before drain was switched on, eight replayed requests landed on the drained server eight times out of eight, whether the binding was a plain cookie hash or an explicit route value. The documentation's claim held exactly.

What it costs on every response

The sticky cookie isn't set once and forgotten. nginx resends Set-Cookie on every single response, not just the one that creates the session. Measuring raw response headers with and without it on an otherwise identical request:

headers, bytes
plain upstream 156
sticky cookie 268

112 extra bytes, on every request, forever, for as long as the session exists. That's nothing on a typical API response, but it's not zero, and nobody mentions it because the feature announcement is naturally about the backend behaviour, not the wire cost of the mechanism that provides it.

What I got wrong on the way

My first attempt at testing drain produced a result that looked like a documentation bug: a session already bound to the draining backend got silently redirected elsewhere, which is the opposite of what nginx's docs promise. I was about to write that up as the most interesting finding in this post.

The actual cause was my test harness, not nginx. I'd established the session against one nginx container and replayed its cookie against a second, separately started container, using Python's http.cookiejar, which enforces domain matching the way a real browser would: a cookie set while talking to host A doesn't get sent when the code then talks to host B, even if both containers run identical config. The cookie I thought I was replaying never left the client. Switching to an explicit Cookie: header, bypassing the jar's domain logic entirely, showed the real behaviour, which matched the documentation.

A second, unrelated harness bug showed up when I tried to reproduce the same test with curl's -b/-c cookie-jar file instead of Python: curl refused to load a saved cookie back for a bare, dotless hostname like nx-sticky, with the error cookie 'srv_id' dropped, domain '[file]' must not set cookies for 'nx-sticky', even though it had just set that exact cookie from a live response moments earlier. That's a curl quirk specific to single-label Docker Compose-style hostnames, not an nginx issue, but it's exactly the kind of thing that will quietly break a local reproduction if you reach for curl's file-based cookie jar rather than an explicit -b "name=value".

Both mistakes map to the same lesson: a tool designed to behave like a cautious real browser will sometimes be more cautious than your test setup expects, and the resulting "failure" is the harness, not the thing you're testing.

Run it yourself

This reproduces the headline comparison. It needs Docker and nothing else.

docker network create stickynet

cat > backend.py <<'PY'
import http.server, socket
class H(http.server.BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200); self.end_headers()
        self.wfile.write(f"{socket.gethostname()}\n".encode())
    def log_message(self, *a): pass
http.server.ThreadingHTTPServer(('0.0.0.0', 8080), H).serve_forever()
PY

for n in b1 b2 b3; do
  docker run -d --name $n --network stickynet \
    -v "$PWD/backend.py:/backend.py:ro" python:3.12-slim python3 /backend.py
done

cat > sticky.conf <<'EOF'
events {}
http {
    upstream backend {
        server b1:8080;
        server b2:8080;
        server b3:8080;
        sticky cookie srv_id expires=1h path=/;
    }
    server { listen 80; location / { proxy_pass http://backend; } }
}
EOF

docker run -d --name nx-sticky --network stickynet \
  -v "$PWD/sticky.conf:/etc/nginx/nginx.conf:ro" nginx:1.29.8

# one client, repeated requests, explicit cookie header (not a jar file --
# curl's jar file rejects bare Docker hostnames, see below)
docker run --rm --network stickynet curlimages/curl:latest sh -c '
cookie=""
for i in $(seq 1 10); do
  curl -s -D /tmp/h.txt -o /tmp/b.txt -b "$cookie" http://nx-sticky/
  cookie=$(grep -oE "srv_id=[a-f0-9]+" /tmp/h.txt | head -1)
  cat /tmp/b.txt
done'
Enter fullscreen mode Exit fullscreen mode

I ran that exact block before putting it in this post; it produced the same backend hostname ten times in a row.

What I measured here applies directly to any upstream block that uses ip_hash for a reason that was never really about load distribution. The failure mode is specific: a shared VPN exit, a corporate NAT gateway, or a mobile carrier's address pool puts several real sessions behind one IP, and ip_hash cannot tell them apart, because it was never given anything to tell them apart with. The fix is a single directive, sticky cookie srv_id expires=1h;, swapped in where ip_hash; used to sit, and as of 1.29.6 it compiles into the open-source binary with no licence file involved. Whether it belongs in a given config is a question about that config's own traffic, not one this post can answer from three containers on one machine.

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.