DEV Community

Libme
Libme

Posted on

Users Randomly Logged Out? Your Redis Eviction Policy Is Deleting Sessions

If a subset of your users get logged out at unpredictable times — no deploy, no session-secret change, worse during traffic peaks — check maxmemory-policy on your Redis instance before you touch cookie code. When Redis hits its memory ceiling with an allkeys-* policy, it deletes whatever keys are least recently used to make room, and it does not care that some of those keys are login sessions or queued jobs. The fix is not a bigger instance; it's separating keys you can afford to lose from keys you can't.

I lost most of a day to this on a side project where sessions, a page cache, and a BullMQ queue all shared one small managed Redis. The cache was doing its job — filling memory — and Redis was doing its job — throwing things out. Together they produced a third behavior nobody asked for: random logouts and jobs that silently never ran.

Why does Redis delete keys that still have time left on their TTL?

Redis with maxmemory set is a bounded store, and maxmemory-policy decides what happens when a write would cross that bound. The policies split into three families:

  • noeviction — refuse the write. The client gets OOM command not allowed when used memory > 'maxmemory'.
  • volatile-lru, volatile-lfu, volatile-random, volatile-ttl — evict only keys that have a TTL set. If there are no such keys, writes fail like noeviction.
  • allkeys-lru, allkeys-lfu, allkeys-random — evict anything, TTL or not.

The trap is that eviction is invisible from the application's side. A deleted session key is indistinguishable from an expired one: your middleware looks up the session, gets nil, and correctly concludes "not logged in." There's no error to catch, no log line, nothing that says this key was taken from you. Same for a queue: BullMQ or Sidekiq asks for a job hash that isn't there anymore and the job just doesn't exist.

Two details make this harder to reason about than it should be. First, volatile-* policies feel safe ("only expiring keys get evicted") right up until you realize sessions almost always have a TTL — they're prime eviction candidates, and the ones with the longest remaining life under volatile-ttl survive while your active-but-recently-renewed ones may not. Second, SELECT-ing a different logical database (db 0 vs db 1) does not isolate anything: all databases on an instance share one memory budget and one eviction policy, and Redis Cluster drops multiple databases entirely. Prefixing keys and switching DB numbers buys you tidiness, not safety.

Takeaway: eviction is a silent DEL issued by the server, so any key whose absence changes correctness must live somewhere eviction cannot reach.

How do I confirm eviction is what's happening?

Three commands, in this order. Don't trust the default documented in the docs — managed providers ship their own parameter defaults, and someone may have changed yours years ago.

# 1. What policy is actually in force, and how close to the ceiling are we?
redis-cli INFO memory | grep -E 'used_memory_human|maxmemory_human|maxmemory_policy'

# 2. The smoking gun: a non-zero, *increasing* counter.
redis-cli INFO stats | grep -E 'evicted_keys|expired_keys'
Enter fullscreen mode Exit fullscreen mode

evicted_keys is a monotonic counter since the last restart. Any non-zero value on an instance holding sessions or jobs is a bug report. Sample it twice a minute apart — a rising number during the window your users complain about is as close to proof as you'll get after the fact.

If you need to catch it live, keyspace notifications will tell you exactly which keys are going:

# 'E' = keyevent channels, 'e' = evicted events
redis-cli CONFIG SET notify-keyspace-events Ee
redis-cli PSUBSCRIBE '__keyevent@*__:evicted'
Enter fullscreen mode Exit fullscreen mode

Each message carries the evicted key name, so you'll see sess:... or bull:mail:... scroll past and the argument ends there. Two honest caveats: notifications are fire-and-forget (a disconnected subscriber misses events entirely, so this is a debugging tool, not an audit log), and CONFIG SET doesn't survive a restart unless you CONFIG REWRITE or change it in your provider's parameter group.

The dead ends I burned time on first, so you can skip them: rotating the session secret, SameSite/Secure cookie flags, sticky-session config on the load balancer, and clock skew on token expiry. All plausible causes of logouts, none of which produce a rising evicted_keys.

Takeaway: evicted_keys > 0 on an instance storing sessions or jobs is not a tuning opportunity, it's a correctness defect.

What policy should each workload get?

Pick per role, not per cluster. The point of the table is that "which policy is best" is the wrong question — the right one is "what does this instance hold."

Instance role Policy Failure mode you're accepting Alert on
Page/query cache allkeys-lru (or allkeys-lfu for skewed access) Cache misses, extra DB load Hit ratio drop
Sessions / auth noeviction + TTL on every key Writes fail loudly at the ceiling evicted_keys > 0, OOM errors
Job queue / broker noeviction New enqueues rejected; existing jobs safe Any OOM error, memory > 70%
Rate limit counters volatile-lru, short TTLs Some limits reset early (fails open) Memory trend only
Locks / idempotency keys noeviction Lock acquisition fails instead of double-granting evicted_keys > 0

noeviction feels scary because it turns a memory problem into visible request failures. That's the point: an OOM error is a page you can respond to, while an evicted lock key is a duplicate charge a customer tells you about. Sidekiq's documentation has recommended noeviction for its Redis for years for exactly this reason.

Enforce it at boot rather than trusting a runbook. This has caught a mis-restored parameter group for me more than once:

import Redis from 'ioredis'

const cache = new Redis(process.env.REDIS_CACHE_URL)      // evictable
const durable = new Redis(process.env.REDIS_DURABLE_URL)  // sessions, queue, locks

export async function assertRedisPolicies() {
  const [, policy] = await durable.config('GET', 'maxmemory-policy')
  if (policy !== 'noeviction') {
    throw new Error(
      `durable Redis has maxmemory-policy=${policy}; sessions and jobs can be silently evicted`
    )
  }
}
Enter fullscreen mode Exit fullscreen mode

Fail the deploy on that, and the class of bug disappears instead of getting rediscovered next year.

Is a second instance worth the cost?

For most small teams, yes, and it's often cheaper than the upsize you were about to buy. Cache is the part that wants headroom; sessions, locks, and queue metadata for a modest app are megabytes, so the durable node can stay tiny.

If you want a Redis-compatible store where eviction isn't a correctness risk at all, Amazon MemoryDB is the one that treats memory as a cache over a durable multi-AZ transaction log rather than as the only copy of your data — it costs meaningfully more than ElastiCache per node, and it inherits the same single-threaded hot-key ceiling, so it solves durability, not throughput. If you'd rather not pay for a second always-on node, Upstash bills per request instead of per instance, which makes creating one database per role effectively free — the tradeoff is that a chatty cache client becomes a line item, so you watch command volume there instead of memory. And if you moved to Valkey after the Redis licensing change, note that it kept the same eviction semantics and INFO fields, so both the bug and every diagnostic above migrate with you unchanged.

Takeaway: the second instance isn't about capacity, it's about giving eviction a blast radius you chose on purpose.

FAQ

Why are my Redis keys disappearing before the TTL expires?
Almost always eviction under maxmemory pressure. Run redis-cli INFO stats | grep evicted_keys — if that number is climbing, Redis is deleting keys to stay under its memory limit, and an allkeys-* policy lets it take keys that have plenty of TTL left.

Is it safe to use the same Redis for caching and sessions?
Only if the policy is noeviction, and then your cache stops absorbing growth gracefully and starts failing writes instead. Separate instances are the safe answer; separate logical databases (SELECT 1) are not, because all databases on one instance share the same memory budget and eviction policy.

What is the best maxmemory-policy for Redis?
There is no single best one — it depends on whether losing a key is acceptable. Use allkeys-lru for pure caches, and noeviction for sessions, job queues, locks, and idempotency keys where a missing key changes program behavior rather than just costing you a recompute.

Bottom line

If users report random logouts or jobs that vanish without a trace, check evicted_keys and maxmemory-policy before you audit a single line of auth code. Caches should be evictable and everything whose absence changes correctness should not be, which in practice means two Redis instances with two policies and a boot-time assertion that nobody has quietly changed them. Keep noeviction on the durable side even though it converts memory pressure into loud request failures — loud is the feature. Bigger instances only move the date this happens.

Related reading

Top comments (0)