A key with a TTL stops existing the moment that TTL passes. GET returns nil, EXISTS returns 0, and as far as anything your application can observe, the key is gone. What I wanted to know was when the memory comes back, because those are not the same event and I had never seen anybody put a number on the gap.
The answer, on an idle server with five million keys expiring together, is a little under twenty-three seconds. For most of that time Redis is holding several hundred megabytes for data that every client already agrees does not exist.
The part I did not expect is that a busy Redis cleans up more than four times faster than a quiet one.
Making the keys die together
My first attempt showed nothing at all, and the reason is worth repeating because it is the sort of mistake that produces a tidy, publishable, wrong conclusion.
I loaded a million keys with SET ... EX 15. Loading them takes about five seconds, so the first key's TTL starts five seconds before the last one's, and the expiries are smeared across that whole window. Redis kept up without breaking a sweat, reclaiming everything 1.1 seconds after the last key died, which gave me a clean graph of nothing happening and very nearly convinced me there was no article here.
The behaviour only appears when the keys expire at the same instant, which you get by giving them all one shared absolute deadline:
deadline = int(time.time() + 30)
for i in range(N):
p.expireat(f"k:{i}", deadline)
That is not an artificial setup. It is what a cache full of hourly rollups looks like, or anything keyed to the top of the hour, or a bulk import that stamps everything with the same end-of-day expiry.
Where the memory goes
With five million keys on Redis 8.10.2, all expiring on the same second:
at the expiry instant : dbsize=4,980,513 extra memory = 810 MB
dbsize reached zero : t+22.8s
DBSIZE still reports nearly five million keys at the moment they all expire, and keeps reporting a declining count for the next twenty-three seconds. Every one of those keys returns nil if you ask for it. The number is not lying exactly, it is counting things that have not been swept up yet.
Redis expires keys two ways. Lazily, when something touches a key and finds it dead, and actively, through a background cycle that runs hz times a second, samples twenty keys with TTLs, deletes the expired ones, and goes round again if more than a quarter of the sample turned out to be dead. That sampling is deliberately bounded so the cycle cannot monopolise the server.
The bit that surprised me
Since the active cycle is the thing doing the work, I assumed turning it up would be the lever. It is a lever:
| time to reclaim everything | |
|---|---|
idle server, default hz 10
|
22.8s |
idle server, hz 100
|
13.7s |
Ten times the cycle frequency bought a 1.7x improvement, which is a smaller return than I expected and suggests the per-cycle work limit matters more than how often it runs.
Then I ran the same test with a single client walking the keyspace with SCAN, doing nothing but reading:
5M keys, one client reading : 5.1s
Four and a half times faster than the idle server, and better than anything I achieved by tuning. Every read that lands on a dead key deletes it immediately, so ordinary traffic does the work that the background cycle is otherwise rationing. The practical shape of this is backwards from how people usually think about load: the quieter your Redis is, the longer dead data sits in it.
Does it push out live data?
This is the part I was most confident about and got wrong.
If several hundred megabytes of dead keys count toward maxmemory, and they do, then writing new data while they linger should force Redis to evict something real. I capped memory just above a four million key dataset, set allkeys-lru, waited for everything to expire, and wrote three hundred thousand fresh keys.
evictions triggered : 334,076
live keys written : 300,000
live keys still present : 300,000
LIVE KEYS LOST : 0 (0.0%)
Three hundred and thirty four thousand evictions, and not one of them cost me live data. The sampler picked dead keys, which is both the correct outcome and completely unsurprising in hindsight, since the dead keys outnumbered the live ones by more than ten to one and were colder by every measure LRU cares about.
So the alarming version of this story is not true, and I would rather say that than leave the implication hanging.
Where it does bite
noeviction is a different matter, because there is nothing Redis is willing to throw away. Same setup, same four million expired keys, policy left at the default:
4,000,000 keys using 504 MB, maxmemory 516 MB, policy noeviction
every key is logically expired. trying to write.
t+ 0.0s dbsize=3,989,147 used= 593MB REFUSED: OOM command not allowed when used memory > 'maxmemory'
t+ 1.2s dbsize=3,723,296 used= 558MB REFUSED: OOM command not allowed when used memory > 'maxmemory'
t+ 2.5s dbsize=3,374,872 used= 512MB write OK
Two and a half seconds of refused writes, with the real error text, caused entirely by memory held for keys that no longer exist in any sense a client can detect. Your application sees OOM command not allowed while the dataset it is worried about has already expired.
noeviction is the Redis default. It is also what DigitalOcean's managed Valkey ships with, which I only found out by asking it.
On a managed instance
I ran the same expiry test against DigitalOcean's managed Valkey 9 to see whether a managed service behaves differently. It does not: a million keys took 7.5 seconds to reclaim on a single vCPU node, against 3.2 seconds for the same test on two local cores, which is the difference you would expect from the hardware rather than from anything the platform is doing differently.
What is different is what you are allowed to do about it. CONFIG and DEBUG are disabled outright, so the hz lever is not available to you from a client at all, and neither is changing the policy that way. The eviction policy lives behind the platform API instead:
GET /v2/databases/{id}/eviction_policy
{"eviction_policy":"noeviction"}
That is the right call for a managed service, since CONFIG SET is a foot-gun on a cluster somebody else is responsible for keeping alive, and routing it through an audited API endpoint is a better answer than leaving it open. It does mean the cheapest fix available to a self-hosted Redis, raising hz, is not a fix you have.
What to do with this
If your keys expire on a schedule rather than individually, size your instance for the peak including the dead ones, because for a window measured in tens of seconds you are paying for both the old generation and whatever is replacing it.
Do not leave a cache on noeviction if it has a maxmemory anywhere near the working set, because the failure mode is refused writes during exactly the moment your keyspace turns over, which is also the moment you are busiest.
And if you have the choice, stagger the deadlines. Adding a random spread of a few minutes to a daily expiry costs nothing and turns a cliff into a slope. Everything in this post exists because I made five million keys die on the same second, and the first version of the experiment, where they died a few microseconds apart, showed no problem whatsoever.

Top comments (0)