DEV Community

Alex Georgiev
Alex Georgiev

Posted on AI-assisted

Valkey 9.2's forkless BGSAVE cuts my memory spike from 350MB to 10MB

Every BGSAVE I have ever watched on a busy Redis or Valkey instance does the same thing: the process forks, and for a few seconds the host's free memory drops like someone pulled a plug. Valkey 9.2.0-rc1, tagged on Docker Hub on 16 September, adds a way to skip the fork entirely. I wanted to know what that actually buys you, not what the release notes say it buys you, so I ran both paths against the same dataset and measured what happened.

The short version: the memory spike nearly disappeared, and the save got noticeably slower doing it.

The setup

I pulled valkey/valkey:9.2.0-rc1 and ran two containers from the identical image, differing only in one setting. The first used the default: bgsave-default-method fork. The second started with --forkless-infrastructure-enabled yes --bgsave-default-method forkless, which is the only way to turn it on — more on that below.

Into each I loaded 3,000,000 keys at 300 bytes each, a little over 1.1GB of data (used_memory reported 1.04GiB). Then, against each container, I ran eight threads hammering SET on random existing keys as fast as they could go, and from a ninth connection I sent a PING every 2ms and timed the round trip. Two seconds into each 12-second run I fired BGSAVE. I repeated this seven times per mode.

This setup matters for one reason: copy-on-write only costs you anything if pages are actually being written while the fork holds them. A BGSAVE against an idle dataset tells you almost nothing. Mine wasn't idle — the eight writer threads pushed roughly 27,000 operations a second into whichever container was running fork mode.

The memory spike, measured twice

I tracked two numbers: the RDB-reported copy-on-write size (rdb_last_cow_size in INFO persistence), and the container's own cgroup memory usage sampled every 150ms through the save.

fork (default) forkless
rdb_last_cow_size, mean of 7 runs 354.7MB 0MB (always)
cgroup memory delta during save, mean of 4 runs 354.4MB 9.9MB
save wall time (rdb_last_bgsave_time_sec) 3s, every run 5-6s

The two memory measurements agree with each other, which is the point of taking both: the kernel's own copy-on-write accounting and an independent cgroup memory sample converge on the same number through two unrelated instruments. On a 1.1GB working set under sustained writes, forking cost roughly a third of the dataset's size in extra memory, every single time, with a tight range of 319.6MB to 367.5MB across the seven runs. Forkless never went above 10.0MB.

That is the headline, and it held up every time I ran it. But it's not the whole story.

What the smoother memory curve cost

The forkless save took 5 to 6 seconds against fork's steady 3. That's not a rounding difference — it's roughly 70% longer for the same data, every time I measured it, over seven runs each.

It also cost write throughput. Across the 12-second window surrounding each save, the eight writer threads delivered a mean of 27,202 operations a second under fork mode and 23,034 under forkless — about 15% less. The background thread doing the serialising in forkless mode isn't free; it competes with the main thread for the same lock and the same CPU core budget, for longer, and the client-facing throughput shows it.

So the trade is real on both sides: fork spikes memory hard and briefly; forkless keeps memory nearly flat but takes longer and leans on your write throughput the whole time it's doing it. Neither is free. If your host is memory-constrained, forkless is the one you want. If your host has memory to spare and your write path cares more about throughput than about a short memory bump, the fork default is doing less damage than its memory graph suggests.

The pause the release notes mention

The one documented caveat I could find for this feature was that if the main thread writes to a key the background serialising thread hasn't reached yet, that client's request briefly stalls while the key gets moved to the front of the queue. I went looking for that stall in my latency samples.

Across seven forkless runs, the worst single round-trip during the 12-second window ranged from 5.1ms to 14.2ms, with a mean of 7.2ms. Fork's worst case ranged from 11.5ms to 25.7ms, mean 20.1ms — consistently about three times rougher. But forkless wasn't perfectly smooth either: one run spiked to 14.2ms, almost twice its own average worst case, which is consistent with hitting that documented stall on an unlucky key. The feature narrows the tail. It doesn't flatten it.

What it refuses, and what else still forks

forkless-infrastructure-enabled cannot be turned on with CONFIG SET against a running server:

$ valkey-cli config set forkless-infrastructure-enabled yes
(error) ERR CONFIG SET failed (possibly related to argument
'forkless-infrastructure-enabled') - can't set immutable config
Enter fullscreen mode Exit fullscreen mode

It has to go on the command line or in the config file at startup, which means adopting this on an existing fleet needs a restart, not a hot config push. Trying to select the method without that flag fails too, with a message that at least tells you why:

$ valkey-cli config set bgsave-default-method forkless
(error) ERR CONFIG SET failed (possibly related to argument
'bgsave-default-method') - 'forkless' can only be selected when the
server was started with 'forkless-infrastructure-enabled yes'
Enter fullscreen mode Exit fullscreen mode

A second BGSAVE while one is already running is refused outright, in both modes, with ERR Background save already in progress — unsurprising, but worth confirming it isn't silently queued.

The more useful limitation is this: enabling forkless-infrastructure-enabled only changes how RDB snapshots are taken. I turned on appendonly on the forkless-enabled instance, which triggers Valkey's automatic initial AOF rewrite, and checked aof_last_cow_size afterwards. It came back at roughly 10MB, not zero. AOF rewrites still fork, forkless RDB setting or not. If your actual pain is AOF rewrite stalls rather than BGSAVE stalls, this feature does not touch that path at all.

Watching it happen

The INFO fields the release notes promised for tracking progress are genuinely there and genuinely populated mid-save, which I didn't expect to hold up as cleanly as it did:

$ valkey-cli info persistence | grep -E 'save_keys|remaining'
current_save_keys_processed:555092
current_save_keys_total:3000000
forkless_estimated_seconds_remaining:3
Enter fullscreen mode Exit fullscreen mode

Polled a second and a half after triggering a save on the same 3,000,000-key dataset, that's a believable in-flight number with a plausible estimate attached, not a placeholder. That's one documented claim that checked out exactly as described, which is worth saying plainly since the other numbers in this post are all pushing back on some part of the pitch.

What I got wrong on the way

My first version of the latency probe recorded each sample's offset by subtracting time.time() (wall clock) from a time.perf_counter() reading (a separate monotonic clock with an arbitrary zero point). The two aren't comparable, so every timestamp in my first batch of CSVs came out as a nonsensical negative number in the billions. The aggregate percentiles I'd already printed to stdout were unaffected, because those only used the latency values, not the timestamps — but the per-sample files were useless for lining a latency spike up against the exact moment BGSAVE started. I fixed it by recording one wall-clock epoch at the start of the run and adding monotonic deltas to it, then re-ran the affected trials. The lesson, again: check that your instrumentation's own clock is the one you think it is before trusting a timing result, even one that looks plausible.

Run it yourself

This is the full loop for one side of the comparison; swap the container flags to switch modes.

docker run -d --name vk-forkless -p 16379:6379 valkey/valkey:9.2.0-rc1 \
  --forkless-infrastructure-enabled yes --bgsave-default-method forkless --save ""

docker run -d --name vk-fork -p 16380:6379 valkey/valkey:9.2.0-rc1 --save ""

python3 - <<'EOF'
import socket, sys
def load(host, port, n=3_000_000, vsize=300):
    s = socket.create_connection((host, port))
    val = b'x' * vsize
    buf, batch = bytearray(), 2000
    for i in range(n):
        key = f'key:{i}'.encode()
        buf += b'*3\r\n$3\r\nSET\r\n$%d\r\n%s\r\n$%d\r\n%s\r\n' % (len(key), key, len(val), val)
        if (i + 1) % batch == 0:
            s.sendall(buf); buf = bytearray()
            got = 0
            while got < batch:
                got += s.recv(65536).count(b'+OK')
load('127.0.0.1', 16379); load('127.0.0.1', 16380)
EOF

docker exec vk-fork valkey-cli bgsave
docker exec vk-forkless valkey-cli bgsave
docker exec vk-fork valkey-cli info persistence | grep -E 'rdb_last_cow_size|rdb_last_bgsave_time_sec'
docker exec vk-forkless valkey-cli info persistence | grep -E 'rdb_last_cow_size|rdb_last_bgsave_time_sec'
Enter fullscreen mode Exit fullscreen mode

Running just this much against an idle dataset gave me an rdb_last_cow_size of about 12MB, nowhere near the 350MB in the table above. That gap is the whole point: copy-on-write only costs you something when pages are being written while the save runs, so to see the real effect you need the eight-thread write load from the full test, not this stripped-down version.

If you're running Valkey on a host where memory headroom during BGSAVE has actually bitten you — not hypothetically, but a real OOM or a real eviction storm triggered by a save — this is worth testing against your own dataset shape before 9.2 ships as stable. If your problem has instead been a save that takes too long and steals too much write throughput while it runs, measure that before switching the default, because this release doesn't make BGSAVE cheaper. It moves where the cost shows up.

Top comments (0)