DEV Community

Alex Georgiev
Alex Georgiev

Posted on AI-assisted

Valkey 9.2's INCREX command removes the crash window between INCR and EXPIRE

I run a lot of rate limiters and quota counters the same way: INCR the key, and if it comes back 1, set an expiry on it so the window resets later. It works, but it is two round trips to the database for something that is conceptually one operation, and there has always been a gap between those two calls where a process could die and leave a counter that never expires. Valkey 9.2.0-rc1 adds a command called INCREX that does both in a single call. I pulled the image, ran it against the pattern it replaces, and tried to find out how much of what I assumed about it was actually true.

Most of it was not. The headline speed story turns out to be mostly about round trips rather than the command being intrinsically faster, and the behaviour I ended up caring about most was not speed at all.

What I tested against what

I ran valkey/valkey:9.2.0-rc1, and for comparison the previous stable valkey/valkey:9.1.2, as plain Docker containers on the same host, talking to them from a Python script using redis-py 8.1.0 over TCP on localhost. This is a release candidate, not a final build: 9.2.0-rc1 is the newest tag Docker Hub has for the valkey/valkey image as of this test, and INCREX is not in 9.1.2 at all:

$ docker exec vk91 valkey-cli INCREX bar 10 EX 30
(error) ERR unknown command 'INCREX', with args beginning with: 'bar' '10' 'EX' '30'
Enter fullscreen mode Exit fullscreen mode

On 9.2.0-rc1 it exists, and COMMAND DOCS confirms when it landed, although the server's own INFO server reports its version as the plain 9.2.0, with no -rc1 suffix, so you cannot tell a candidate build from a final one by querying it:

$ docker exec vk92 valkey-cli COMMAND DOCS INCREX
summary  Increments the numeric value of a key with an option to expire.
         Uses 0 as initial value if the key doesn't exist.
since    9.2.0
Enter fullscreen mode Exit fullscreen mode

The grammar is INCREX key [NX|XX] [EX seconds|PX ms|EXAT ts|PXAT ts] [BYINT n|BYFLOAT n]. Leave out BYINT/BYFLOAT and it increments by 1.

The round-trip number, and why it is smaller than it looks

The pattern I actually write in application code is two separate, blocking round trips:

val = r.incr(key)
if val == 1:
    r.expire(key, 60)
Enter fullscreen mode Exit fullscreen mode

I ran 2,000 of these against 2,000 distinct keys, three times, and 2,000 calls to INCREX key EX 60 BYINT 1 against fresh keys, three times:

run old (INCR + EXPIRE), mean new (INCREX), mean speedup
1 0.487 ms 0.258 ms 1.89x
2 0.460 ms 0.289 ms 1.59x
3 0.572 ms 0.283 ms 2.02x

That is a real, repeatable 1.6x to 2x. But it is the saving from removing a round trip, not evidence that INCREX is a faster command on the server side. I checked that directly by pipelining the old pattern instead of sending it as two blocking calls:

pipe = r.pipeline(transaction=False)
pipe.incr(key)
pipe.expire(key, 60)
pipe.execute()
Enter fullscreen mode Exit fullscreen mode
run old, pipelined new (INCREX) ratio
1 0.325 ms 0.297 ms 1.09x
2 0.343 ms 0.307 ms 1.12x
3 0.361 ms 0.299 ms 1.21x

Once the old pattern stops waiting on the network twice, the gap drops to 9-21%. Anyone already pipelining this, or wrapping it in MULTI/EXEC, is not going to notice much of a speed difference. The 2x figure is real, but it is a tax on code that does not pipeline, and in my experience most rate-limiter code does not, because pipelining a conditional increment (if val == 1) is awkward enough that people skip it.

The actual reason to switch: nothing is left half-done

The round trip is not the part of this that matters. The gap between INCR and EXPIRE is a window where a request can time out, crash, or get killed by a deploy, leaving a key that was incremented but never got an expiry. I simulated that directly: 500 requests, each with a 2% chance of "crashing" immediately after the first command succeeds, before the second one is sent.

old: 500 requests, 11 simulated crashes, 489 completed, 11 keys stuck with no TTL
new: 500 requests, 11 simulated crashes, 489 completed, 0 keys stuck with no TTL
Enter fullscreen mode Exit fullscreen mode

Under the old pattern, every simulated crash leaves exactly one permanently orphaned key, because the increment already landed and nothing will ever set its expiry. Under INCREX, a crash either happens before the single command is sent, in which case nothing happened at all, or after it completed, in which case the key already has both its new value and its TTL. There is no state in between for a crash to land in.

I also ran 20 threads against one shared key, 100 increments each, to check INCREX does not drop updates under contention:

expected 2000, got 2000, ttl=30
Enter fullscreen mode Exit fullscreen mode

No lost updates, and the TTL set once by the first caller was left alone by the other 1,999. That is exactly what an atomic increment is supposed to do, so it was not a surprise, but it was worth confirming on a command this new rather than assuming it.

What it refuses, and one thing I had backwards

I assumed the NX/XX flags on INCREX meant what they mean on the standalone EXPIRE command: only touch the TTL if the key does, or does not, already have one. That assumption is wrong. INCREX follows SET's convention instead, where the condition is about the key's existence, not its TTL:

$ valkey-cli SET baz 5
OK
$ valkey-cli INCREX baz NX EX 100 BYINT 1
1) (integer) 5
2) (integer) 0
Enter fullscreen mode Exit fullscreen mode

baz already existed, so the NX condition failed and the whole operation was refused, increment included: the value stayed 5, no TTL appeared, and the second element of the reply, the amount actually applied, is 0. The same call against a key that does not exist yet succeeds and sets both:

$ valkey-cli INCREX newkey NX EX 50 BYINT 7
1) (integer) 7
2) (integer) 7
Enter fullscreen mode Exit fullscreen mode

If you are porting a rate limiter that currently uses EXPIRE ... NX to mean "do not reset a window that is already running," reaching for INCREX ... NX is the wrong instinct. What you actually want is no condition at all: I checked separately that calling INCREX with no expiration clause on a key that already has a TTL leaves that TTL exactly as it was, rather than clearing it. A key with 100 seconds left kept its 100 seconds after an increment-only call.

INCREX refuses the same inputs INCR refuses, with the same error text:

$ valkey-cli SET strval hello
OK
$ valkey-cli INCREX strval EX 30
(error) ERR value is not an integer or out of range
Enter fullscreen mode Exit fullscreen mode

A wrong-type key gets the usual WRONGTYPE Operation against a key holding the wrong kind of value, and asking for BYINT and BYFLOAT on the same call is a syntax error.

Here is the gap that bothered me. Push plain INCR past the signed 64-bit maximum and it refuses outright, with an error that names the problem:

$ valkey-cli INCR bignum2
(error) ERR increment or decrement would overflow
Enter fullscreen mode Exit fullscreen mode

Push INCREX to the same ceiling and nothing throws. The value just stops moving, and the only sign anything went wrong is a 0 sitting in the second slot of the reply:

$ valkey-cli SET bignum 9223372036854775807
OK
$ valkey-cli INCREX bignum EX 30 BYINT 1
1) (integer) 9223372036854775807
2) (integer) 0
Enter fullscreen mode Exit fullscreen mode

A try/except around INCR that catches the overflow error will not catch anything here, because INCREX never raises it. Whatever is calling this command has to look at that second number itself, every time, to know its counter is stuck.

Watching it in production

Monitoring does not need any new instrumentation to see this. INFO commandstats already tracks INCREX under its own key, apart from INCR and EXPIRE, so a dashboard built on the old counters will show whether a service has actually switched over just by adding one more line:

cmdstat_increx:calls=24710,usec=151237,usec_per_call=6.12,rejected_calls=0,failed_calls=4
cmdstat_incrby:calls=24700,usec=84758,usec_per_call=3.43,rejected_calls=0,failed_calls=0
cmdstat_expire:calls=24689,usec=119775,usec_per_call=4.85,rejected_calls=0,failed_calls=0
Enter fullscreen mode Exit fullscreen mode

The failed_calls=4 there is my own error tests from above, not a hidden problem with the command.

What I got wrong on the way

My first few calls failed with a plain syntax error, because I assumed the increment amount was a bare positional argument (INCREX foo 10 EX 30), which matches nothing in the real grammar; it needs the BYINT or BYFLOAT token in front of the number. I also lost more time than I should have on the NX test, because I read the (integer) 0 second reply element as a failure with no explanation, rather than as "zero was applied." It took checking the key's value and TTL directly, and comparing against a fresh key where NX did succeed, to see that the command had refused cleanly rather than broken.

Run it yourself

docker run -d --name vk92 -p 16392:6379 valkey/valkey:9.2.0-rc1
docker exec vk92 valkey-cli COMMAND DOCS INCREX
docker exec vk92 valkey-cli --no-raw INCREX somekey EX 30 BYINT 5
Enter fullscreen mode Exit fullscreen mode

The round-trip benchmark:

import redis, time, statistics

r = redis.Redis(host='127.0.0.1', port=16392, decode_responses=True)
N = 2000

def old_pattern():
    times = []
    for i in range(N):
        key = f'bench_old_{i}'
        t0 = time.perf_counter()
        val = r.incr(key)
        if val == 1:
            r.expire(key, 60)
        times.append(time.perf_counter() - t0)
    return times

def new_pattern():
    times = []
    for i in range(N):
        key = f'bench_new_{i}'
        t0 = time.perf_counter()
        r.execute_command('INCREX', key, 'EX', 60, 'BYINT', 1)
        times.append(time.perf_counter() - t0)
    return times

for name, fn in (('old', old_pattern), ('new', new_pattern)):
    times_ms = sorted(t * 1000 for t in fn())
    print(name, 'mean', statistics.mean(times_ms))
Enter fullscreen mode Exit fullscreen mode

And the crash simulation:

import redis, random

r = redis.Redis(host='127.0.0.1', port=16392, decode_responses=True)
N, CRASH_RATE = 500, 0.02

def run(pattern, seed):
    r.flushall()
    random.seed(seed)
    crashed = 0
    for i in range(N):
        key = f'rl_{i}'
        if pattern == 'old':
            r.incr(key)
            if random.random() < CRASH_RATE:
                crashed += 1
                continue
            r.expire(key, 60)
        else:
            if random.random() < CRASH_RATE:
                crashed += 1
                continue
            r.execute_command('INCREX', key, 'EX', 60, 'BYINT', 1)
    no_ttl = sum(1 for i in range(N) if r.exists(f'rl_{i}') and r.ttl(f'rl_{i}') == -1)
    return crashed, no_ttl

for pattern in ('old', 'new'):
    print(pattern, run(pattern, seed=42))
Enter fullscreen mode Exit fullscreen mode

What I'd do with this

If you have a counter that needs a TTL and you are writing it as two commands today, check first whether you already pipeline them. If you do, INCREX buys you correctness and one fewer thing to maintain, not much speed. If you do not pipeline, you get both. Either way, do not reach for NX or XX assuming they mean what they mean on EXPIRE, and if your counters can plausibly reach the signed 64-bit ceiling, check the second element of the reply rather than trusting an exception to tell you it stopped.

Top comments (0)