DEV Community

Davi Orlandi
Davi Orlandi

Posted on

Redis Queues in Practice: Lists, Streams, and When Each One Wins

The first Redis queue I shipped looked almost too good to be true. We already had Redis for sessions. Someone pasted two commands into a design doc: LPUSH on the producer, BRPOP on the worker. Reviewers nodded. The first week felt like a small miracle. Jobs flowed. Dashboards stayed green. I remember telling a teammate that we had "solved" background work without introducing another broker.

Then a worker pod got OOM-killed mid-send. Redis had already removed the job from the list. The email never went out. Nobody noticed until a customer asked, days later, why the receipt never arrived. Over the next month we found a thin trail of vanished reconciliations: not catastrophic, just enough to make me distrust every green latency chart.

That is the personal stake behind this essay. Lists are wonderful and Lists are incomplete. Streams are more ceremony and Streams remember what you asked them to remember. Choosing between them is less about microbenchmarks and more about whether you can sleep when a pod dies at the wrong millisecond.

What BRPOP actually promises

BRPOP is honest about one thing: it removes the element and returns it in a single step. After that handoff, the job lives only in the worker's memory. If the process dies before the side effect finishes, Redis has already forgotten the work. That is at-most-once delivery. It is fine for warming a cache. It is the wrong contract for billing, outbound email, or anything a human will notice missing.

payload, err := rdb.BRPop(ctx, 5*time.Second, "jobs:email").Result()
if err != nil {
    return err
}
return sendEmail(payload[1]) // die here and the job is gone
Enter fullscreen mode Exit fullscreen mode

I used to treat that snippet as "simple Redis." It is simple. It is also a silent hole in the floor.

Making Lists honest without leaving Redis

You can keep Lists and still recover from crashes. The pattern is familiar once you have lived through the first lost job: move each item atomically into a processing list, finish the side effect, then remove it. A reaper returns stale processing entries when a worker disappears.

moved, err := rdb.BLMove(ctx,
    "jobs:email", "jobs:email:processing",
    "RIGHT", "LEFT", 5*time.Second).Result()
if err != nil {
    return err
}
if err := sendEmail(moved); err != nil {
    return err // leave it for the reaper, or push back to ready
}
return rdb.LRem(ctx, "jobs:email:processing", 1, moved).Err()
Enter fullscreen mode Exit fullscreen mode

This works. You now own heartbeats, per-worker keys, reaper timing, and awkward edge cases like identical payloads (LREM matches by value, so every job needs a unique id). For a single service with a moderate backlog and one worker pool, that trade can be rational. You stay on familiar commands. You accept that reliability is application code.

Streams as bookkeeping you stop reinventing

Streams change the story. XADD appends. A consumer group tracks deliveries in the Pending Entries List. XACK clears a pending entry. XAUTOCLAIM reassigns work that sat idle too long. The crash recovery you would otherwise write by hand lives closer to Redis itself.

streams, err := rdb.XReadGroup(ctx, &redis.XReadGroupArgs{
    Group:    "workers",
    Consumer: "worker-1",
    Streams:  []string{"jobs:email", ">"},
    Count:    10,
    Block:    5 * time.Second,
}).Result()

for _, msg := range streams[0].Messages {
    if err := handle(msg); err != nil {
        continue // no XACK; reclaim will retry
    }
    rdb.XAck(ctx, "jobs:email", "workers", msg.ID)
}
Enter fullscreen mode Exit fullscreen mode

The production habits that matter are not exotic. Cap retries, then copy poison messages to a dead-letter stream. Treat delivery as at-least-once and make handlers tolerate duplicates. Trim with MAXLEN ~ so the stream cannot grow forever. Alert on PEL size and oldest pending idle time. Those are the difference between "we use Streams" and "we understand Streams."

The comparison that actually decides the rewrite

Aspect List + BLMOVE + reaper Stream + consumer group
Crash recovery You build it PEL + XAUTOCLAIM
In-flight visibility Scan processing keys XPENDING
Replay / audit Gone after ack Until trimmed
Multiple consumer groups Awkward fan-out copies Native
Ops complexity Low until the reaper Medium

Throughput for small payloads is usually similar. Redis's single-threaded command path dominates. The decision is correctness tooling, not microseconds. If two subsystems need the same message independently, Streams with separate groups win immediately. If a worker crash after receive is acceptable, Lists may still be enough. If Redis is already the shared critical path for everything in the company, consider whether messaging deserves a dedicated broker while Redis stays on cache and sessions.

The gotchas I keep writing down

Trimming unacked stream entries with MAXLEN follows position, not acknowledgement. Size the bound above your largest backlog. Ack after the durable side effect, never before. Prefer stable consumer names; pod names as consumer ids litter the group until someone cleans them up. And please do not use Pub/Sub as a job queue. Offline consumers lose messages. Live fan-out is a different problem.

Closing

I still reach for Lists when the queue is small, owned by one service, and the team understands the processing-list pattern. I reach for Streams when disappearance is unacceptable, when multiple groups need the same log, or when I am tired of reinventing the reaper. The painful lesson from that first lost receipt was not "Redis is bad at queues." It was that a queue without acknowledgement is a delayed apology waiting for the right crash.

Top comments (0)