DEV Community

subashthiruppathy
subashthiruppathy

Posted on

Why Your BullMQ / Redis Job Queue is Silently Leaking Memory (And How to Fix It)

It starts with a Slack alert at 3 AM: Redis memory at 85%.

You check your queue dashboard. Zero waiting jobs. Zero active jobs. Workers are healthy. Throughput looks normal. And yet used_memory has been climbing a little every day for weeks, like a slow tire leak.

If you run BullMQ in production, there's a good chance you've been here, or you're about to be. BullMQ is excellent, but its defaults favor debuggability over memory hygiene. Left unconfigured, it will happily keep every job you ever ran.

In this post we'll go through the six most common causes of "silent" memory growth in BullMQ and Redis, how to confirm which one is biting you, and how to fix each.


TL;DR

# Cause Fix
1 Completed/failed jobs are kept forever (default) removeOnComplete / removeOnFail with age + count
2 Huge job payloads and return values Store references, not blobs
3 Orphaned repeatable jobs Use job schedulers, clean up stale keys
4 Unbounded event streams Cap streams.events.maxLen
5 Creating Queue / QueueEvents instances per request Create once, reuse, close on shutdown
6 Wrong Redis eviction policy or fragmentation maxmemory-policy noeviction, monitor RSS vs used memory

1. The big one: jobs are never removed by default

This is the cause behind the majority of "my Redis is full of nothing" reports.

By default, BullMQ keeps completed and failed jobs in Redis indefinitely. Each job is stored as a hash (with its data, options, return value, stack traces, timestamps) plus entries in sorted sets for the completed or failed state.

If you process 50,000 jobs a day at ~2 KB each, that's roughly 100 MB a day of data nobody will ever read again.

Confirm it

import { Queue } from 'bullmq';

const queue = new Queue('emails', { connection });
console.log(await queue.getJobCounts('completed', 'failed', 'waiting', 'active'));
// { completed: 4_812_903, failed: 31_220, waiting: 0, active: 3 }
Enter fullscreen mode Exit fullscreen mode

There's your leak. Nearly 5 million completed jobs sitting in Redis.

Fix it

Set retention at the queue level so every job inherits it:

const queue = new Queue('emails', {
  connection,
  defaultJobOptions: {
    removeOnComplete: {
      age: 60 * 60,   // keep for 1 hour (seconds)
      count: 1000,    // ...but never more than 1000
    },
    removeOnFail: {
      age: 7 * 24 * 60 * 60, // keep failures for 7 days for debugging
      count: 5000,
    },
  },
});
Enter fullscreen mode Exit fullscreen mode

A few notes:

  • removeOnComplete: true deletes the job immediately. Great for fire-and-forget work, but you lose the ability to inspect results.
  • A number (removeOnComplete: 100) keeps the latest N jobs.
  • The object form ({ age, count }) is usually what you want: bounded by time and count.
  • Keep failed jobs longer than completed ones. That's the data you actually debug with. But still bound them.

Clean up the existing mess

Setting the option only affects future jobs. To purge the backlog:

// Remove completed jobs older than 1 hour, up to 10,000 at a time
let removed;
do {
  removed = await queue.clean(60 * 60 * 1000, 10_000, 'completed');
} while (removed.length > 0);

// Same for failed jobs older than 7 days
do {
  removed = await queue.clean(7 * 24 * 60 * 60 * 1000, 10_000, 'failed');
} while (removed.length > 0);
Enter fullscreen mode Exit fullscreen mode

⚠️ Always pass a limit and loop. Cleaning millions of jobs in one call blocks Redis, and Redis is single-threaded. Do it in batches, ideally off-peak.


2. Fat payloads and return values

Even with retention configured, what you put in a job matters. BullMQ stores job.data and the worker's return value (returnvalue) in Redis. Developers often do this:

// ❌ Don't do this
await queue.add('generate-report', {
  userId,
  rows: hugeArrayOfTenThousandRows,
});

// and in the worker
async function processor(job) {
  const pdf = await buildPdf(job.data.rows);
  return { pdfBase64: pdf.toString('base64') }; // 5 MB return value in Redis
}
Enter fullscreen mode Exit fullscreen mode

A single 5 MB return value multiplied by a few thousand retained jobs is gigabytes.

Fix it: pass references, not payloads

// ✅ Store the big thing elsewhere (S3, Postgres, disk)...
const key = await storage.put(`reports/${reportId}/input.json`, rows);

// ...and put only the pointer in the job
await queue.add('generate-report', { userId, inputKey: key });

// Worker returns a pointer too
async function processor(job) {
  const rows = await storage.get(job.data.inputKey);
  const pdfKey = await storage.put(`reports/${job.id}.pdf`, await buildPdf(rows));
  return { pdfKey };
}
Enter fullscreen mode Exit fullscreen mode

Rule of thumb: if a job's data or result is bigger than a few KB, it probably belongs somewhere else.

Also watch out for:

  • Stack traces on failures. Every failed attempt appends a stack trace to the job hash. A job with attempts: 10 and a deep stack trace can get surprisingly big. Bound failed job retention (see #1).
  • job.log() abuse. Logs are stored in a Redis list per job. Logging every row you process adds up fast. Use keepLogs inside removeOnComplete / removeOnFail to cap how many log lines are retained:
removeOnComplete: { age: 3600, count: 1000, keepLogs: 20 }
Enter fullscreen mode Exit fullscreen mode

3. Orphaned repeatable jobs

Repeatable jobs (cron-like) are a sneaky source of leaks. In older BullMQ patterns, a repeatable job is identified by a key derived from its name, repeat pattern, and job ID. Change any of those between deploys and BullMQ treats it as a brand-new repeatable job, while the old one keeps firing and enqueuing forever.

Typical story: you tweak a cron from */5 * * * * to */10 * * * *, redeploy, and now both schedules are active. Do that a few times and you have a pile of zombie schedulers each producing jobs.

Confirm it

const repeatables = await queue.getRepeatableJobs();
console.log(repeatables.length);
console.table(repeatables.map(r => ({ name: r.name, pattern: r.pattern, key: r.key })));
Enter fullscreen mode Exit fullscreen mode

If you see multiple entries for the same logical job, that's your culprit.

Fix it

Remove the stale ones:

for (const r of await queue.getRepeatableJobs()) {
  if (isStale(r)) await queue.removeRepeatableByKey(r.key);
}
Enter fullscreen mode Exit fullscreen mode

For the long term, use Job Schedulers (available in recent BullMQ v5 releases), which are keyed by a stable scheduler ID you choose, so redeploying with a new pattern updates the schedule instead of creating a second one:

await queue.upsertJobScheduler(
  'daily-digest',                    // stable ID
  { pattern: '0 8 * * *' },          // change this freely between deploys
  { name: 'send-digest', data: {} },
);
Enter fullscreen mode Exit fullscreen mode

Check the BullMQ docs for your installed version, since the API is still evolving.


4. Unbounded event streams

BullMQ publishes queue events (completed, failed, progress, and so on) to a Redis Stream at bull:<queue>:events. This is how QueueEvents works. Streams don't expire on their own, so BullMQ trims them to approximately a max length, which defaults to 10,000 entries.

That default is fine for a quiet queue. For a high-throughput queue, 10k events per queue can still be a noticeable chunk of memory (event payloads include job IDs and, for completed, the return value), and if you have many queues it multiplies.

Fix it

If you don't consume events, or only need recent ones, lower the cap:

const queue = new Queue('emails', {
  connection,
  streams: {
    events: { maxLen: 1000 },
  },
});
Enter fullscreen mode Exit fullscreen mode

If you don't use QueueEvents at all and only rely on workers, a small maxLen is a free win.


5. Leaks on the Node.js side: instances created per request

Not every leak is in Redis. A very common bug in Express/Fastify/Nest apps:

// ❌ A new Queue (and a new Redis connection) on EVERY request
app.post('/signup', async (req, res) => {
  const queue = new Queue('emails', { connection: { host, port } });
  await queue.add('welcome', { email: req.body.email });
  res.sendStatus(202);
});
Enter fullscreen mode Exit fullscreen mode

Every new Queue() opens Redis connections and registers listeners. Do this under load and you'll see climbing Node heap, climbing Redis connection count (INFO clients), and eventually MaxListenersExceededWarning or ECONNRESET errors.

The same goes for QueueEvents and Worker instances created inside handlers or loops.

Fix it: create once, share, and close on shutdown

// queues.ts
import IORedis from 'ioredis';
import { Queue } from 'bullmq';

// maxRetriesPerRequest: null is required for Workers
export const connection = new IORedis(process.env.REDIS_URL!, {
  maxRetriesPerRequest: null,
});

export const emailQueue = new Queue('emails', { connection });
Enter fullscreen mode Exit fullscreen mode
// graceful shutdown
process.on('SIGTERM', async () => {
  await worker.close();      // waits for active jobs to finish
  await emailQueue.close();
  await queueEvents.close();
  await connection.quit();
  process.exit(0);
});
Enter fullscreen mode Exit fullscreen mode

Other Node-side things to check:

  • Attaching listeners inside the processor. queueEvents.on('completed', ...) inside a job handler adds a new listener on every job.
  • Closures capturing job.data. Holding references to big job objects in module-level caches or arrays prevents garbage collection.
  • Not closing workers in tests. Open handles pile up and hide real leaks.

6. Redis config: eviction policy and fragmentation

Two last Redis-level gotchas.

Use noeviction

BullMQ requires Redis to be configured with:

maxmemory-policy noeviction
Enter fullscreen mode Exit fullscreen mode

If you use an eviction policy like allkeys-lru, Redis can evict BullMQ's internal keys when memory is tight, causing missing jobs and corrupted queue state. A leak that should have been loud (writes failing) instead becomes silent data loss. Better to fail loudly: with noeviction, Redis returns errors on writes when full, and your alerts fire.

Watch fragmentation

Deleting millions of jobs frees memory logically, but Redis may not return it to the OS. Compare:

redis-cli INFO memory | grep -E "used_memory_human|used_memory_rss_human|mem_fragmentation_ratio"
Enter fullscreen mode Exit fullscreen mode

If mem_fragmentation_ratio is well above ~1.5 after a big cleanup, you're holding memory the allocator hasn't released. Options:

  • Enable active defragmentation (activedefrag yes) on supported builds
  • Restart Redis during a maintenance window (or fail over to a replica)

A diagnostic playbook

When Redis memory is growing and you don't know why, work through this in order.

1. Find the big keys

redis-cli --bigkeys
# or sample safely with a memory-aware scan
redis-cli --memkeys
Enter fullscreen mode Exit fullscreen mode

2. Count keys by queue (use SCAN, never KEYS * in production)

redis-cli --scan --pattern 'bull:emails:*' | wc -l
Enter fullscreen mode Exit fullscreen mode

3. Inspect job counts per state

console.log(await queue.getJobCounts());
Enter fullscreen mode Exit fullscreen mode

4. Spot-check a single job's size

redis-cli MEMORY USAGE bull:emails:12345
Enter fullscreen mode Exit fullscreen mode

5. Check the event stream size

redis-cli XLEN bull:emails:events
Enter fullscreen mode Exit fullscreen mode

6. Check connection count

redis-cli INFO clients | grep connected_clients
Enter fullscreen mode Exit fullscreen mode

Match the symptom to the cause:

  • Huge completed/failed counts → #1
  • A few enormous job hashes → #2
  • Many repeatable entries → #3
  • Large events stream → #4
  • Rising connected_clients or Node heap → #5
  • RSS far above used_memory → #6

Make it hard to regress

Fixing it once isn't enough. Put guardrails in place:

  1. Set defaultJobOptions on every queue, ideally via a shared factory so nobody forgets:
export function createQueue(name: string) {
  return new Queue(name, {
    connection,
    defaultJobOptions: {
      removeOnComplete: { age: 3600, count: 1000 },
      removeOnFail: { age: 7 * 24 * 3600, count: 5000 },
      attempts: 3,
      backoff: { type: 'exponential', delay: 2000 },
    },
    streams: { events: { maxLen: 1000 } },
  });
}
Enter fullscreen mode Exit fullscreen mode
  1. Alert on Redis memory growth rate, not just the absolute threshold. A steady upward slope catches leaks weeks before a full-memory outage.
  2. Export queue metrics (job counts by state) to Prometheus or your APM and graph completed and failed over time. They should plateau, not climb.
  3. Add a lint/test rule that fails CI if a queue is created without retention options.

Wrapping up

BullMQ isn't broken. It's just conservative: it assumes you want to keep everything until you say otherwise. In development that's a feature. In production, at scale, it's a slow-motion incident waiting to happen.

If you take only one thing from this post, make it this: set removeOnComplete and removeOnFail with both age and count on every queue. That alone resolves the vast majority of cases.

Then audit payload sizes, repeatable jobs, and your Node-side instance lifecycle, and make sure Redis is on noeviction with fragmentation monitored.

Have you been bitten by a BullMQ memory leak that isn't on this list? Drop it in the comments. I'd love to add it.


Further reading: the BullMQ docs sections on "Auto-removal of jobs", "Job Schedulers", and "Going to production".

Top comments (2)

Collapse
 
kashif_manzer profile image
Kashif Manzer •

Good list. One more Node-side trap worth adding to section 5: in serverless setups, QueueEvents instances and listeners created in a module initializer run on every cold start, and each one opens a blocking Redis connection that never closes. The queue looks healthy while connected_clients climbs, which is why step 6 of your playbook deserves to run first there. Seconding the advice to alert on the growth rate rather than an absolute threshold. That caught our leak weeks before a static limit would have fired.

Collapse
 
subashthiruppathy_5e0f532 profile image
subashthiruppathy •

Thanks Kashif, great addition. That's a nasty one because everything on the queue side looks fine while connections pile up. In serverless I'd keep the function producer-only (a plain Queue with a shared, reused connection) and run QueueEvents and Worker in a long-lived process instead. If you do need them in a function, cache the instances on the module or global scope so warm invocations reuse them, and check connected_clients first when debugging. Glad the growth-rate alert caught your leak early, and I'll fold this into section 5.