DEV Community

Geminate Solutions
Geminate Solutions

Posted on

Move Slow LLM Calls Off the Request Path with BullMQ in Node.js

If an LLM call can take 20 seconds or more, keep it out of your HTTP handler. Put the work on a BullMQ queue backed by Redis, return a job ID right away with a 202 Accepted, and let a separate worker make the slow call while the client gets progress updates over Server-Sent Events.

The rest of this post is working code for that setup. It covers retries with backoff, concurrency and rate limits, dead-letter handling, and sending progress to the browser.

Why the request path is the wrong place

An inline LLM call looks fine in development. Under real traffic, these things go wrong:

  • Timeouts stack up. Your load balancer, reverse proxy and client each have their own idle timeout. A long generation can go past one of them, and the user sees an error even though the model call worked.
  • Retries create duplicates. When a request times out, the browser or the user tries again. Now you pay for two generations and may write two results.
  • Provider limits hit everyone. A burst of traffic sends a burst of calls straight to the provider. The 429s come back to users as failures.
  • Deploys kill work in progress. Restarting the API process drops every generation that was still running.

With a queue, the API stays fast and the slow part has its own retry policy, its own concurrency cap and its own failure handling.

The shape

  1. The client sends POST /summaries. The API adds a job and returns { jobId }.
  2. A worker process picks up the job, calls the model and saves the result.
  3. The client opens GET /summaries/:jobId/events and receives progress until the job is done.
  4. Jobs that fail permanently go to a dead-letter queue so someone can look at them.

Install the packages:

npm install bullmq ioredis express
Enter fullscreen mode Exit fullscreen mode

Shared queue setup

// queue.js
import { Queue, QueueEvents } from 'bullmq';
import IORedis from 'ioredis';

export function makeConnection() {
  // BullMQ workers need maxRetriesPerRequest set to null
  return new IORedis(process.env.REDIS_URL, { maxRetriesPerRequest: null });
}

export const llmQueue = new Queue('llm-summaries', { connection: makeConnection() });
export const deadLetterQueue = new Queue('llm-dead-letter', { connection: makeConnection() });
export const llmEvents = new QueueEvents('llm-summaries', { connection: makeConnection() });
Enter fullscreen mode Exit fullscreen mode

Workers and QueueEvents use blocking Redis commands. Giving each one its own connection keeps them from getting in each other's way.

Enqueue and return immediately

// api.js
app.post('/summaries', async (req, res) => {
  const { documentId } = req.body;

  const job = await llmQueue.add(
    'summarize',
    { documentId },
    {
      jobId: 'summary-' + documentId, // same document, same job
      attempts: 5,
      backoff: { type: 'exponential', delay: 2000 },
      removeOnComplete: { age: 3600 },
      removeOnFail: { age: 24 * 3600 },
    }
  );

  res.status(202).json({ jobId: job.id });
});
Enter fullscreen mode Exit fullscreen mode

The custom jobId is the cheapest deduplication you will get. If a user clicks the button three times, BullMQ will not add a second job with the same ID while the first still exists. Choose the ID based on what makes the work unique, for example the document plus the prompt version.

With attempts: 5 and exponential backoff starting at 2 seconds, retries wait roughly 2, 4, 8 and 16 seconds. That spacing gives a rate-limited provider time to recover.

The worker: concurrency and rate limits

// worker.js
import { Worker, UnrecoverableError } from 'bullmq';
import { makeConnection, deadLetterQueue } from './queue.js';

const worker = new Worker(
  'llm-summaries',
  async (job) => {
    await job.updateProgress({ stage: 'loading', pct: 10 });
    const doc = await loadDocument(job.data.documentId);
    if (!doc) throw new UnrecoverableError('Document not found');

    await job.updateProgress({ stage: 'generating', pct: 30 });
    const summary = await callModel(doc.text, 60000);

    await job.updateProgress({ stage: 'saving', pct: 90 });
    await saveSummary(job.data.documentId, summary);

    return { documentId: job.data.documentId };
  },
  {
    connection: makeConnection(),
    concurrency: 8,
    limiter: { max: 50, duration: 60000 },
  }
);
Enter fullscreen mode Exit fullscreen mode

These two settings control different things:

  • concurrency is how many jobs one worker process runs at the same time. With 3 worker processes at a concurrency of 8, you have at most 24 model calls in flight.
  • limiter applies to the whole queue, across all workers. Here it allows 50 jobs per minute. Set it a bit under your provider quota so other services that share the key still have room.

If you mix cheap and expensive tasks, such as classification and long-form generation, give them separate queues. Otherwise a pile of slow jobs uses up all the slots and the quick jobs wait behind them.

Retry the right errors only

Retrying a malformed request five times wastes money and slows down the feedback. Sort errors into two groups at the point where you call the model:

async function callModel(text, timeoutMs) {
  try {
    // llm.summarize is a stand-in for your provider SDK call
    return await llm.summarize(text, { signal: AbortSignal.timeout(timeoutMs) });
  } catch (err) {
    const status = err.status ?? 0;
    if ([400, 401, 403, 404, 422].includes(status)) {
      throw new UnrecoverableError('Model rejected request: ' + err.message);
    }
    throw err; // 429, 5xx, network errors and timeouts: let BullMQ retry
  }
}
Enter fullscreen mode Exit fullscreen mode

UnrecoverableError tells BullMQ to skip the remaining attempts and fail the job right away. Use it for anything that will fail the same way on every attempt.

Always set your own timeout on the model call. If the provider hangs and nothing times out, that job keeps its concurrency slot forever.

Dead-letter handling

BullMQ has no built-in dead-letter queue, but it takes a few lines to add one. The worker's failed event fires on every failed attempt, so check whether this was the last one:

worker.on('failed', async (job, err) => {
  if (!job) return;
  const outOfAttempts = job.attemptsMade >= (job.opts.attempts ?? 1);
  const unrecoverable = err.name === 'UnrecoverableError';

  if (outOfAttempts || unrecoverable) {
    await deadLetterQueue.add('dead', {
      originalJobId: job.id,
      data: job.data,
      reason: err.message,
      attempts: job.attemptsMade,
      failedAt: Date.now(),
    });
  }
});
Enter fullscreen mode Exit fullscreen mode

A separate dead-letter queue gives you one place to alert on, inspect and replay from. Nobody has to search through the failed set of every queue.

One limitation: this handler runs inside the worker process. If the process crashes right after a final failure, the dead-letter write never happens. The job still stays in the failed set for 24 hours because of removeOnFail, so a scheduled sweep can fill in anything that was missed.

For replays, add a small admin script that reads from the dead-letter queue, fixes the cause and re-adds the job under its original ID.

Sending progress back to the client

Server-Sent Events work well here. They use plain HTTP, the browser reconnects on its own, and you only need data to flow one way.

app.get('/summaries/:jobId/events', async (req, res) => {
  const { jobId } = req.params;
  res.writeHead(200, {
    'Content-Type': 'text/event-stream',
    'Cache-Control': 'no-cache',
    Connection: 'keep-alive',
  });

  const send = (event, data) =>
    res.write('event: ' + event + '\ndata: ' + JSON.stringify(data) + '\n\n');

  let closed = false;
  const onProgress = ({ jobId: id, data }) => { if (id === jobId) send('progress', data); };
  const onCompleted = ({ jobId: id }) => {
    if (id === jobId) { send('done', { resultUrl: '/summaries/' + jobId }); cleanup(); }
  };
  const onFailed = ({ jobId: id, failedReason }) => {
    if (id === jobId) { send('failed', { reason: failedReason }); cleanup(); }
  };

  function cleanup() {
    if (closed) return;
    closed = true;
    llmEvents.off('progress', onProgress);
    llmEvents.off('completed', onCompleted);
    llmEvents.off('failed', onFailed);
    res.end();
  }

  llmEvents.on('progress', onProgress);
  llmEvents.on('completed', onCompleted);
  llmEvents.on('failed', onFailed);
  req.on('close', cleanup);

  // The job may have finished before the client connected
  const job = await llmQueue.getJob(jobId);
  const state = job ? await job.getState() : 'missing';
  if (state === 'completed') onCompleted({ jobId });
  else if (state === 'failed' || state === 'missing') onFailed({ jobId, failedReason: state });
});
Enter fullscreen mode Exit fullscreen mode

People often skip the state check at the end. Fast jobs can finish between the POST and the moment the browser opens the stream. Without the check, that client waits forever.

The done event sends a URL, not the result itself. That keeps large outputs out of Redis event payloads, and the client can fetch the result from your database with normal auth checks.

Each API process needs only one QueueEvents instance. Many open streams will add many listeners, so call llmEvents.setMaxListeners(0) or send events to clients through your own lookup keyed by job ID.

Production checklist

  • Handle SIGTERM with await worker.close() so in-flight jobs finish before a deploy.
  • Run workers as a separate process or container from the API, and scale them separately.
  • Turn on Redis persistence. Without it, a Redis restart loses every queued job.
  • Alert on dead-letter queue depth and on the wait time of the oldest job, not only on error counts.
  • Make saveSummary idempotent. A retry after a partial write should overwrite, not duplicate.
  • Log the job ID, attempt number and token usage on every model call.

When you do not need a queue

If the call usually finishes in a second or two, as with short classification, and the user is waiting on the screen anyway, streaming the response inline is simpler. Use a queue when calls are slow, happen in bursts, cost a lot to repeat, or must survive a deploy.

At Geminate Solutions we have built queue-heavy systems, including an exam platform that handled 10M+ requests per minute in production. Most of the problems came from the details above, such as retrying the wrong errors or missing fast completions, more than from the queue library.

For the wider picture of adding model calls to a Node.js backend, from provider choice to streaming, see the Node.js AI integration guide. Once these jobs are live, monitoring AI agents in production explains what to watch.

The full Node.js AI integration guide is there if you want to keep going.

Top comments (0)