DEV Community

Lacey Glenn
Lacey Glenn

Posted on

Why Your Social App's Notification System Will Break at 10K Users (And How to Fix It Before It Happens)

If you've ever shipped a "like," "comment," or "follow" notification feature, you know how deceptively simple it looks in the demo. One user does something, another user gets pinged. Done.

Then you hit 10,000 users and everything falls apart at once: notifications arrive 20 minutes late, your push provider starts throttling you, your database CPU spikes every time someone posts, and your on-call engineer is debugging a queue that's somehow 400,000 jobs deep.

This isn't a hypothetical. It's the single most predictable failure pattern in social app backends, and it happens because notification systems are usually built for the happy path of one user, not the fan-out path of thousands. Here's exactly where it breaks, and how to fix each piece before it becomes a 2 a.m. incident.

Why 10K Users Specifically?

There's nothing magic about the number 10,000 — it's just where three things usually converge for the first time:

  • Your average "popular" user now has enough followers that a single action fans out to hundreds or thousands of notification events
  • Your traffic is high enough that naive polling or synchronous writes start queuing up
  • You're still on the architecture you built in week one, because it worked fine until now

Below are the five places that break first, roughly in the order teams discover them.


1. The Fan-Out Problem: One Action, Thousands of Writes

The naive version:

// DON'T DO THIS AT SCALE
async function onNewPost(post) {
  const followers = await db.query(
    'SELECT user_id FROM followers WHERE following_id = $1',
    [post.authorId]
  );

  for (const follower of followers) {
    await db.insert('notifications', {
      userId: follower.user_id,
      type: 'new_post',
      postId: post.id,
    });
    await pushService.send(follower.user_id, 'New post!');
  }
}
Enter fullscreen mode Exit fullscreen mode

This works beautifully for a user with 20 followers. For a user with 50,000 followers, this is 50,000 sequential database writes and 50,000 sequential HTTP calls to your push provider — inside a single request handler. Your API times out, your DB connection pool exhausts, and every other request queues behind it.

Why it breaks specifically at 10K users: early on, nobody has enough followers for fan-out to matter. Once a handful of accounts become "popular" relative to your user base, their actions become mini denial-of-service events against your own infrastructure.

The fix — push the fan-out off the request path:

// Producer: just enqueue, respond fast
async function onNewPost(post) {
  await queue.publish('fan-out-notifications', { postId: post.id });
  return; // API responds in milliseconds
}

// Consumer: does the heavy lifting asynchronously, in batches
async function fanOutWorker({ postId }) {
  const post = await db.getPost(postId);
  const followerIds = await db.getFollowerIds(post.authorId); // batched query

  const BATCH_SIZE = 500;
  for (let i = 0; i < followerIds.length; i += BATCH_SIZE) {
    const batch = followerIds.slice(i, i + BATCH_SIZE);
    await db.bulkInsertNotifications(batch.map(userId => ({
      userId, type: 'new_post', postId,
    })));
    await queue.publish('send-push-batch', { userIds: batch, postId });
  }
}
Enter fullscreen mode Exit fullscreen mode

The API request now does one cheap thing (publish a message) instead of one expensive thing (fan out synchronously). Everything expensive happens in a worker that can scale horizontally and retry independently of your user-facing traffic.


2. The N+1 Query You Don't Notice Until It's Too Late

Fetching a user's device tokens one row at a time is invisible at low volume and catastrophic at scale:

// N+1 pattern — one query per follower
for (const follower of followers) {
  const tokens = await db.query(
    'SELECT device_token FROM devices WHERE user_id = $1',
    [follower.user_id]
  );
}
Enter fullscreen mode Exit fullscreen mode

At 500 followers, that's 500 round-trips to the database for a single post. Multiply by concurrent posts and you've turned your database into the bottleneck for a feature that has nothing to do with your core data model.

The fix: batch it.

const tokens = await db.query(
  'SELECT user_id, device_token FROM devices WHERE user_id = ANY($1)',
  [followerIds]
);
Enter fullscreen mode Exit fullscreen mode

One query, one round-trip, regardless of whether you're notifying 10 people or 10,000.


3. Push Provider Rate Limits You Didn't Know Existed

FCM and APNs both have throughput limits, and they will silently drop or delay your notifications if you slam them with individual requests instead of using their batch/multicast APIs.

// Inefficient: one HTTP call per device
for (const token of tokens) {
  await fcm.send({ token, notification: { title: 'New post!' } });
}
Enter fullscreen mode Exit fullscreen mode

At scale this hits rate limits, and worse, a single slow or failed call blocks the loop if you're not handling concurrency correctly.

The fix: use multicast/batch sending, and respect the provider's batch size limits (FCM caps multicast at 500 tokens per call):

const BATCH_SIZE = 500;
for (let i = 0; i < tokens.length; i += BATCH_SIZE) {
  const batch = tokens.slice(i, i + BATCH_SIZE);
  const response = await fcm.sendMulticast({
    tokens: batch,
    notification: { title: 'New post!' },
  });

  // Critical: prune dead tokens so you stop paying this cost forever
  response.responses.forEach((res, idx) => {
    if (!res.success && res.error?.code === 'messaging/registration-token-not-registered') {
      deadTokens.push(batch[idx]);
    }
  });
}
await db.deleteDeviceTokens(deadTokens);
Enter fullscreen mode Exit fullscreen mode

That last part — pruning dead tokens — is the piece almost every team skips, and it's why notification volume keeps growing even as engagement doesn't: you're paying to "notify" phones that were factory-reset eight months ago.


4. No Backpressure: The Queue That Ate Your Server

Once you move to a queue (good), the next failure mode is not thinking about backpressure (bad). If your worker can process 200 jobs/second but a viral post generates 50,000 fan-out jobs instantly, you need the queue to absorb that burst — not your worker process to fall over trying to drain it all at once.

Symptoms this is happening to you:

  • Memory usage on worker nodes climbs steadily and never recovers
  • Notifications for older events "jump the queue" in front of newer, more relevant ones
  • One viral post delays every other user's notifications by 30+ minutes

The fix — bound your concurrency and prioritize:

// Using a library like BullMQ
const worker = new Worker('send-push-batch', processJob, {
  concurrency: 20,          // bounded, not unlimited
  limiter: {
    max: 1000,
    duration: 1000,         // 1000 jobs/sec ceiling, matches provider limits
  },
});

// Separate high-priority queue for time-sensitive notifications
// (DMs, mentions) so they don't sit behind a mass fan-out
await queue.add('send-push-batch', data, { priority: isDirectMessage ? 1 : 10 });
Enter fullscreen mode Exit fullscreen mode

Backpressure isn't a nice-to-have at scale — it's the difference between "the queue is a bit backed up" and "the queue took down the box it runs on."


5. Real-Time Delivery: The WebSocket Server That Can't Scale Out

If you're pushing live in-app notifications over WebSockets, there's a subtler trap: a naive implementation keeps each user's connection and the logic to notify them on the same server process, in memory.

// Breaks the moment you run more than one server instance
const connectedUsers = new Map(); // userId -> socket

function notifyUser(userId, payload) {
  const socket = connectedUsers.get(userId);
  if (socket) socket.emit('notification', payload);
}
Enter fullscreen mode Exit fullscreen mode

This works great with one server and falls apart the moment you need two, because the user your worker wants to notify might be connected to a different instance that has no idea about that in-memory map.

The fix — decouple "who's connected" from "who's sending":

// Each server subscribes to a shared pub/sub channel (Redis, NATS, etc.)
redisSub.subscribe('notifications');

redisSub.on('message', (channel, message) => {
  const { userId, payload } = JSON.parse(message);
  const socket = connectedUsers.get(userId); // only checks local connections
  if (socket) socket.emit('notification', payload);
});

// Any server, anywhere, can trigger a notification
function notifyUser(userId, payload) {
  redisPub.publish('notifications', JSON.stringify({ userId, payload }));
}
Enter fullscreen mode Exit fullscreen mode

Now it doesn't matter which server instance holds the user's socket — every instance is listening, and only the one holding the actual connection acts on it. This is what lets you add a second, third, or tenth WebSocket server without a rewrite.


The Architecture That Actually Scales

Putting it together, a notification system that survives past 10K users generally looks like this:

[API request] --> [Queue: fan-out-notifications]
                          |
                          v
                 [Worker: batch DB writes]
                          |
                          v
        +-----------------+------------------+
        |                                     |
        v                                     v
[Queue: send-push-batch]           [Pub/Sub: realtime-notifications]
        |                                     |
        v                                     v
[Worker: batched provider calls]   [WebSocket servers, any instance]
        |
        v
[Prune dead tokens back to DB]
Enter fullscreen mode Exit fullscreen mode

The common thread across every fix above: get expensive work off the request path, batch everything that touches an external system, and never assume state lives on just one server.

The Checklist Before You Hit 10K Users

  • [ ] Fan-out happens in a background worker, not inside the API request
  • [ ] Follower/device-token lookups are batched, never looped one row at a time
  • [ ] Push sends use the provider's multicast/batch endpoints, respecting size limits
  • [ ] Failed/dead device tokens are pruned automatically, not left to accumulate forever
  • [ ] Your queue has bounded concurrency and a rate limiter matching your provider's real limits
  • [ ] Time-sensitive notifications (DMs, mentions) are prioritized over mass fan-out events
  • [ ] Real-time delivery uses pub/sub so it works across multiple server instances, not just one

None of these fixes are exotic — they're all standard distributed-systems patterns. The trick is applying them before your first viral post, not during the incident review after it.


Have you hit a different notification scaling wall — email digest storms, SMS costs exploding, or something else entirely? I'd love to hear what broke first for you in the comments.

Top comments (0)