DEV Community

snehaup1997
snehaup1997

Posted on

Your Queue Is Just Borrowed Time

When a system starts receiving more work than it can process, we usually reach for a queue. Instead of making the producer wait, we put the work somewhere safe and let consumers catch up. It feels like the obvious solution and, for a while, everything looks fine.

The API stays responsive. Consumers keep processing. It feels like success and I think this is where we sometimes misunderstand what a queue actually gives us.

A queue doesn't create processing capacity. It converts a rate mismatch into waiting time. And waiting time is not free. Sounds obvious, but it has interesting consequences once you look at what happens to that waiting work.

Consider a service that normally processes:

8,000 events / second
Enter fullscreen mode Exit fullscreen mode

Suddenly, traffic increases to:

9,000 events / second
Enter fullscreen mode Exit fullscreen mode

The queue starts growing. Nothing is obviously broken yet. The API is still responding. Consumers are running. The queue is accepting messages. The team notices the growing backlog and increases consumer concurrency.

The consumers depend on a database that is already close to its limit. More workers create more database contention. Database latency increases, consumer throughput falls. Retries start appearing and those retries create even more work. The system now looks like this:

Traffic increase
      ↓
Queue backlog
      ↓
Database pressure
      ↓
Consumer throughput ↓
      ↓
More backlog
      ↓
More workers
      ↓
More database contention
      ↓
More failures
      ↓
More retries
      ↓
Even more work
Enter fullscreen mode Exit fullscreen mode

The queue wasn't the original problem. It was where the capacity problem became visible. This is why I don't think "the queue is growing" is enough information to decide what to do next.

The useful question is:

Why is the queue growing?

Consider a simpler system:

Producer
   ↓
 Queue
   ↓
Workers
   ↓
Database
Enter fullscreen mode Exit fullscreen mode

Suppose producers send:

10,000 messages / second
Enter fullscreen mode Exit fullscreen mode

while consumers can process:

7,000 messages / second
Enter fullscreen mode Exit fullscreen mode

The deficit is:

10,000 - 7,000 = 3,000 messages / second
Enter fullscreen mode Exit fullscreen mode

So the queue grows by roughly 3,000 messages every second. After five minutes:

3,000 × 300 = 900,000 messages
Enter fullscreen mode Exit fullscreen mode

The queue has protected the consumers from immediate overload. But nothing has changed about the underlying capacity mismatch. If traffic falls below processing capacity, the backlog can drain. If it doesn't, the queue keeps accumulating work. This gives us two questions when designing a queue:

How much burst can we absorb?
How much waiting can the workload tolerate?
Enter fullscreen mode Exit fullscreen mode

The second question is easy to miss.
A queue containing 100,000 messages doesn't tell us enough by itself. We also need to know how quickly the backlog is being processed and how old the waiting work has become.

Queue Depth Is Not the Same as Queue Health. Little's Law gives us a useful relationship:

L = λW
Enter fullscreen mode Exit fullscreen mode

where:

L = average number of items in the system
λ = throughput
W = average time an item spends in the system
Enter fullscreen mode Exit fullscreen mode

So:

W = L / λ
Enter fullscreen mode Exit fullscreen mode

For example, if a system has approximately:

50,000 messages in the system
5,000 messages / second of throughput
Enter fullscreen mode Exit fullscreen mode

then, under steady-state conditions:

50,000 / 5,000 = 10 seconds
Enter fullscreen mode Exit fullscreen mode

That represents roughly 10 seconds of work at the current throughput. It is a useful mental model, not a precise prediction of when a queue will empty. If arrivals continue to exceed processing capacity, the backlog will keep growing. This is why I would monitor queue age alongside queue depth.

A large backlog that is draining quickly can be temporary. A smaller backlog that keeps getting older can be much more concerning.

For example:

P50 message age:    1 second
P95 message age:    4 seconds
P99 message age:   12 seconds
Max age:           10 minutes
Enter fullscreen mode Exit fullscreen mode

is a very different situation from:

P50 message age:    8 seconds
P95 message age:   90 seconds
P99 message age:    5 minutes
Max age:            10 minutes
Enter fullscreen mode Exit fullscreen mode

Both have a ten-minute-old message. Only the second suggests that a large portion of the workload is experiencing significant delay.

Suppose consumers can sustainably process:

8,000 messages / second
Enter fullscreen mode Exit fullscreen mode

and normal traffic is:

7,600 messages / second
Enter fullscreen mode Exit fullscreen mode

The system is already operating at:

7,600 / 8,000 = 95%
Enter fullscreen mode Exit fullscreen mode

Technically, there is capacity remaining. Operationally, there isn't much room for reality. Traffic and processing time vary. Databases slow down, networks experience latency, retries appear, dependencies occasionally fail. So capacity isn't simply:

Arrival rate < Processing rate
Enter fullscreen mode Exit fullscreen mode

It is also about how much headroom exists before normal variation starts creating backlog. A queue can absorb that variation. But only for a while.

Suppose:

Normal traffic:      5K/s
Burst traffic:      10K/s
Consumer capacity:   8K/s
Enter fullscreen mode Exit fullscreen mode

During the burst, the queue absorbs:

10K - 8K = 2K/s
Enter fullscreen mode Exit fullscreen mode

If the burst lasts 60 seconds:

2K × 60 = 120K messages
Enter fullscreen mode Exit fullscreen mode

Now suppose traffic falls to 6K/s while consumers can process 10K/s. The backlog drains at:

10K - 6K = 4K/s
Enter fullscreen mode Exit fullscreen mode

So recovery takes:

120K / 4K = 30 seconds
Enter fullscreen mode Exit fullscreen mode

This gives us two different properties:

  • Burst tolerance — how much temporary excess work the queue can absorb.
  • Recovery capacity — how quickly the system can eliminate that backlog afterward.

Both belong in capacity planning. A system that survives a spike but takes ten minutes to recover has a very different operational profile from one that clears the backlog in thirty seconds.

More Consumers Are Not Always More Capacity. Suppose the queue is growing, so we increase workers:

50 workers
   ↓
100 workers
   ↓
200 workers
Enter fullscreen mode Exit fullscreen mode

It sounds reasonable. But imagine the workers all depend on:

Database
   ↓
50 connections
Enter fullscreen mode Exit fullscreen mode

At some point, the database becomes the bottleneck. Adding workers doesn't create more database capacity. It can actually make things worse through connection contention, lock contention, CPU pressure, or increased I/O. Real pipeline looks more like:

Queue
   ↓
Consumer Pool
   ↓
Database
   ↓
External API
Enter fullscreen mode Exit fullscreen mode

The sustainable throughput is constrained by the bottleneck in that path. So when a queue grows, I would ask:

Are consumers under-provisioned?
Or are consumers waiting on something else?
Enter fullscreen mode Exit fullscreen mode

That distinction matters. The queue may simply be the place where a downstream bottleneck becomes visible.

Retries can turn a problem into a feedback loop because they can increase load precisely when the system is already struggling.

Suppose a downstream service starts timing out:

Message
   ↓
Timeout
   ↓
Retry
   ↓
Timeout
   ↓
Retry
Enter fullscreen mode Exit fullscreen mode

Now the system isn't just processing the original workload.

It is processing the original workload plus recovery work.

That can create:

Higher latency
      ↓
More timeouts
      ↓
More retries
      ↓
More work
      ↓
More pressure
      ↓
Higher latency
Enter fullscreen mode Exit fullscreen mode

Retries therefore need their own controls:

  • exponential backoff;
  • jitter;
  • maximum attempts;
  • retry budgets;
  • dead-letter handling;
  • idempotent operations.

A retry mechanism should recover from transient failures without becoming another source of load.

This is another reason why queue depth alone isn't enough.

A queue might be growing because traffic increased.

Or because consumers slowed down.

Or because retries doubled the effective workload.

The metrics need to tell us which one.

Not All Queued Work Has the Same Value

There is another question I think queue designs should ask:

Does every message still have value by the time we process it?

Imagine a queue containing:

Update product 123 → $10
Update product 123 → $11
Update product 123 → $12
Update product 123 → $13
Enter fullscreen mode Exit fullscreen mode

If the only thing that matters is the current price, processing all four events may be unnecessary.

We could potentially coalesce them:

$10
$11
$12
$13
 ↓
$13
Enter fullscreen mode Exit fullscreen mode

Now the system has reduced four units of work to one.

We're not processing the backlog faster.

We're deciding that some of the backlog no longer needs to be processed.

The same idea can apply to:

  • repeated cache invalidations;
  • rapidly changing user preferences;
  • superseded state updates;
  • duplicate notifications;
  • refresh requests;
  • intermediate synchronization events.

The exact semantics depend on the application. Some events absolutely cannot be discarded or coalesced.

But where the business semantics allow it, eliminating stale work can be more effective than simply adding processing capacity.

Backpressure Is the Other Half

A queue becomes much more useful when it is paired with backpressure.

Instead of:

Producer
   ↓
Unlimited queue
   ↓
Workers
Enter fullscreen mode Exit fullscreen mode

I prefer:

Producer
   ↓
Admission control
   ↓
Bounded queue
   ↓
Workers
   ↓
Downstream
Enter fullscreen mode Exit fullscreen mode

When downstream capacity falls, that constraint should eventually reach the producer.

Depending on the system, that might mean:

  • rate limiting;
  • concurrency limits;
  • delayed retries;
  • HTTP 429 responses;
  • load shedding;
  • priority-based admission.

The implementation varies, but the principle is simple:

When the system cannot safely accept more work, it needs a mechanism to say no.

An unbounded queue can hide overload until the system runs out of memory, storage, or patience.

A bounded queue makes the constraint visible.

Not All Work Should Be Treated Equally

Once admission control exists, another question appears:

Which work should be protected when capacity runs out?

Imagine one queue contains:

Password reset
Payment processing
Analytics events
Search indexing
Report generation
Enter fullscreen mode Exit fullscreen mode

During normal operation, treating them equally may be fine.

During severe overload, it may not be.

A more deliberate architecture could separate workloads:

                 Incoming Work
                       │
          ┌────────────┼────────────┐
          ↓            ↓            ↓
       Critical      Normal        Bulk
          │            │            │
       Queue A       Queue B      Queue C
Enter fullscreen mode Exit fullscreen mode

Different workloads can then have different:

  • capacity limits;
  • priorities;
  • retry policies;
  • SLOs;
  • scaling rules.

This means overload behavior is designed rather than discovered in production.

Overload Should Be Deliberate

I don't want production systems to have only two states:

Healthy
   ↓
Broken
Enter fullscreen mode Exit fullscreen mode

There should be intermediate behavior:

Normal
  ↓
Queue absorbs burst
  ↓
Queue approaching limit
  ↓
Throttle
  ↓
Reject low-priority work
  ↓
Protect critical traffic
Enter fullscreen mode Exit fullscreen mode

An analytics event might be dropped during extreme overload.

A payment event should not necessarily be treated the same way.

The exact policy depends on the application, but the important thing is to decide what happens before the queue reaches its limit.

What Should We Measure?

If I were operating a queue-backed system, I'd monitor:

Queue depth
Queue capacity utilization

Arrival rate
Processing rate

P50 / P95 / P99 message age

Consumer concurrency
Consumer utilization
Processing latency

Retry rate
Dead-letter rate
Rejection rate

Backlog recovery rate
Enter fullscreen mode Exit fullscreen mode

The important part isn't just collecting these numbers.

It's understanding their relationship.

For example:

Queue depth ↑
Message age ↑
Arrival rate > processing rate
Enter fullscreen mode Exit fullscreen mode

suggests a capacity mismatch.

But:

Queue depth ↑
Processing rate ↓
Database latency ↑
Enter fullscreen mode Exit fullscreen mode

points toward a downstream bottleneck.

And:

Queue depth ↑
Retry rate ↑
Downstream errors ↑
Enter fullscreen mode Exit fullscreen mode

suggests retries may be amplifying the original failure.

So the useful question isn't simply:

"Is the queue growing?"

It's:

"Why is the queue growing?"

The Architecture I Would Build

Putting those ideas together, I'd want the queue surrounded by explicit controls:

                    Incoming Work
                         │
                         ↓
                ┌─────────────────┐
                │ Admission       │
                │ Rate / Priority │
                └────────┬────────┘
                         ↓
                ┌─────────────────┐
                │ Bounded Queue   │
                │ Depth / Age     │
                └────────┬────────┘
                         ↓
                ┌─────────────────┐
                │ Consumer Pool   │
                │ Concurrency     │
                └────────┬────────┘
                         ↓
                ┌─────────────────┐
                │ Downstream      │
                │ Capacity Limit  │
                └────────┬────────┘
                         │
              ┌──────────┴──────────┐
              ↓                     ↓
        Success / ACK          Failure / Retry
                                    │
                                    ↓
                           Backoff / Retry Budget
                                    │
                                    ↓
                              Dead Letter
Enter fullscreen mode Exit fullscreen mode

Each part has a different responsibility:

  • Admission control limits incoming pressure.
  • The queue absorbs temporary rate differences.
  • Concurrency limits protect downstream dependencies.
  • Priority controls protect valuable work.
  • Retry budgets prevent failures from multiplying load.
  • Dead-letter handling isolates work that cannot be processed.
  • Queue age metrics expose the latency created by backlog.
  • Coalescing or expiration can eliminate work that no longer has value.

The queue is still valuable.

We're simply giving it a defined job.

The Tradeoff

There is no universally correct queue size.

A larger queue gives you more room to absorb bursts, but also allows more work to accumulate before the system pushes back.

A smaller queue creates earlier backpressure, but provides less burst tolerance.

So I wouldn't choose queue capacity simply because the underlying technology supports a particular number of messages.

I'd start with the workload.

Suppose:

Consumer capacity:      8,000/s
Burst arrival rate:    10,000/s
Burst duration:             60s
Enter fullscreen mode Exit fullscreen mode

The excess work is:

(10,000 - 8,000) × 60
= 120,000 messages
Enter fullscreen mode Exit fullscreen mode

That gives us an estimate of the buffer required to absorb that particular burst.

Then I'd ask:

How old can the work become?
How quickly can we recover?
What happens when the buffer fills?
Which work can be discarded?
Which work must be protected?
Enter fullscreen mode Exit fullscreen mode

Those answers determine the actual queue design.

The Bigger Lesson

I like queues because they make systems tolerant of temporary differences in processing rates.

The problem begins when a temporary buffer becomes a permanent holding area for work the system cannot keep up with.

A queue can absorb a burst.

It can decouple producers from consumers.

It can protect downstream services from sudden traffic.

But it cannot change the fundamental relationship between incoming work and processing capacity.

And once work starts waiting, another resource appears:

time.

That work may become slower to process, more expensive to recover, or eventually not worth processing at all.

So I think about queues less as storage for work and more as a bounded buffer between systems operating at different rates.

Use them to absorb bursts.

Use backpressure when the buffer fills.

Protect important work.

Control retries.

Remove stale work when the business semantics allow it.

Measure age and recovery, not just depth.

And always know what happens when the queue reaches its limit.

Because ultimately:

A queue doesn't give you more capacity.

It gives you time to respond to a capacity problem.

And sometimes that time is exactly what you need.

Sometimes it is just borrowed time.

Top comments (0)