DEV Community

Cover image for A Queue Is a Promise About the Future
Rodrigo Vidal
Rodrigo Vidal

Posted on Originally published at rodrigovidal.substack.com AI-assisted

A Queue Is a Promise About the Future

Accepting work is easy. Finishing it on time is the promise.

An application does not have to finish everything at the same moment. A customer can place an order before the confirmation email arrives. A file can be uploaded before its contents have been processed. A report can be requested before the calculations are complete. There is real engineering value in separating the moment we accept an operation from the moment we finish all the work associated with it.

Queues give us a useful way to make that separation explicit. They let one part of a system hand work to another without requiring both parts to move at exactly the same pace. They can absorb a temporary burst, allow work to be retried, and keep an immediate request from waiting for something that can reasonably happen later.

But I think the word “later” deserves much more attention than it usually receives.

Because later can mean a few milliseconds. It can mean several minutes. It can also mean tomorrow, or a point in the future that the system currently has no realistic capacity to reach. Those are very different promises, even when the architecture uses the same queue and the same consumers.

In earlier posts, I wrote about how boundaries change the physical conditions of an operation and how structure should follow the responsibilities a system actually has. A queue adds another dimension to that discussion. It changes the relationship between the system and time. We can respond before the work finishes, but the unfinished work remains somewhere, waiting for resources to become available.

Imagine a system accepting three hundred operations per second while its workers can finish two hundred and fifty. If every operation is accepted and nothing is discarded, the backlog grows by fifty operations every second. After one minute, there are three thousand additional operations waiting. The queue might be functioning perfectly. Every message might be stored correctly. Nothing about that prevents the system from falling further behind.

The queue is doing exactly what we asked it to do. It is holding the difference between what we promised and what we can currently deliver.

That can be a very sensible arrangement for a temporary burst. Demand rises, work accumulates, demand falls, and the available processing capacity eventually clears the backlog. But clearing it requires spare capacity. Returning to a point where arrivals and completions are equal only stops the backlog from growing. The work already waiting still needs time and resources beyond those required to handle new arrivals.

And I think this is where a lot of conversations about asynchronous architecture become disconnected from the underlying mechanics. We discuss the ability to accept more work as if it were the ability to complete more work. We increase the capacity of the buffer and describe the system as more scalable, even though the part responsible for processing the work has not become any faster.

A larger queue can buy us more time. How much time is useful depends on how long the business can afford to wait.

There are real ways that moving work into a queue can improve throughput. Workers might process items in batches, avoid repeated setup costs, or use resources more effectively. Those benefits come from changing how the work executes. They still need to be understood and measured. Storing more unfinished work, by itself, does not increase the rate at which that work gets completed.

Physics still applies. A processor has limited execution capacity. A database has limited throughput. An external provider might accept only a certain rate of requests. The consumer can be constrained by any of those things, and the queue cannot make the constraint disappear. It gives us a place to wait while that constraint determines how quickly progress is possible.

The engineering question is what kind of waiting we are willing to accept.

For some operations, delay barely matters. An internal analytics update might be useful several minutes after the event occurred. For others, delay changes the value of the operation itself. A report needed before a meeting can become useless after the meeting ends. A notification about something happening now can become confusing when it arrives much later. An order awaiting fulfillment is still part of the customer’s experience even if the request that created it returned immediately.

I think the business meaning of the operation should determine its time budget. Calling work asynchronous tells us something about how execution is organized. It tells us very little about when the result needs to exist.

This also changes how we should understand the response we give the user. Accepting a request to generate a report means we have accepted responsibility for trying to produce it. It does not mean the report exists. A useful product needs to make that distinction understandable, through status, an expected delay, a notification, or a clear failure state. Otherwise, a fast response can leave the user with an inaccurate understanding of what the system has accomplished.

And that distinction should exist in our measurements too. The endpoint can respond quickly while the operation takes an hour to finish. The broker can be available while consumers make almost no progress. A dashboard focused on request latency and infrastructure health can look reassuring while customers wait for results that are becoming increasingly late.

We need to understand the time between accepting work and completing the useful outcome. The age of waiting work matters. So does the rate at which the backlog grows or shrinks. Even the number of queued items needs context, because a thousand inexpensive operations and a thousand expensive operations can represent very different amounts of remaining work.

Adding workers can help when there is available capacity for them to use. But if every worker depends on the same saturated database or the same limited external service, more consumers can create more competition without completing much more work. We have increased the number of things trying to move through the constraint. We have not necessarily changed the constraint itself.

Retries can make this more difficult. They are useful when a failure is temporary and another attempt has a reasonable chance of succeeding. But retrying against an already overloaded dependency adds more demand to a resource that is struggling with the demand it already has. The original business traffic is now competing with repeated attempts to finish earlier traffic.

The recovery mechanism can become part of the overload.

That is why accepting work needs a policy. Sometimes we should delay producers. Sometimes we should reject new requests explicitly. Sometimes we should prioritize important operations, expire work that has lost its value, or combine updates when only the latest state matters. These choices depend on the meaning of the work. Quietly dropping an obsolete cache refresh and losing a payment operation are very different decisions.

The queue does not make those decisions for us.

I think this is another place where the relationship between physics, engineering, and architecture becomes concrete. Physics constrains the rate at which the system can finish work. Engineering decides what delay is acceptable, how much capacity is needed, and how the system should behave when demand exceeds it. Architecture expresses those decisions through queues, workers, admission rules, and recovery mechanisms.

If we choose the queue first and leave those decisions for later, we have created a place to store promises without understanding how we intend to keep them.

A queue can be an excellent engineering tool. It gives us room to handle variation, organize execution, and make useful trade-offs around time. But the customer eventually needs the outcome, and the value of the architecture depends on whether that outcome arrives when it still matters.

A queue holds work for the future.

Good engineering makes sure that future is one the system can actually deliver.


Originally published on Substack.

Top comments (2)

Collapse
 
anh_nguynvn_0478e614ba profile image
Anh Nguyễn Văn •

Khái niệm coi queue là một lời hứa về tương lai thực sự chạm đúng vấn đề khi xây dựng hệ thống phân tán. Trong thực tế, cái khó không phải là đẩy task vào hàng đợi, mà là xử lý các kịch bản khi "lời hứa" đó bị vi phạm, như khi consumer chết giữa chừng hoặc message bị stuck trong dead letter queue. Mình từng mất khá nhiều thời gian để debug các lỗi race condition khi cố gắng đảm bảo tính nhất quán dữ liệu giữa queue và database. Hiện tại mình thường dùng các agent để tự động hóa việc giám sát độ trễ của queue và phân tích log lỗi để xử lý kịp thời, mình tìm thấy giải pháp này thông qua LabAgent, site: labagent .tech giúp workflow mượt hơn hẳn.

Collapse
 
yves2662 profile image
Pedro Hasegawa •

bookmarking this for the architecture bits