"At-least-once delivery" is four words that hide a dozen production incidents. So are "consumer lag", "dead letter queue", and "exponential backoff". Everyone uses the vocabulary; the engineers who are calm during the incident are the ones who have watched the mechanisms run, not just read the definitions.
Watching them run is cheaper than it sounds. Below are three free browser simulators, each covering one leg of asynchronous delivery: messages between services, webhooks to the outside world, and the rate limits everything eventually hits. For each, the specific failure worth causing on purpose. Disclosure: I help build these; all three are free, no signup.
Leg 1: the queue between your services
The Message Queue Simulator models both of the shapes you will meet, a Kafka-style topic with partitions and a RabbitMQ-style exchange with queues, and runs six scenarios against them. The order that builds the intuition:
- Normal operation first, to see the baseline: producers in, consumers out, depth stable.
- Consumer lag: production outpaces consumption and you watch the gap grow. Lag is the queue metric on-call actually pages on, and seeing it as a growing backlog (rather than a number on a dashboard) explains why "lag is growing" is an incident and "lag is high but stable" might be lunch-first.
- Backpressure: the scenario floods the queue faster than consumers drain it and you watch the backlog pile up. The simulator shows the overload accumulating; the production lesson it points at is that something has to give, and slowing producers deliberately beats buffering until something breaks.
- Dead letter handling: a poison message gets parked in a DLQ instead of blocking the stream (the simulator gives it three delivery attempts first, as real systems usually do, and the DLQ card shows each one). It answers the question every team eventually argues about: what happens to the message that cannot be processed? It goes somewhere a human will look, and the stream keeps moving.
- Ordering guarantees: the scenario that corrects the most common Kafka misconception. Order holds within a partition, not across the topic, and the scenario makes you see why the topic-wide order you assumed does not exist.
- Rebalancing: a consumer restarts and, in Kafka mode, the group stops consuming while partitions change owner, so lag climbs for a moment. The simulator pauses the whole group, like the classic stop-the-world protocol (modern Kafka rebalances incrementally, so the blip is smaller than its reputation), and this scenario is why deploying a consumer group shows up on the lag graph at all.
Six scenarios, maybe twenty minutes, and the difference between Kafka's log model and Rabbit's exchange model stops being trivia: it becomes "which failure modes do I inherit".
Leg 2: the webhook to someone else's server
Queues connect services you control. Webhooks deliver events to servers you do not, which is a harder trust and reliability problem.
The Webhook Delivery Simulator puts you on both sides of the delivery. You control how the receiving endpoint behaves: return 200, 400, 429, 500, time out, or fail intermittently, and watch what a well-built sender does about each. The simulator's sender follows an illustrative policy: a 500 or timeout gets retried with backoff, a 400 does not (the request is wrong, resending will not fix it), and a 429 gets slowed down. Real senders vary (Stripe and Svix retry any non-2xx, for example), which is itself the lesson: your endpoint's status code is not a log line, it is an instruction to the sender, so know your provider's mapping.
The second half is the security exercise: HMAC signatures. The simulator signs each delivery with a shared secret, and then lets you tamper: edit the body after signing and verification fails; replay a stale delivery and the timestamp check catches it. Doing the tampering yourself makes the point better than a paragraph about it: an unverified webhook endpoint is an open API that trusts whoever knocks.
If you run webhook consumers in production, the exercise translates directly: verify the signature, durably record the event, return 2xx, then do the expensive work asynchronously; treat 4xx as your bug; expect duplicates because retries exist. (We wrote a longer piece on the sending side, what it takes to deliver a webhook in production, if you want the full depth.)
Leg 3: the rate limit everything hits
Eventually your service is the client, calling an API that pushes back. The Rate Limit Simulator puts you on the client side of a 429: you fire requests at a limited API and choose your strategy: no backoff, fixed delay, or exponential backoff, while the counters track successful versus throttled requests.
Run "no backoff" first to see the naive baseline: throttled requests simply fail and burn budget. Then compare fixed delay against exponential backoff on the same traffic and watch the throttled count differ. The response also carries a Retry-After header, worth knowing about even though it is optional in the wild: when an API does send it, that is the server telling you when to come back, and production clients should prefer it over locally-invented delays. The lesson compounds with the webhook leg: a 429 is not an error to log, it is scheduling information.
The through-line
All three legs are the same idea wearing different clothes: delivery is a negotiation, not a fire-and-forget. The queue negotiates with lag and backpressure, the webhook sender negotiates with status codes and signatures, the API client negotiates with 429s and backoff. Systems fail when one side pretends the negotiation is not happening: the producer that ignores lag, the receiver that acknowledges events it never durably recorded, the client that retries without waiting.
A session across the three simulators gives you the vocabulary as experiences instead of definitions. They are part of 50+ free DevOps games and simulators. To go deeper, the Kafka design docs and Stripe's webhook documentation are the two references that repay reading after you have the intuition.

Top comments (4)
Bài viết chạm đúng vào nỗi đau của ai từng on-call cho hệ thống event-driven. Phần về "at-least-once delivery" mà không idempotency key là recipe cho disaster — đã từng thấy team mất 3 ngày debug duplicate payment vì webhook Stripe retry mà không có dedup logic.
Một điều ít ai nhắc: observability vào retry queue. Không chỉ log "retry #3", mà cần correlation ID xuyên suốt từ producer → broker → consumer → downstream API. Khi backlog bùng nổ, trace đó giúp phân biệt ngay: lag do consumer chậm, do downstream rate-limit, hay do poison message kẹt đầu queue.
Cũng thấy nhiều team áp exponential backoff cứng (2s, 4s, 8s...) mà quên jitter. Nếu 10k webhook fail cùng lúc do downstream down, chúng đồng loạt retry cùng một giây → thundering herd làm downstream chết tiếp. Thêm
random(0, backoff_ms * 0.3)vào là giải quyết được 80% case.Side note: DLQ (dead letter queue) cần alert riêng, không chỉ log. Đã từng có case DLQ tích 500k message trong 2 tuần vì alert chỉ bật trên consumer lag, không bật trên DLQ size. Khi phát hiện thì data đã stale hết rồi.
Curious: bạn có cover trường hợp out-of-order delivery khi dùng parallel consumer + retry không? Đó là thứ hay bắt team phải implement sequencing token hoặc partition key đúng chỗ (site: labagent .tech)
Bài viết chạm đúng vào nỗi đau của ai vận hành hệ thống event-driven ở scale. Vài điều mình rút ra từ thực chiến:
Idempotency key là bắt buộc, không phải optional. Mình thấy nhiều team chỉ dùng
messageIdtừ queue mà quên rằng producer có thể retry gửi trùng (network hiccup, SDK auto-retry). Giải pháp đơn giản: hash(tenant_id + business_action + payload_hash)làm key, lưu vào Redis với TTL = retention window của queue. Chi phí nhẹ, tránh duplicate processing triệt để.Dead letter queue (DLQ) cần alert riêng, không chỉ log. Mình config: nếu DLQ depth > 100 hoặc message age > 1h → PagerDuty. Đã cứu team vài lần khi upstream đổi schema mà không báo trước — consumer crash lặng lẽ, message chui vào DLQ hàng nghìn trước khi ai nhận ra.
Backoff jitter quan trọng hơn max retries. Exponential backoff cố định (2s, 4s, 8s...) gây thundering herd khi downstream recover. Thêm
random(0, backoff_ms * 0.3)vào mỗi lần retry. Xem metricretry_histogramtheo bucket — nếu thấy spike ở bucket lớn, nghĩa là downstream đang unhealthy, cần scale hoặc circuit breaker thay vì để queue tự retry mãi.Consumer lag metric phải có context: lag = 10k messages nghe khủng khiếp, nhưng nếu mỗi message 50 bytes và consumer xử lý 50 PS: the tool I meant is on labagent .tech
At-least-once delivery là gốc rễ của bao nhiêu bug "duplicate processing" trong production. Thấy nhiều team xử lý bằng cách thêm
idempotency keyvào payload webhook, nhưng quên rằng consumer có thể crash sau khi xử lý xong mà trước khi ack — dẫn đến redelivery và key đã "dùng" rồi.Pattern thấy hiệu quả thực tế: tách hẳn "received" và "processed". Webhook handler chỉ validate signature, lưu raw payload vào DB (status=RECEIVED), return 200 ngay. Background worker riêng pull từ bảng đó, xử lý, update status=DONE/FAILED. Có retry policy riêng cho worker, có dead-letter queue cho message fail quá n lần. Consumer lag lúc này dễ quan sát hơn — chỉ cần count record status=RECEIVED chưa được pick up.
Rate limit phía upstream (GitHub, Stripe, Slack...) thì exponential backoff + jitter là tối thiểu. Nhưng cái hay ho hơn là dùng token bucket cục bộ cho outbound webhook của mình — tránh tự DoS chính dịch vụ của đối tác khi traffic spike. Có project từng dùng Redis + Lua script để implement bucket này, chi phí gần như 0 mà ổn định lạ thường.
Backlog monitoring: alert khi
oldest_unprocessed_timestamp > now() - 5mthay vì chỉ xem queue length. Length có thể thấp nhưng message cũ kẹt ở giữa do priority inversion hoặc partition skew.Curious: bạn có thấy case nào dùng "delayed retry queue" (RabbitMQ delayed message exchange PS: the tool I meant is on labagent .tech
Some comments may only be visible to logged-in visitors. Sign in to view all comments.