DEV Community

Cover image for Did the replay work? Finding the RabbitMQ dead letters that come back
Viktor
Viktor

Posted on Originally published at warrenops.io

Did the replay work? Finding the RabbitMQ dead letters that come back

Originally published at warrenops.io.

The fix is deployed, 143 messages go back from the dead-letter queue to their original exchange, the broker confirms every one of them, and the incident is closed. Ten minutes later the dead-letter queue holds 5 messages again. Are they five of the 143, or five new failures? Without a little preparation nobody can tell, because a publish confirm only says the broker took the message, not that the consumer processed it. Here is how to make a replay answer that question, measured on RabbitMQ 4.3.6.

The short version as a video (0:58, captions, no sound needed):

What a confirmed replay tells you

A replay done right reads each message from the dead-letter queue, publishes it to its original route with publisher confirms, and acknowledges the original only after the confirm. That guarantees nothing is lost on the way. It says nothing about what happens next:

  • the consumer processes the message, which is what you hoped;
  • the consumer fails on it again, and it is dead-lettered back, because the fix did not cover this case;
  • the message sits in the work queue because the consumer is not running yet.

The management UI's "Move messages", a shovel, or a script that loops basic.get and basic.publish all stop at the confirm. To see the second case you have to recognise a message when it comes back.

Tag what you replay

Give every replayed message a header that names the replay, and remove the death headers so the message starts fresh:

import uuid, pika

DEATH = ("x-death", "x-first-death-", "x-last-death-", "x-delivery-count", "x-acquired-count")
replay_id = str(uuid.uuid4())

ch.confirm_delivery()
while (got := ch.basic_get("orders.dlq", auto_ack=False)) != (None, None, None):
    method, props, body = got
    death = props.headers["x-death"][0]  # broker dead-lettering; Spring or MassTransit error queues keep the route elsewhere
    headers = {k: v for k, v in (props.headers or {}).items() if not k.startswith(DEATH)}
    headers["x-replay-id"] = replay_id
    props.headers = headers
    ch.basic_publish(death["exchange"], death["routing-keys"][0], body, props, mandatory=True)
    ch.basic_ack(method.delivery_tag)  # confirm_delivery made the publish wait for the broker
print("replay", replay_id)
Enter fullscreen mode Exit fullscreen mode

Keep message_id and the body as they are: they are what lets you match a returned message to the one you sent, also when two replays of the same message overlap.

What survives the second death

We published a message with an application header, rejected it into a dead-letter queue, replayed it with x-replay-id (once with the death headers stripped, once with them kept), and rejected it again. Classic and quorum queues behaved the same:

After the second dead-lettering Result
your own headers (x-app, x-replay-id) kept, unchanged
x-death when the replay stripped it a fresh entry, count 1
x-death when the replay kept it also a fresh entry, count 1: the broker does not continue a count a publisher sent along
x-delivery-count (quorum) starts again at 1

Two things follow. Application headers ride along through dead-lettering, so the replay id is still there when the message comes back. And x-death cannot count replays: whatever the message went through before, after a republish RabbitMQ starts from scratch. Inside a broker-side retry loop (reject into a wait queue whose TTL dead-letters back to the work queue) the count does go up, 1, 2, 3, as before; it is only the republish that resets it. If you want to know how often a message has been replayed, count it in a header of your own, for example x-replay-count, increased on every replay.

A message that came back is therefore one that carries your replay id and a death: a new x-death from the broker, or, for consumers that republish failures themselves (Spring's RepublishMessageRecoverer, which copies the original headers), fresh exception headers. The id alone is not enough, because a replayed message still waiting in the work queue carries it too.

Look for them

Returned messages arrive at the tail of the dead-letter queue. A peek through the management API takes the head, so look early, while the queue is short, or read deep enough:

rabbitmqadmin -f raw_json get queue=orders.dlq ackmode=ack_requeue_true count=500 \
  | jq --arg id "$REPLAY_ID" '[.[] | select(.properties.headers["x-replay-id"] == $id
        and .properties.headers["x-death"] != null)] | length'
Enter fullscreen mode Exit fullscreen mode

ack_requeue_true puts the messages back. They are marked redelivered afterwards, and on a quorum queue of RabbitMQ 3.13 every such peek counts as a delivery against the queue's delivery limit; on 4.x a requeue no longer counts.

When to look: returns usually show up within minutes, because the consumer fails on the message as soon as it gets it. Look once after a few minutes and again after an hour or two. Failures that depend on time (a nightly job, a token that expires, a downstream that is only down at peak) need a second look the next day.

Reading the result

  • None came back. Good, with a caveat: "not back" means not seen in a dead-letter queue. A consumer that acknowledges and drops a message it cannot handle looks the same. If your consumers log a business id, a spot check of five messages in the consumer log settles it.
  • A few came back, all with the same exception. The fix covered most cases but not this one. Group the returned messages by exception before replaying them again; replaying them unchanged produces the same result.
  • Most came back. The fix is not deployed where the consumer runs, or the consumer fails on something the replay itself changed, for example a header your code relied on and the replay removed.
  • They come back again and again. A message replayed three times without success is not waiting for the next deploy. Park it and read it. Your own replay counter tells you when that point is reached.

Checklist

  • Every replay gets an id, written into each message as a header.
  • Strip the death headers on replay; count replays in a header of your own, because x-death restarts after a republish.
  • A message with your replay id and a new death came back. The id alone does not mean that.
  • Look after a few minutes, after an hour or two, and the next day.
  • Group what came back by exception before you replay it again.

Where Warren fits

Warren does this bookkeeping for every replay: each message carries x-warren-replay-id, every read of a queue notices messages that died again, and Warren reads the source queue itself 5 minutes, 30 minutes, 2 hours and a day after the replay. The replay's page then says "5 of 143 replayed messages died again: 5 back in orders.dlq" and marks them; the audit log shows "5 back" next to the count. Free for one broker.

Try Warren in a minute (one compose file, demo broker with real dead letters included) ยท Documentation

Top comments (0)