DEV Community

Cover image for One bad payload took down our kafka consumer group for 45 Minutes
Kamen
Kamen

Posted on Originally published at kamenivanov.substack.com

One bad payload took down our kafka consumer group for 45 Minutes

The alert fired in the afternoon: consumer lag climbing on a topic that normally sits close to zero, not a spike - a wall. Messages were arriving, nothing was being processed, and the lag graph looked like a cliff instead of a curve.

The first assumption was a downstream dependency, the usual suspect for this kind of stall, a slow database call or an unresponsive third-party API holding up the consumer thread, but none of that was the reason. The consumer wasn't slow, it was crashing, over and over, on the same message, at the same offset, every time it restarted.

The payload was a JSON event from an upstream service, a schema change that had gone out that morning, one new field added to a nested object. Nothing about the change looked breaking. It wasn't breaking for the producer, and it wasn't breaking for most consumers of that topic. It was breaking for this one, because the deserializer was configured to fail closed on an unrecognized field instead of ignoring it, a setting picked long ago for a reason nobody currently on the team remembered.

Every time the consumer restarted, it picked up from the last committed offset, hit the same malformed-for-us message, threw a deserialization exception, and died before committing anything past it, there was no dead letter queue. There was no path around the message. The consumer group was structurally unable to make progress past that offset, and every restart just replayed the same failure in a loop, which is why 45 minutes passed before someone manually skipped the offset in production, a fix that felt more like surgery than an incident response.

DLQ

What made this expensive wasn't the schema change itself, changes like that happen constantly and usually pass through without anyone noticing. What made it expensive was that the failure mode had no drain. A single unreadable message had the same blast radius as an outage, because the only two options the consumer had were process it or die trying. There was no third option that said set it aside and keep going. The fix wasn't the deserializer setting, though that got revisited too, it was giving the consumer a way to fail a single message without failing the group: a dead letter queue wired into the exception handler, so a poison message gets published somewhere inspectable, the offset commits, and the consumer moves on. The failure becomes a queue with one message in it instead of a stalled pipeline and a page at 2 in the morning.

Kafka Consumer

If your consumers don't have a DLQ, the question worth asking isn't whether a malformed message will show up, it's what happens to the entire group the day it does. A queue with no drain for its own failures isn't resilient, it's just quiet until the first message it can't handle.


More like this?
I write short production war stories like this one, plus a deep-dive architecture series. Subscribe here if you want both in your inbox.


Recommended Resources for Java & Spring Boot Engineers

If you are preparing for Senior/Lead Java interviews or looking to solidify your Spring Boot & Architecture skills, check out these highly-rated resources from the Javarevisited publication (Use promo code friends20 for an exclusive 20% discount automatically applied at checkout):

Top comments (0)