DEV Community

Samandar Xusenov
Samandar Xusenov

Posted on

Why a 200 OK is not proof of delivery: 4 lessons from routing 100k+ webhook deliveries

I maintain a service that takes Facebook and Instagram Lead Ads and delivers each new lead, in real time, to wherever the business actually works — a CRM, Telegram, Google Sheets, a CPA network, or a plain webhook. Over 50 destination types, all of them somebody else's API, all of them free to break at 3am.

The naive version of this is one HTTP POST. The real version is about 80% error handling, and almost every line of that 80% exists because something failed silently in production first.

Here are the four lessons that cost me the most.

1. A 200 OK is not proof of delivery

This one still stings.

We had an integration reporting a ~99% success rate while users complained that rows were missing. The destination returned 200 OK every single time. It was also writing empty rows.

The cause was a field mapping that pointed at a form field's label instead of its value. We sent a well-formed request full of empty strings, and the API cheerfully accepted it.

The fix was to stop trusting the status code and start reading the response body. For Google Sheets, the append call echoes back an updatedRange:

"updates": { "updatedRange": "Sheet1!A2:C2", "updatedRows": 1 }
Enter fullscreen mode Exit fullscreen mode

If that range is missing, or the row count is zero, it was not a delivery — no matter what the status line said. Every adapter now has a "what does success actually look like in the body?" check, not just res.ok.

Rule of thumb: for any write API, find the field in the response that proves the write happened, and assert on that.

2. Retries have to survive an outage, not a blip

Most retry examples you find online do 3 attempts with exponential backoff over ~30 seconds. That handles a dropped packet. It does not handle a CRM that is down for six hours, which is completely normal.

We separate failures into two classes:

  • Validation failures — duplicate, blacklisted, malformed phone number. Retrying is pointless; the payload will never become valid. Fail once, record why, move on.
  • Infrastructure failures — timeouts, 5xx, DNS, connection reset. These are worth retrying for a long time.

Infra retries now stretch across roughly 26 hours with widening backoff. An overnight outage stops being lost revenue and becomes a delay.

The important part is that this split is decided at the point of failure, by inspecting the error, not by a global retry count. Getting that classification wrong in either direction is expensive: retry a validation error forever and you burn the queue; give up on a 503 after 30 seconds and you lose the lead.

3. Circuit breakers belong per destination, not globally

The first breaker I wrote was global. One dead endpoint slowed everyone down, because the queue kept marching into the same wall.

Breakers are now keyed per destination. But that raised a subtler question: when does a breaker close again?

A breaker that opens because a remote host is unreachable is telling you something about the network. A breaker that opens because the API returned 401 is telling you something about one user's credentials. Those need completely different recovery:

  • Network-class breakers can heal on their own, and they can also be healed by evidence — if any other tenant successfully reaches that same host, the network is clearly fine, so parked breakers can be moved from OPEN to HALF_OPEN immediately instead of waiting out a timer.
  • Auth-class failures should never silently retry into a lockout. They should flip the connection's health to "broken" and tell the human, with the specific reason.

Conflating those two was one of our longest-lived bugs. A transient 502 from one provider once flipped a perfectly good connection to error and quietly stopped delivering. Now transient failures leave health unchanged, and only a provider-rejected credential marks a connection broken.

4. Report your success rate honestly

Our dashboard used to show a delivery-rate gauge computed as delivered / total. It looked bad, and unfairly so — because total included duplicates we intentionally blocked, blacklisted numbers, and phone numbers that were never valid.

Those are the system working, not failing.

So validation-class outcomes are now excluded from the denominator. The gauge answers a real question — "of the leads that could be delivered, how many were?" — instead of a meaningless one.

This is worth saying out loud because the temptation runs the other way. It is very easy to build a metric that makes your product look good and tells your user nothing. A number your user cannot act on is decoration.

The boring conclusion

Nothing here is clever. It is all just refusing to believe optimistic signals: the status code, the retry count, the global breaker, the flattering percentage.

If you are building anything that hands data to somebody else's API, I would spend your first week of hardening on lesson 1. It is the one that hides the longest.


I work on Targenix, which does this for Facebook and Instagram lead ads. The list of destinations is here if you are curious what 50+ different failure modes looks like in practice.

Top comments (0)