Last month I broke my own webhook catcher with the exact thing it was built to catch.
The service is simple: it mints a single-use HTTPS URL, records whatever POST lands on it — headers, body, timestamps — and lets you read the events back for 24 hours. I ended up packaging it as a standalone endpoint (https://x402.freeq.one/tools/webhook_catch.html) so agents can smoke-test webhook integrations without standing up their own receiver.
The incident: I was reloading config on the box and for about 90 seconds the endpoint answered 502. I didn't think much of it. The repository host I was testing against thought a lot of it. Its delivery system retried on a backoff schedule — +1 min, +2, +4, +8 — and by the time the config settled I had 41 stored events. Forty of them were the same delivery. The 100-event cap had gone from a design detail to a live eviction problem, with the retry flood crowding out everything a test session actually wanted to inspect.
The lesson I'd skipped: a receiver's job is to answer 2xx immediately, not to answer "correctly." Every non-2xx is read by the sender as a failure and an invitation to resend — on the sender's schedule, not mine. A test endpoint that leaks 5xx under load is a retry magnet.
Fixes, in the order I applied them:
- Acknowledge first, process later. Return 200 as soon as the raw body is safely buffered; validate and parse asynchronously.
- Dedupe on the delivery ID. Every retry carries the same GUID header. Identical IDs now get marked as retries instead of counted as new events.
- Never leak a 5xx from internal errors. If a storage write hiccups, still return 200. Losing one record beats triggering a storm of forty.
- Bonus win: the retry timestamps turned out to be a free trace of the sender's backoff curve. I now keep them as diagnostics.
If you operate any webhook receiver — test hook or production — answer fast and judge late. The sender will happily bury you otherwise.
Top comments (0)