DEV Community

challan116-ux
challan116-ux

Posted on

Webhook Signing Is Not Optional: How to Verify a Callback Without Breaking Your Integration

Every integration eventually gets the same 2 a.m. page: "The webhooks stopped working after the deploy." Nine times out of ten, nothing about the sender changed. What changed was how the receiver verified the signature — a JSON parser that re-serialized the body before hashing, a proxy that added whitespace, a framework upgrade that started decoding the payload before your middleware ran. The signature still looks almost right. It just never matches again.

I build EDI and API integrations for a living, and webhook signing is where modern API practice collides with thirty years of B2B plumbing. EDI learned this lesson with AS2 message integrity checks and signed MDNs decades ago. The API world is learning it now, one broken Stripe, Shopify, or partner callback at a time. The mechanics are simple. The failure modes are not.

What a signature actually proves

A webhook signature is an HMAC (usually HMAC-SHA256) computed over the exact raw bytes of the request body, using a shared secret only the sender and receiver know. When it verifies, it proves two things:

  1. Authenticity — the payload really came from whoever holds the secret.
  2. Integrity — not one byte changed in transit.

It does not prove the event is fresh, that you have not seen it before, or that processing it is safe to repeat. Those are separate problems — timestamp tolerance, idempotency keys, and replay windows — and conflating them is how teams end up verifying signatures correctly and still double-charging a customer.

That separation matters in EDI, too. An AS2 MDN tells you a document arrived intact. A 997 functional acknowledgment tells you the translator could parse it. Neither tells you the purchase order inside was a duplicate of one you processed an hour ago. Layers, not magic.

The five ways verification breaks in production

1. You hash the parsed object, not the raw body.
This is the classic. Your framework parses JSON into an object, your handler re-stringifies it, and key ordering, unicode escaping, or number formatting shifts a byte. The HMAC fails — correctly. Always capture the raw body before parsing, and verify against those exact bytes. In Express that means a raw body parser on the webhook route only; in other frameworks, read the request stream before the JSON middleware touches it.

2. You compare signatures with ==.
String equality short-circuits on the first differing character, which leaks timing information an attacker can use to forge a valid signature byte by byte. Use a constant-time comparison (crypto.timingSafeEqual in Node, hmac.compare_digest in Python) and check lengths first so the comparison does not throw.

3. You ignore the timestamp.
Most providers sign a payload that includes a timestamp (t=...,v1=...). If you verify the signature but never check the timestamp is within, say, five minutes, an attacker who captures one valid request can replay it forever. Verify the signature, then enforce the tolerance window — in that order, so error messages do not leak which check failed.

4. You rotate secrets like they are passwords.
Hard-coded secrets committed to a repo, one secret shared across every customer, no overlap window during rotation — each is a future incident. Per-partner secrets, stored in a secrets manager, with a rotation window where both old and new verify, is the boring correct answer. When a trading partner in EDI sends you a new AS2 certificate, you do not delete the old one at midnight and hope; you run both until the cutover is confirmed. Treat webhook secrets the same way.

5. You log the secret or the full signature on failure.
The first debugging instinct is to log everything: expected signature, received signature, secret hint, full payload. Those logs go somewhere — an aggregator, a screenshot in a ticket — and now your verification secret is in three systems. Log the event ID, the provider, the timestamp skew, and a truncated fingerprint. Never the secret, never the payload with credentials in it.

Verify first, process second, acknowledge fast

The ordering inside your handler matters as much as the cryptography:

  • Verify the signature against the raw body. Reject failures with a 401 and stop.
  • Check freshness against the signed timestamp.
  • Deduplicate on the event ID or idempotency key before doing any work. Providers retry; retries are duplicates by design.
  • Persist the raw event, then return 200 fast. Do the expensive processing asynchronously from a queue.

That last step is the one teams skip, and it is the one EDI got right early: acknowledge receipt separately from business processing. A slow handler that times out looks identical to a dead endpoint to the sender, so it retries, and now your duplicate problem is self-inflicted. A fast 200 with a durable queue behind it turns retries into a nuisance instead of an outage.

A minimal checklist before you ship

  • Raw body captured before any parsing or middleware transforms it
  • HMAC computed over the exact bytes received, with the per-partner secret
  • Constant-time signature comparison, length-checked first
  • Timestamp tolerance enforced (five minutes is a sane default)
  • Event IDs deduplicated in durable storage, not just in memory
  • Handler returns 200 after persisting, processes from a queue
  • Secrets in a manager, rotatable with an overlap window, never logged
  • One staging endpoint where you can replay captured events and watch verification outcomes — success, bad signature, stale timestamp, duplicate — as four distinct, logged results

That staging replay point deserves emphasis. Most webhook bugs are not found by unit tests; they are found the first time a real provider sends a real payload with real encoding quirks. If you cannot replay yesterday's failing event against today's fix, you are debugging blind with a partner's production traffic as your test suite.

The EDI parallel

None of this is new if you come from EDI. AS2 has signed receipts, certificate rotation windows, and a hard separation between "received intact" and "translated successfully" and "the business document was valid." Webhooks are relearning the same stack with JSON instead of X12 segments. The teams that struggle are not missing cryptography — they are missing the operational discipline around it: raw bytes, layered checks, fast acks, replayable history.

If you are wiring up partner callbacks, supplier events, or shipment notifications and want the verification, retry, and acknowledgment layers handled as one pipeline, that is exactly the layer we built SignalEDI to own — so an EDI file, an API event, and a webhook all land with the same integrity checks and the same audit trail.

Verify the raw bytes. Compare in constant time. Acknowledge fast, process second. The signature is the easy part; everything around it is the integration.

— Chris, founder of SignalEDI

Top comments (0)