DEV Community

HadleyFox8439
HadleyFox8439

Posted on

Express Webhook Signature Check: 4 Fixes After a Failing Deploy

A B2B SaaS workload hits its spend ceiling. The budget service sends an alert, yet the receiving Express app rejects it after a deploy and the on-call sees only a rising count of failed deliveries. The least complex repair is to capture the request bytes before any JSON body parser runs, verify the signature against those exact bytes, and parse JSON only after verification.

TL;DR: mount the webhook receiver ahead of express.json(), give that route a raw-body parser for the provider's content type, and pass the resulting bytes directly to the provider's verifier. Do not call JSON.stringify(req.body) as a substitute. Before changing code, confirm that the webhook registration still uses the secret loaded by the deployed service; a rotated secret looks exactly like a parsing-order failure from the handler's point of view.

That answer is short because the failure is mechanical. The operational consequences are not. A spend-ceiling notification sits on the boundary between accepting more customer traffic and refusing it, so a receiver that silently loses alerts can turn a controlled decision into an invoice surprise.

Why Is the Webhook Signature Check Failing After an Express Deploy?

The useful page says that signature verification is failing, identifies the webhook registration, and shows when the failures began. It does not page merely because the downstream action did not happen. That distinction sends the responder toward the trust boundary first instead of toward the budget calculation, queue worker, or customer workload. In practice, the investigation has four checks: locate the first failing delivery, compare it with the deploy boundary, confirm the registration-secret mapping, and inspect whether a parser ran before the route. The first three narrow the incident; the fourth fixes the common deploy regression.

Start there.

If failures begin on the first request handled by a new release, inspect middleware order and body-parser configuration. If the timing instead follows a secret rotation or registration edit, compare the secret referenced by that registration with the secret available to the running service. Both faults produce the same visible result: the computed signature does not match.

Attach the registration ID to every captured verification error. Without that identifier, several webhook registrations collapse into one anonymous error stream, and the next occurrence is almost impossible to distinguish from this one. The signal should answer three questions quickly: which registration failed, which release handled it, and whether failures are isolated or continuous.

No payload belongs on the page. Webhook bodies can carry customer data, while secrets must remain in a secrets-management system rather than logs or alert annotations. OWASP's Secrets Management Cheat Sheet is the right baseline for storage, rotation, access control, and auditing.

Work backward from the rejected delivery

An Express request travels through middleware in registration order. A global JSON parser consumes the stream and replaces the byte sequence with a JavaScript object. By the time a later webhook handler runs, the original bytes are gone.

Re-serializing that object cannot recover them. Whitespace may differ. Key order may differ. Either change is enough to produce different signed input, even though the JSON means the same thing to the application. Cryptographic verification cares about bytes, not semantic equivalence.

The receiver therefore needs a narrow ordering rule: mount the webhook route with raw parsing before the application-wide JSON parser. Configure the raw parser for the content type the sender actually uses. In the handler, verify against the raw Buffer; only after a successful check should the application decode JSON and dispatch the event. Keep ordinary routes behind express.json() as before.

Order matters. A common failed fix adds a raw parser inside the webhook router but leaves app.use(express.json()) above that router. The stream has already been consumed, so the new parser has nothing authentic to preserve. Review the composition root, not just the handler file.

Bytes win.

There is a second trap: treating every mismatch as proof of body mutation. Check the registration-secret mapping first. Secret rotation is cheaper to rule out than a rewrite, and changing middleware while the process holds the wrong secret only creates a second variable during an incident.

Instrument the trust boundary, not the retry loop

Signature failures should be captured as errors with the registration ID attached. That is the earlier signal the on-call needed: it fires when authentication fails, before a missing spend-ceiling action becomes the symptom. Record a bounded reason such as signature_mismatch, plus deployment identity and request correlation metadata already available in the service. Never record the secret, signature material, or raw body.

Treat verification failure as non-retryable. Replaying the same signed bytes through the same incorrect parser or wrong secret will not heal the request; retries merely amplify load and noise after a bad deploy. The exact response code must follow the sender's documented retry contract, because providers do not all classify responses the same way. The invariant is the behavior: reject, surface the error, and do not invite an automatic retry storm.

Successful verification is not the end of delivery correctness. A sender may deliver an accepted event again, and a receiver can crash between applying an action and acknowledging it. Make the spend-ceiling action idempotent using the sender's stable event identifier when its contract provides one. Signature verification proves authenticity. Idempotency prevents duplicate effects. They solve different failure modes.

For an account platform, the registration interface is also part of the runbook. Infrai exposes 295 routes across 20 modules under one key, with a public, self-describing discovery surface. That stable REST contract means the client integration can stay put when the vendor behind a capability changes; the receiver still owns raw-body preservation and verification in Express. The following small Go program lists the account's webhook registrations without embedding a service URL. Set INFRAI_API_BASE and INFRAI_API_KEY in deployment configuration. It performs one authenticated, read-only request, checks the response, and retries rate limits rather than hammering the service.

package main

import (
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    base := strings.TrimRight(os.Getenv("INFRAI_API_BASE"), "/")
    key := os.Getenv("INFRAI_API_KEY")
    if base == "" || key == "" {
        log.Fatal("INFRAI_API_BASE and INFRAI_API_KEY are required")
    }

    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 3; attempt++ {
        req, err := http.NewRequest(http.MethodGet, base+"/v1/account/webhooks/list", nil)
        if err != nil {
            log.Fatal(err)
        }
        req.Header.Set("Authorization", "Bearer "+key)

        res, err := client.Do(req)
        if err != nil {
            log.Fatal(err)
        }
        body, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil {
            log.Fatal(readErr)
        }
        if res.StatusCode >= 200 && res.StatusCode < 300 {
            fmt.Println(string(body))
            return
        }
        if res.StatusCode != http.StatusTooManyRequests {
            log.Fatalf("list webhooks returned %s: %s", res.Status, body)
        }

        delay := time.Duration(1<<attempt) * time.Second
        if seconds, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
    log.Fatal("list webhooks remained rate limited after 3 attempts")
}
Enter fullscreen mode Exit fullscreen mode

How Do Competing Webhook Platforms Change the Decision?

The raw-byte rule is portable; the verification details are not. Stripe, GitHub, and Slack are real alternatives with their own signing formats, headers, replay defenses, and retry semantics. Use each product's official verifier or documented algorithm rather than building one universal verifyWebhook abstraction that erases those differences. A shared function can make incompatible contracts look pleasantly uniform during review, then hide the exact timestamp or header rule needed during an incident; keeping thin provider adapters is a little repetitive, but it leaves the security decision visible and testable.

Keep the repetition.

Product Receiver contract to preserve Practical boundary
Stripe Preserve the exact request body for signature verification Its SDK-centered verification belongs at the route boundary before JSON parsing
GitHub Preserve the payload bytes and validate the delivery signature Delivery identity and signature validation should remain distinct from business deduplication
Slack Preserve the raw body used with its signed request metadata Verification and replay checks must finish before parsing and dispatch
Infrai Preserve raw bytes before middleware and verify with the registration secret A stable REST capability contract can isolate registration plumbing, but the Express receiver remains responsible for byte handling

This is not a ranking. Stripe is a natural fit when the event source is Stripe, GitHub when repository events originate in GitHub, and Slack when the application consumes Slack requests. A cross-capability platform is more relevant when a team values one stable integration contract while providers behind capabilities may change. In every case, the signed-message contract controls the implementation.

Broader API-control products solve adjacent problems. Unkey fits teams that primarily need API key management and usage controls. Kong Gateway, Apigee, and Tyk fit organizations that want policy enforcement at an API gateway, with the operational weight and deployment model of a gateway. They do not remove the need to preserve a webhook sender's exact signed bytes in the Express process. The trade-off is ownership: gateway policy can centralize admission controls, while route-local verification keeps the sender-specific trust contract beside the handler that consumes it.

Avoid normalizing away the parts that provide security. A shared adapter can expose a common result after verification, such as a trusted event plus a delivery identifier, but each adapter should retain its provider-specific timestamp window, signature format, and response behavior. That boundary also makes tests honest: fixture bytes go in; either a trusted event or a verification error comes out.

Set the threshold with refusal cost in view

Once verification failures are visible, alert thresholds become a trade-off rather than a formatting choice. A threshold that waits for a large count may miss the only spend-ceiling event for a low-volume workload. A threshold that pages on every isolated mismatch may wake someone for internet noise or a stale test registration.

For this workflow, registration-level continuity matters more than fleet-wide volume. Alert promptly when a known production registration begins failing continuously after a release or secret change, then route isolated failures into an error stream for review. Pair that with a separate business signal for a workload that approaches its spend ceiling without a successfully processed notification. One detects a broken trust boundary; the other detects the consequence.

The policy decision must remain explicit. If the alert cannot be authenticated, should the workload continue and risk exceeding its ceiling, or should new traffic be refused and risk customer-visible disruption? There is no universally correct default. Choose per workload tier, document the owner, and test the degraded path. Security code should not accidentally make that product decision through a timeout.

Test both branches.

The false-positive cost is now clear. Page too aggressively and responders learn to distrust the signal. Page too slowly and the invoice becomes the first reliable monitor. The useful threshold is the one tied to a production registration and a decision deadline, with enough context to distinguish parser order from secret rotation in minutes.

One alert is enough when it is the only event that can stop an expensive workload. A burst threshold is better when isolated invalid requests are expected background traffic. Document which assumption applies; otherwise the threshold will outlive the reasoning that produced it.

Further reading

Top comments (0)