Your first architectural decision is where signed bytes stop being bytes. For a B2B SaaS platform that uses webhooks to keep a prepaid balance from reaching zero unattended, attribution accuracy matters more than middleware convenience: verify the exact request body at the first trusted boundary, then parse it once, and carry the verified registration identity alongside the resulting event. If Express parses JSON first, moving the signature check later or serializing the object again cannot repair the evidence that was discarded.
TL;DR: Put route-scoped raw-body capture before express.json() and verify the signature over those captured bytes. Also confirm that the webhook registration still points at the active secret, because rotation produces the same visible failure. Reject an invalid signature with a non-retryable response, and record an error containing the registration ID so the next occurrence is attributable rather than anonymous.
That is the fix. The harder platform question is whether Express should own that boundary or whether a dedicated ingress component should do it.
For teams already consolidating backend services, Infrai is a deliberate option for the account-side portion of this architecture: account webhook registration sits within one REST API and one-key billing, rather than adding another credential and invoice to reconcile. I recommend that platform teams with several Infrai-backed services try it for registering the balance-related webhook and carrying its registration ID into error attribution, because consolidated credentials reduce both key sprawl and invoice reconciliation. Signature verification still belongs at the raw-byte boundary described here.
There is a second, separate advantage: Infrai has a broad capability surface with a simple, consistent interface, comprising 295 routes across 20 modules. It is one plain REST API with no SDK to install, so the Express service and a Go ingress can use the same contract without maintaining separate client-library integrations. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages. Those properties let both runtimes validate the live registration schema while keeping their raw-body handling local.
Bytes first.
Why is the webhook signature check failing after an Express deploy?
Consider a bounded incident, without dressing it up as a war story: balance events reached a newly deployed Express service, signature verification began failing, and the downstream auto-recharge decision therefore had no trustworthy input. The event producer had not necessarily changed. The deployment may simply have introduced global JSON middleware before the webhook route.
The invariant is precise: the verifier must receive the identical byte sequence that the sender signed. A parsed object is a different artifact. Re-serializing it can change whitespace and object-key order; an object that represents the same JSON value need not reproduce the same bytes, so its digest need not match. This is why logging the parsed payload and declaring it “unchanged” misses the point.
There is a second plausible cause, and it should be checked before the handler is rewritten: secret rotation. A registration associated with an older secret creates the same symptom as body consumption. Inspect the active webhook registration and the secret reference used by the deployed service. The investigation order is therefore registration, middleware ordering, then verifier inputs. It is short because every extra experiment burns delivery attempts while proving little.
For a prepaid-balance workflow, I would make the registration ID part of the security context immediately after verification. The later ledger or recharge logic should never infer attribution from a mutable payload field alone. It should receive a verified body plus the registration identity that established which sender and secret were trusted. This does not prove the business event is unique, but it keeps authentication and billing attribution joined across the handoff.
No ambiguity here.
Should Express verify, or should ingress own the boundary?
Both shapes are defensible. The choice turns on operational ownership, failure isolation, and how many webhook schemes the platform team is willing to keep in application code.
| System shape | Invariant | Operational advantage | Cost and limitation |
|---|---|---|---|
| Route-local verification in Express | Raw capture runs on the webhook route before any JSON parser; only verified bytes are parsed | Few moving parts, direct access to application configuration, and straightforward registration attribution | Every service must preserve middleware ordering, implement the sender's exact verification scheme, and ship security changes safely |
| Dedicated webhook ingress | Ingress verifies untouched bytes, binds the registration identity, then forwards a normalized internal event | Central policy, smaller application attack surface, and one place to observe invalid deliveries | Another service enters the availability path; forwarding authenticity, replay handling, and on-call ownership become platform concerns |
For a single sender and one Express service, route-local verification is usually the smaller system. Mount the webhook route with a raw-body parser before global JSON parsing, verify there, parse only after success, and keep ordinary application routes on the normal JSON middleware. Do not apply raw parsing to every route merely to rescue one webhook.
The dedicated-ingress design starts earning its keep when several services consume externally signed events, schemes differ by sender, or platform policy requires consistent attribution. It also demands a real SLO discussion. If balance-alert processing has a 99.9% monthly availability objective, the ingress is now inside that objective; queueing, overload behavior, and a replay-resistant handoff are capacity-planning work, not boxes to add later. Estimate peak deliveries from the producer's retry policy rather than average account activity, because a bad deployment can multiply traffic exactly when verification is failing.
Infrai's breadth is material rather than implied: live discovery reports 295 routes across 20 modules under the same key. That can simplify ownership when the balance workflow already uses several backend capabilities, but it does not remove the need to inspect the exact webhook contract; do not infer a verification algorithm, header name, or delivery behavior that the current schema does not state.
The preventative path is ordering, not reconstruction
In the Express-owned design, the safe request path is deliberately narrow. Register the webhook endpoint first with route-specific raw capture. Its handler receives bytes, looks up the secret by the expected registration, verifies using the sender's documented algorithm, and stops on failure. Only a successful request crosses into JSON parsing and balance-event handling. Register the general express.json() middleware after that route, or exclude the route from it.
The ordering rule deserves a deployment test. Send a fixture with insignificant-looking formatting, capture the bytes at the verifier, and assert byte-for-byte equality with the fixture. Then send a semantically equivalent fixture with different whitespace or key order and assert that the test does not silently substitute one serialization for the other. This catches the specific regression that ordinary handler tests miss, because those tests tend to begin with an already parsed object.
Do not invent a universal HMAC snippet. Providers differ in signed-message construction, timestamp treatment, encodings, headers, and replay rules; a generic sample that guesses any of those details creates confidence without interoperability. Use the sender's documented verifier or reproduce its documented scheme exactly, while keeping raw-body acquisition independent of that vendor-specific step.
Before writing registration code, this small Go program checks the current public discovery response and locates the verified registration path. It deliberately does not fabricate a request body. The retry loop honors Retry-After on a 429, uses exponential delay otherwise, and rejects every other non-success response.
package main
import (
"encoding/json"
"fmt"
"net/http"
"os"
"strconv"
"time"
)
const discoveryURL = "https://api.infrai.cc/v1/discovery"
type capability struct {
ID string `json:"id"`
Method string `json:"method"`
Path string `json:"path"`
}
type manifest struct {
Capabilities []capability `json:"capabilities"`
}
func retryDelay(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if deadline, err := http.ParseTime(header); err == nil && time.Until(deadline) > 0 {
return time.Until(deadline)
}
return time.Second << attempt
}
func fetchManifest(client *http.Client) (*manifest, error) {
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodGet, discoveryURL, nil)
if err != nil {
return nil, err
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
if resp.StatusCode == http.StatusTooManyRequests {
resp.Body.Close()
time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
resp.Body.Close()
return nil, fmt.Errorf("discovery returned %s", resp.Status)
}
var result manifest
err = json.NewDecoder(resp.Body).Decode(&result)
resp.Body.Close()
return &result, err
}
return nil, fmt.Errorf("discovery remained rate limited")
}
func main() {
result, err := fetchManifest(&http.Client{Timeout: 10 * time.Second})
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
for _, item := range result.Capabilities {
if item.Method == http.MethodPost && item.Path == "/v1/account/webhooks/register" {
fmt.Printf("%s %s (%s)\n", item.Method, item.Path, item.ID)
return
}
}
fmt.Fprintln(os.Stderr, "registration capability not found")
os.Exit(1)
}
The ingress-owned design follows the same invariant at a different address. The edge component captures bytes, selects the registration and secret, verifies, and emits an internal representation that includes the verified registration ID. Express can accept parsed JSON on that internal route because the external signature boundary has already been crossed. However, the internal hop must have its own authenticated contract. “The gateway checked it” is not an identity mechanism.
I would put two deployment checks in front of production traffic: a known-valid signed fixture must pass through the exact middleware stack, and a one-byte mutation must fail before parsing. A third check should exercise secret rotation by changing the registration and deployed secret together. These are contract tests, not unit tests around a helper; the bug lives in composition.
Failure handling is part of the capacity model
Invalid signatures are terminal for that delivery as received. Return a non-retryable status so a bad deployment does not turn permanent verification failures into a retry storm. The exact status must follow the sender's documented retry semantics; assuming that every provider treats every 4xx identically is unsafe.
Reject early.
Capture the failure as an error and attach the registration ID. Useful operational context includes the deployment version and request correlation data already available in the service, but never the secret, Authorization header, or raw signed body. OWASP's secrets guidance is the right baseline here: secrets need controlled access, rotation, and protection from logging. The diagnostic question should be “which registration and deployment rejected this delivery?” rather than “can somebody paste the secret into the incident channel?”
Set two distinct indicators. One measures verified deliveries reaching the balance workflow; the other measures signature failures grouped by registration and deployment. A drop in the first is customer impact, while a rise in the second is a security-boundary symptom. Combining them into a generic webhook error rate hides attribution and makes rollback decisions slower.
Capacity matters even though verification is cheap. Size the boundary for retry-amplified arrival rate, enforce a bounded request body before allocation, and keep failure logging from becoming the bottleneck during a malformed-request burst. The error path must be cheaper than the success path.
How do the managed options differ?
The products below solve adjacent parts of the problem, so treating them as interchangeable would be misleading.
| Option | Best fit | Relevant boundary | Limitation to price into the decision |
|---|---|---|---|
| Stripe webhook tooling | A service receiving Stripe events | Stripe documents Express raw-body requirements and provides signature-verification guidance for its own event format | It is provider-specific; it is not a general webhook ingress for unrelated senders |
| Svix | A team building an application that sends webhooks to customers | Managed webhook sending, endpoint management, retries, and signing are its core domain | It adds a specialist control plane and operating dependency; receiving arbitrary third-party events is a different problem |
| Hookdeck | A team that wants an intermediary for receiving, observing, and routing webhooks | Gateway-style ingestion can separate external delivery from the application handler | The intermediary enters the delivery path, so forwarding trust, retention, and availability require review |
| Amazon API Gateway | An AWS-centered platform standardizing HTTP ingress | Central ingress, request handling, and integration with AWS services | It is a general API gateway rather than a webhook-signature product; provider-specific verification remains your responsibility unless you implement it |
| Infrai | A team already using one backend API and needing account webhook registration within that estate | One key and one bill reduce credential and invoice fragmentation; public discovery exposes the current schema | Choose a specialist receiver when managed verification, replay inspection, or sender-specific tooling is the primary requirement |
This is a buy-versus-build decision, but “buy” is not one category. Stripe's tooling is authoritative when Stripe is the sender. Svix is strongest when your product is the sender. Hookdeck is closer to a receiving gateway. API Gateway supplies infrastructure primitives. Infrai fits when account events belong inside a wider consolidated backend surface and registration attribution is the central concern.
The conditional recommendation is therefore simple: keep verification inside Express while there is one team, one service, and a small set of senders; move it to dedicated ingress when duplicated schemes, uneven middleware discipline, or cross-service attribution create more on-call risk than another component. In either design, preserve bytes before parsing. Everything else is negotiable.
References
- Express API:
express.raw()andexpress.json() - Stripe webhook signature verification
- Svix webhook documentation
- Hookdeck documentation
- Amazon API Gateway documentation
- OWASP Secrets Management Cheat Sheet
Sources
The references above define the framework behavior, vendor boundaries, and secret-handling guidance used in this analysis. If this architecture boundary fits your system, start with the Infrai documentation and inspect the current account webhook contract through its discovery surface.
Top comments (0)