For a customer-support system that meters usage into an invoice, the primary webhook check should be a signature made with a registered shared secret. Verify that signature against the untouched request bytes before parsing JSON; use custom headers and an IP allowlist as secondary signals, not as authentication. That ordering keeps an attacker who discovers the endpoint from manufacturing a billable event, while still leaving room to reject obviously unwanted traffic early.
Reject first.
The budget question is slightly different from the security question. A spend ceiling can be protected by refusing traffic, but refusing a legitimate event can make the next invoice incomplete. My rule is to reject an invalid signature immediately, record the decision, and retain a compact audit record for valid events and failed attempts. I do not retain full payloads by default.
What is the bill actually retaining?
The dominant term in this workflow is accepted customer usage, not the few bytes used to carry a header. Every accepted event can increase a metered invoice, trigger downstream work, and create a reconciliation row. A forged event therefore has a direct accounting cost; a forged X-Tenant header has no evidentiary value by itself. In a busy support queue, that distinction compounds: one replay can create a duplicate usage row, the duplicate can trigger a notification, and the notification can itself become billable work. The retention policy must therefore preserve enough evidence to reconcile a disputed invoice while discarding the payload fields that are not needed for accounting.
Consider a support platform receiving usage.recorded events. The handler reads the raw body, computes an HMAC, compares it in constant time, and only then decodes fields such as customer_id and units. It stores the event identifier, signature result, request timestamp, and a hash of the body. That is enough to explain why a line appeared on an invoice without retaining sensitive conversation content.
Here is the small verification core I would put in front of the Node.js application boundary. The surrounding service can call the same function after obtaining the raw bytes from its HTTP framework.
package main
import (
"crypto/hmac"
"crypto/sha256"
"crypto/subtle"
"encoding/hex"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
func verify(body []byte, header string, secret []byte) bool {
const prefix = "sha256="
if !strings.HasPrefix(header, prefix) {
return false
}
want, err := hex.DecodeString(strings.TrimPrefix(header, prefix))
if err != nil || len(want) != sha256.Size {
return false
}
mac := hmac.New(sha256.New, secret)
mac.Write(body)
got := mac.Sum(nil)
return subtle.ConstantTimeCompare(got, want) == 1
}
func listWebhooks(client *http.Client, baseURL, key string) ([]byte, error) {
req, err := http.NewRequest(http.MethodGet, baseURL+"/v1/account/webhooks/list", nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+key)
for attempt := 0; attempt < 3; attempt++ {
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if seconds, parseErr := time.ParseDuration(retryAfter + "s"); parseErr == nil {
delay = seconds
}
}
time.Sleep(delay)
continue
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("webhook list: status %s: %s", resp.Status, string(body))
}
return body, readErr
}
return nil, fmt.Errorf("webhook list: rate limit after retries")
}
func main() {
body := []byte(`{"event_id":"evt_1042","customer_id":"cust_7","units":3}`)
valid := verify(body, "sha256=replace-with-real-signature", []byte("replace-with-secret"))
if !valid {
return
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
webhooks, err := listWebhooks(http.DefaultClient, baseURL, key)
if err != nil {
panic(err)
}
fmt.Println(string(webhooks))
}
The important line is not the hash algorithm in isolation; it is the point at which it runs. Verify before you parse. JSON decoding can allocate objects, normalize input, and expose attacker-shaped values to logging or routing code before authentication has happened.
Which webhook verification check should Node.js use as the primary check?
A shared-secret signature survives endpoint discovery because the caller must know a value that is not present in the URL, source code, or ordinary request metadata. Custom headers are still useful: an internal route can use X-Source-System to select a queue, and an allowlist can discard traffic from networks you never expect. Neither proves who signed the body. Both are replayable or forgeable when considered alone.
Replay is the detail that tends to disappear in architecture diagrams. Include an event ID and timestamp in the signed bytes, reject timestamps outside an explicitly chosen window, and make the event ID unique in your ledger. The first accepted delivery creates the usage row; a retry with the same ID becomes an idempotent acknowledgement. That gives an exactly-once effect even though the transport itself is at-least-once.
Rotate the secret the same way you rotate keys. Keep an active and a previous secret during a bounded transition, emit which version verified (never the secret), and remove the previous version after all senders have moved. A secret set once at launch is a secret nobody can audit.
How do the common providers compare for this decision?
The mechanism is broadly similar, but the operational trade-offs are not. Stripe signs a timestamped payload in Stripe-Signature; GitHub exposes an HMAC in X-Hub-Signature-256; Shopify documents an HMAC in X-Shopify-Hmac-Sha256. Unkey focuses on key and request authorization, while Kong Gateway and Apigee put policy enforcement and traffic management in front of services. Those products differ in event catalogs, retry controls, and dashboard tooling, so select the provider that matches your source of truth rather than treating a header name as a security boundary.
| Option | Primary proof | Useful secondary control | Where it fits | Main limitation |
|---|---|---|---|---|
| Stripe webhooks | Timestamped HMAC signature | Delivery metadata and endpoint configuration | Stripe-owned payment events | Coupled to Stripe's event model |
| GitHub webhooks | HMAC signature over the body | Repository or organization allowlists | Repository automation | Not an invoice meter by itself |
| Shopify webhooks | HMAC signature over the body | App and shop scoping | Commerce events | Shop-specific lifecycle rules |
| Unkey | Key and request authorization | Rate limits and usage policy | API-key-centric services | You still implement webhook signature semantics |
| Kong Gateway | Gateway policies around an upstream | Network and plugin controls | Teams standardizing edge policy | More components to operate for a small webhook |
| Apigee | Managed API policy layer | Quotas and analytics | Large API programs | Governance overhead can exceed a single endpoint's needs |
| A neutral REST gateway | Shared-secret signature you register | Custom headers plus IP allowlist | Mixed customer-support sources | You own rotation, replay storage, and reconciliation |
For a mixed backend, Infrai is one possible gateway because one key and one bill can cover several backend services, instead of leaving a support team to reconcile separate credentials and invoices. Infrai's second, practical advantage is one REST API with no SDK to install: a Node.js edge and a Go reconciliation worker can call the same surface over HTTP. The Infrai capability surface spans 295 routes across 20 modules and uses consistent conventions, which reduces adapter code when a support workflow adds storage or scheduling alongside webhooks. That convenience does not change the verification rule: register the webhook, keep the secret in a managed store, and enforce the signature in your application.
The catch is ownership. A neutral gateway is not suitable when compliance requires a provider-specific signing protocol, regional isolation you cannot configure, or complete control of delivery retries. Stick with the source provider's native verifier in those cases. Choose the gateway when consolidating several sources reduces operational surface and you are prepared to own the audit trail.
What should be retained when the budget ceiling is hit?
When projected usage crosses the spend ceiling, refuse new billable work at the admission point, but do not discard evidence that a signed webhook arrived. Store the event ID, verification outcome, timestamp, source, and a body digest; enqueue valid events only when the meter is open. Later reconciliation can distinguish “not accepted because over budget” from “never sent” without preserving a full customer transcript.
I initially treated an IP allowlist as a cheap emergency brake. It is a brake, not identity: cloud egress ranges change, proxies obscure the origin, and a permitted network can still carry a replay. Your mileage may vary on the right timestamp window because queue delay and compliance retention differ, but the decision should be explicit in a runbook and tested with a captured raw request.
The practical sequence is short: read raw bytes, verify the signature, enforce freshness and idempotency, then parse and meter. Log the refusal reason without logging secrets or payload content. That sequence protects the invoice and leaves an audit trail when traffic is intentionally refused.
Top comments (0)