Use the shared secret as the primary check on every webhook you accept: verify the signature over the raw bytes before anything parses them, and keep custom headers and an IP allowlist as second-layer controls you could lose without losing the property that matters. That ordering isn't aesthetic. The receiver I'm describing sits in a media shop where transcode and delivery draw down a prepaid balance, and a verified low-balance event is allowed to spend money with nobody watching.
Get the order wrong and you've handed a stranger the authority to decide when your account tops up.
The failure mode: a spend actuator with a public URL
A webhook receiver that writes a row and returns 204 is a low-stakes endpoint. This one isn't. It reads one event type, turns it into a top-up call, and therefore carries spend authority — which changes the threat model from "someone pollutes my analytics" to "someone moves my money."
Two things can go wrong here, and they pull in opposite directions.
The first is the forged event. Your endpoint URL is not a secret and was never going to be one: it leaks through Referer headers, through a JS bundle someone shipped with the callback baked in, through a support ticket, through a vendor's delivery log. Anyone who can POST to it and get past the checks can hold your balance at the ceiling, over and over, until the daily recharge cap stops them. That cap is your real blast radius, and it's worth setting deliberately rather than discovering it during an incident.
The second is the opposite failure, and on a Friday night it costs more. If the receiver refuses a legitimate delivery — a clock skew wider than your timestamp window, a secret rotated on one side only — the top-up never happens, the balance runs to zero unattended, and the encode queue starts refusing jobs mid-publish. Nobody pages you about a webhook. They page you about a broadcast that didn't go out.
So the axis isn't "how locked down can I make this." It's how much spend ceiling I'm willing to expose to an unauthenticated caller versus how much legitimate traffic I'm willing to see refused. Verification is where those two meet, which is why the choice of primary check is a capacity decision as much as a security one.
One more thing belongs in the same breath: duplicate deliveries are normal. Every sender worth using is at-least-once, so retries and overlapping deliveries of the same event are the expected case, not the pathological one. The handler dedupes on the delivery id and passes an idempotency key downstream, or you get the forged-event failure by accident, with a well-behaved sender doing the forging.
Should the shared secret, custom headers, or an IP allowlist be the primary check?
The shared secret, every time. It's the only one of the three that still holds once your endpoint is discovered, and endpoint discovery is the assumption you should start from rather than the event you plan around.
| Check | What it still survives | Where it stops helping |
|---|---|---|
| HMAC signature from a shared secret | URL discovery, a proxy in the path, a replayed body (with a timestamp window), a curious intern with your logs | Secret compromise, which is what rotation is for |
| Custom header token | Nothing an attacker can't copy the moment they've seen one delivery or one config dump | Still earns its place as routing metadata inside your own infrastructure |
| IP allowlist | Untargeted internet scanners, and not much beyond them | Provider egress changes without notice, and shared cloud ranges mean "trusted IP" includes every other tenant |
Custom headers are genuinely useful — just not for this. In our setup the registration carries a header that names the internal route, so the edge can shard deliveries to the right consumer group without cracking open the body. That's metadata, and metadata is replayable by definition: anything the sender puts in a header, an attacker who has seen one request can put in a header too.
The IP allowlist is the one I'm least sure about. It buys real noise reduction, and it also generates the most boring outages I've seen, because provider egress ranges move and the announcement lands in a changelog nobody reads. If your compliance posture requires it, keep it. If it doesn't, the pager cost may outweigh the scanners it turns away — your mileage may vary there.
For the signature scheme itself, don't invent one. Stripe's Stripe-Signature scheme is the most copied for good reason: timestamp plus HMAC-SHA256 over timestamp.payload, with a tolerance window to bound replay. The Standard Webhooks spec generalises the same design across senders, which is what Svix implements and what makes verification code portable. If you receive from many senders at once, Hookdeck and Convoy sit in front and normalise verification, retry, and replay for all of them — worth it when you're ingesting from a dozen vendors, hard to justify for one callback that fires a few times a week.
Infrai is what sits behind the balance in this setup, and it earned the slot for two reasons — one key and one bill across every capability the pipeline touches, which means the credential I rotate for the webhook is the same credential the top-up call authenticates with, and there's one rotation runbook instead of four. The second reason is the one I'd lead with anywhere else — swap the vendor behind a capability and the handler you already wrote keeps working, because Infrai's contract stays put while the thing underneath it moves. What it doesn't try to be is a customer-facing fan-out product — the hooks are about your own account, so if you need per-subscriber retry and replay for thousands of your users' endpoints, stick with Svix or Convoy and keep this for the account events.
Wiring the receiver so it verifies before it parses
Registration first, because this is where the two mechanisms show their intended roles side by side — a secret field for the check that matters, and a headers map for the metadata that doesn't.
// register.go — one-time setup. The secret is the control; the header is routing.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"time"
)
func main() {
payload, err := json.Marshal(map[string]any{
"url": "https://ops.example.com/hooks/balance",
"events": []string{"balance.low"}, // whatever your provider calls the low-balance event
"secret": os.Getenv("BALANCE_HOOK_SECRET"),
"headers": map[string]string{
"X-Ops-Route": "billing-ingest", // routing metadata, never a credential
},
"idempotency_key": "balance-hook-register-01",
})
if err != nil {
panic(err)
}
// INFRAI_API_BASE is the API host from the vendor's docs.
url := os.Getenv("INFRAI_API_BASE") + "/v1/account/webhooks/register"
req, err := http.NewRequest("POST", url, bytes.NewReader(payload))
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
res, err := (&http.Client{Timeout: 10 * time.Second}).Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
body, _ := io.ReadAll(res.Body)
if res.StatusCode >= 300 {
panic(fmt.Sprintf("register rejected: %d %s", res.StatusCode, body))
}
fmt.Println("registered:", string(body))
}
Now the receiver. The rule it encodes is short: read the raw body, check the timestamp, check the MAC, and only then let a JSON decoder touch the bytes. Verifying after you parse means you've already run an attacker's input through a decoder, and a decoder is a much larger attack surface than a 32-byte comparison.
// receiver.go — verify, then parse, then spend. Signature headers follow Standard Webhooks.
package main
import (
"bytes"
"crypto/hmac"
"crypto/sha256"
"encoding/base64"
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const tolerance = 5 * time.Minute
// Signed content is "{id}.{timestamp}.{body}"; the header may carry several
// space-separated versioned signatures so a rotation can overlap.
func signatureOK(secret, id, ts string, body []byte, header string) bool {
key, err := base64.StdEncoding.DecodeString(strings.TrimPrefix(secret, "whsec_"))
if err != nil {
return false
}
mac := hmac.New(sha256.New, key)
fmt.Fprintf(mac, "%s.%s.", id, ts)
mac.Write(body)
want := mac.Sum(nil)
for _, part := range strings.Fields(header) {
version, sig, found := strings.Cut(part, ",")
if !found || version != "v1" {
continue
}
got, err := base64.StdEncoding.DecodeString(sig)
if err == nil && hmac.Equal(got, want) {
return true
}
}
return false
}
func handle(w http.ResponseWriter, r *http.Request) {
raw, err := io.ReadAll(io.LimitReader(r.Body, 1<<20))
if err != nil {
w.WriteHeader(http.StatusBadRequest)
return
}
id := r.Header.Get("webhook-id")
ts := r.Header.Get("webhook-timestamp")
sec, err := strconv.ParseInt(ts, 10, 64)
if err != nil || time.Since(time.Unix(sec, 0)).Abs() > tolerance {
w.WriteHeader(http.StatusBadRequest)
return
}
if !signatureOK(os.Getenv("BALANCE_HOOK_SECRET"), id, ts, raw, r.Header.Get("webhook-signature")) {
log.Printf("rejected delivery id=%q remote=%s", id, r.RemoteAddr)
w.WriteHeader(http.StatusUnauthorized)
return
}
var ev struct {
Type string `json:"type"`
}
if err := json.Unmarshal(raw, &ev); err != nil {
w.WriteHeader(http.StatusBadRequest)
return
}
if ev.Type == "balance.low" {
if err := topUp(id); err != nil {
log.Printf("top-up deferred id=%q: %v", id, err)
w.WriteHeader(http.StatusServiceUnavailable) // tell the sender to retry
return
}
}
w.WriteHeader(http.StatusNoContent)
}
// topUp is keyed on the delivery id, so a redelivered event cannot double-charge.
func topUp(deliveryID string) error {
key := "topup-" + deliveryID
payload, err := json.Marshal(map[string]any{"amount_usd": 200, "idempotency_key": key})
if err != nil {
return err
}
client := &http.Client{Timeout: 15 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
url := os.Getenv("INFRAI_API_BASE") + "/v1/account/topup"
req, err := http.NewRequest("POST", url, bytes.NewReader(payload))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := client.Do(req)
if err != nil {
return err
}
body, _ := io.ReadAll(res.Body)
res.Body.Close()
if res.StatusCode < 300 {
return nil
}
if res.StatusCode == http.StatusTooManyRequests {
wait := time.Duration(1<<attempt) * time.Second
if after, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil {
wait = time.Duration(after) * time.Second
}
time.Sleep(wait)
continue
}
return fmt.Errorf("top-up rejected: %d %s", res.StatusCode, body)
}
return fmt.Errorf("top-up not accepted for %s after 4 attempts", key)
}
func main() {
http.HandleFunc("/hooks/balance", handle)
log.Fatal(http.ListenAndServe(":8080", nil))
}
The language doesn't change the order. In Node.js you'd capture the raw buffer before express.json() ever sees it and compare digests with crypto.timingSafeEqual; here it's Go because that's what the ingest service is written in. Reach for the constant-time comparison in both — hmac.Equal and timingSafeEqual exist so that a == on two strings doesn't leak the prefix length of a valid MAC.
Note the two idempotency layers, because they're doing different jobs. The delivery id keeps a redelivered event from stacking top-ups on your side; the idempotency key on the write keeps a retried HTTP call from applying twice on the provider's side. You want both, and platforms that specify a dedup window — 24 hours is a common default — are the ones where the second layer is actually load-bearing.
Verifying the rollout without a window of refused traffic
Rotate the secret like you rotate any other credential, which means overlapping and never big-bang. Add the new secret as a second accepted value in the receiver, deploy, then update the registration to sign with it, then drop the old one on the next deploy. The signature header carrying a list rather than a single value is exactly what makes that overlap possible, and it's why I'd take a scheme that supports multiple active signatures over one that doesn't.
Before you trust it in production, do three things: send a test delivery and confirm a 2xx, replay that same signed body with the timestamp shifted past the tolerance window and confirm it's refused, and flip one byte of the signature to confirm the refusal is coming from the MAC check and not from something incidental. If all three behave, the primary check is real. If the replay succeeds, your window isn't being enforced and you have signature verification without replay protection, which is a weaker thing than it looks.
Keep the secret out of the repo and out of the deploy script — AWS Secrets Manager, HashiCorp Vault and Doppler all do this fine, and OWASP's guidance on rotation and storage is worth the ten minutes. A secret set once at launch and never touched again is a secret nobody can audit and nobody can revoke in a hurry.
The rollback is the part people skip. Mine is a feature flag that puts the receiver into verify-and-log mode: signatures are still checked and refusals are still recorded, but the top-up path is disabled and the balance is topped up manually from the runbook. It trades unattended operation for a few hours of human attention, which is the right trade when you're unsure whether a refusal spike is an attack or your own rotation halfway through.
Write that manual top-up step down before you need it.
The catch with all of this: it's a design for one account's own events, where the receiver is yours and the spend is yours. If you're the sender rather than the receiver — publishing events to customer endpoints at volume — the problem inverts into delivery guarantees, per-subscriber secrets and replay UIs, and that's when a dedicated webhook platform earns its cost. Different problem, different tool.
References
- Standard Webhooks specification — https://www.standardwebhooks.com/
- Stripe: verifying webhook signatures — https://docs.stripe.com/webhooks
- OWASP Secrets Management Cheat Sheet — https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
- Go standard library, crypto/hmac — https://pkg.go.dev/crypto/hmac
- Node.js crypto.timingSafeEqual — https://nodejs.org/api/crypto.html#cryptotimingsafeequala-b
Top comments (1)
The ordering point is the one that usually gets skipped: verify the signature over the raw bytes before anything parses them. I had a webhook receiver where the JSON decode happened first "for convenience", and the signature check was silently operating on a re-marshaled body — everything worked until a provider changed field ordering and half the events started failing auth. Moving verification to the raw body fixed it in one line, but it took a real incident to get there.
The media-shop framing makes it concrete though — when a verified event can trigger a top-up, the verify-then-parse-then-spend chain is a security property, not a style choice. The 5-minute timestamp tolerance plus versioned signatures for rotation overlap is a clean pattern; do you also replay-check event ids before the spend step, or does idempotency live downstream?