TL;DR: reject a signed webhook unless its MAC matches the exact bytes received, and make that check happen before any JSON middleware. For an edtech account platform, rehearse key leakage by overlapping old and new secrets for a bounded window, measuring refused valid traffic, then removing the old secret. Set a spend ceiling for retained requests and duplicate processing before the drill; do not quietly weaken verification to keep acceptance high.
The reason is mechanical. A sender signs bytes, not the object those bytes might represent. Parsing and re-encoding JSON can change whitespace, key order, escaping, or numeric representation while preserving the apparent data, so a verifier that sees reconstructed JSON is checking a different message. Worse, a global parser may consume the request stream before route-specific verification gets a chance to read it.
For Express, mount the webhook route with a raw-body parser before mounting the application's JSON parser. The verifier should receive a Buffer, authenticate it, and only then call JSON parsing. Keep this route narrow. Applying raw buffering to every account-platform request expands memory pressure without improving the signature boundary.
Why must webhook signature verification use the raw body?
Suppose a lesson-completion event arrives twice during a retry, once formatted compactly and once with spaces. Those bodies can decode to the same object, but they are different byte sequences. A MAC calculated over one cannot authenticate the other. This is desirable: the signature binds the transmitted representation as well as its meaning.
Middleware order therefore becomes security policy. The safe sequence is body-size enforcement, raw-byte capture, signature parsing, MAC calculation, constant-time comparison, JSON decoding, schema validation, and finally business processing. Authentication does not prove that an event is fresh, unique, or valid for the account; timestamp bounds and an event identifier belong in the surrounding protocol when the sender defines them.
Fail closed. A missing or malformed signature should not reach enrollment, roster, or identity state, even if the JSON itself looks plausible. Record the rejection category without recording the secret or full student payload.
Bytes first.
Put the byte boundary in one small component
The following Go middleware shows the boundary without tying it to a commercial service. It assumes the sender's documented contract is HMAC-SHA-256 with a lowercase hexadecimal signature. If the sender specifies another encoding or signed-message format, implement that contract exactly; guessing is not interoperability.
package webhook
import (
"crypto/hmac"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"io"
"net/http"
)
const maxBody = 1 << 20 // 1 MiB operational limit for this receiver.
func Handler(secret []byte, accept func(map[string]any) error) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
body, err := io.ReadAll(http.MaxBytesReader(w, r.Body, maxBody))
if err != nil {
http.Error(w, "invalid body", http.StatusBadRequest)
return
}
provided, err := hex.DecodeString(r.Header.Get("X-Webhook-Signature"))
if err != nil {
http.Error(w, "invalid signature", http.StatusUnauthorized)
return
}
mac := hmac.New(sha256.New, secret)
_, _ = mac.Write(body)
if !hmac.Equal(provided, mac.Sum(nil)) {
http.Error(w, "invalid signature", http.StatusUnauthorized)
return
}
var event map[string]any
if err := json.Unmarshal(body, &event); err != nil {
http.Error(w, "invalid json", http.StatusBadRequest)
return
}
if err := accept(event); err != nil {
http.Error(w, "processing failed", http.StatusInternalServerError)
return
}
w.WriteHeader(http.StatusNoContent)
})
}
Two limits matter here. The one-mebibyte cap is an example operating decision, not a universal safe value; derive the real limit from the sender's documented maximum plus measured headroom. Also, this sample has one secret to keep the byte-handling path visible. The drill needs a key set, with explicit identifiers or a bounded attempt against active keys, so rotation does not create an outage.
In Express, avoid a global express.json() call before this route. Mount express.raw({ type: ... }) on the webhook path, verify the resulting bytes, then decode them. A parser's verify hook can expose the original buffer too, but it couples authentication to shared parser configuration; a dedicated route makes ownership and capacity easier to audit.
The limitation of route-local raw buffering is that the application now owns body limits, signature-format compatibility, rotation state, rejection telemetry, and the failure behavior between authentication and durable processing. It is not appropriate when the team cannot carry that on-call load. A managed ingress can move some of those duties outside the application, while an owned verifier preserves control over protocol details and data custody; either choice still needs evidence that exact bytes are preserved and invalid messages fail closed. The trade-off is concrete: operational delegation reduces code and pager surface, while another network hop and control plane can add dependency, cost, and lock-in. Do not infer the right answer from a feature list. Estimate peak request bytes during dual-key verification, the retry burst after a school-system outage, the retention period allowed for diagnosis, and the staffing needed to rotate a key under pressure, then choose the boundary that fits both the error budget and spend ceiling.
Run the leaked-key drill against an error budget
Use a synthetic tenant containing no learner data. At minute 0, declare the current signing secret exposed and start the clock. Install a replacement through the same secret-management path used in production, allow both keys only for the predeclared overlap, send fixtures signed by each key, and then retire the exposed key. After retirement, its fixtures must be refused while replacement-key fixtures continue through schema validation.
That sounds obvious. It still needs numbers.
One overlooked edge can invalidate the exercise: if synthetic events bypass the same ingress and secret-distribution path as production events, the drill proves the fixture harness, not the rotation procedure.
Track accepted events by key identifier, authentication refusals by reason, body-limit refusals, end-to-end latency, duplicate event identifiers, and queue depth. The drill passes only if the service SLO remains inside its error budget and the observed resource use stays below the spend ceiling. A rise in refused traffic is not automatically a failure: refusals from the deliberately retired key are the expected security result. Valid replacement-key refusals are the damaging signal.
Use a decision table before anyone is under pressure:
| Decision | Prefer managed handling when | Prefer owned handling when |
|---|---|---|
| Secret custody | Existing controls already cover access, audit, and rotation | The team can operate equivalent controls and needs local custody |
| Verification boundary | The service preserves exact bytes and exposes rejection telemetry | Protocol variation requires code-level control |
| On-call load | Rotation and retry behavior are contractually clear | The team accepts pager ownership for parser, queue, and key-set failures |
| Lock-in | Exportable events and portable signature contracts exist | A standards-based internal interface is more valuable than reduced operations |
| Spend ceiling | Refusal and replay costs are observable and capped | Predictable volume justifies reserved capacity and engineering time |
This is not a feature contest. The right side changes with event volume, staffing, regulatory obligations, and the cost of refusing a legitimate school update. Capacity planning must include the overlap window because every additional active key may add verification work, and retained raw payloads increase both storage exposure and cost.
Verify rollback without reopening the leak
Rollback should restore delivery, not restore trust in the exposed secret. If replacement-key traffic fails, pause downstream mutation, preserve bounded metadata for diagnosis, and correct the new-key distribution path. Re-enabling the leaked credential converts an availability incident into a known authentication weakness.
Before ending the drill, prove four conditions: altered-body fixtures are refused; malformed signatures do not panic or reach JSON decoding; duplicate valid events do not repeat account mutations; and logs contain neither signing material nor full student payloads. Then remove the overlap configuration and confirm that only the intended active-key set remains.
The operational rule is blunt: spend may cap retention, concurrency, and replay depth, but it must not buy acceptance of unauthenticated traffic. If the platform cannot stay within both that ceiling and its valid-event SLO during rotation, reduce the overlap workload or add capacity before the next drill. Do not move the parser ahead of the verifier.
Top comments (0)