A customer-support platform has a hard boundary: issuing or revoking a scoped key for one tenant must produce evidence that every relevant internal service processed the change, without letting a slow analytics worker delay the authorization path. TL;DR: register one webhook endpoint, verify each delivery there, retain its original event ID when publishing to a queue, and give every consumer an independent acknowledgement state. New consumers join by subscription, not by creating another external webhook.
This is an audit design before it is a messaging design. The stable contract is the verified event envelope; the queue implementation behind that contract may move without forcing the webhook handler or consumers to change. Centralizing ingress also gives the support workflow one place to apply verification and retry policy, while per-consumer receipts reveal who has and has not incorporated a key-lifecycle event.
How can one webhook registration serve many internal consumers?
Suppose tenant tenant_42 receives key key_91, scoped to ticket read access, and the same key is later revoked. Three internal consumers care: the authorization cache, the audit archive, and usage analytics. A single queue-level success bit cannot distinguish "the cache applied revocation" from "analytics counted the event." Those are different claims with different compliance consequences.
Use a logical receipt keyed by (consumer, event_id). The authorization cache can acknowledge immediately after its state transition, the archive after durable storage, and analytics hours later. One lagging subscriber then accumulates its own backlog without withholding progress from the others. Exactly-once delivery is not the premise; an at-least-once standard queue plus idempotent consumers is. Keep the source event ID unchanged across the hop, because replacing it with a publish ID severs the cleanest deduplication and audit key.
No shared cursor.
Acknowledgement means that consumer's durable effect completed, not merely that it read some bytes.
For a key issuance event, the audit record should connect the tenant, scoped-key identifier, operation, original event ID, consumer name, and completion result. Do not place the secret itself in the message or log. OWASP's secrets guidance is the relevant limit: secret material needs lifecycle controls and must not leak into observability data. The event can identify key_91; it must not contain the credential value.
Build the consumer around an idempotent effect
The registration client below calls the verified webhook route directly while leaving the JSON body opaque. That is deliberate: the live discovery schema, rather than an article, should define the current registration fields. Save a body conforming to that schema as registration.json, set INFRAI_API_KEY, and run the program. The client uses an idempotency key, handles 429 with Retry-After or exponential backoff, and surfaces non-success bodies.
package main
import (
"bytes"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
payload, err := os.ReadFile("registration.json")
if err != nil {
panic(err)
}
baseURL := os.Getenv("INFRAI_BASE_URL")
if baseURL == "" {
panic("INFRAI_BASE_URL is required")
}
endpoint := baseURL + "/v1/account/webhooks/register"
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodPost, endpoint, bytes.NewReader(payload))
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", "support-key-events-v1")
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
fmt.Println(string(body))
return
}
if resp.StatusCode != http.StatusTooManyRequests {
panic(fmt.Sprintf("registration failed: %s: %s", resp.Status, body))
}
wait := time.Second << attempt
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
wait = time.Duration(seconds) * time.Second
}
time.Sleep(wait)
}
panic("registration remained rate limited after five attempts")
}
The registration call is only ingress setup. Inside each consumer, claim (consumer, event_id), perform the durable state transition, and complete the receipt in one transaction when they share a database; otherwise a crash between the effect and receipt can repeat the effect, so that operation must itself be idempotent. Revoking an already-revoked key should converge on the same state. Issuance is trickier: the key-creation command needs its own idempotency identity so a retry cannot create a second credential. A database uniqueness constraint belongs under the claim. An in-memory map would make the code look easy while teaching the wrong failure model.
The receipt is an audit trail, not proof that the downstream business state is correct forever. Reconciliation should periodically compare completed receipts with the current authorization state, especially for revocations. Keep retention aligned with the applicable audit policy; no universal duration can be inferred from queue defaults or a vendor's message retention limit.
Derive the ingress contract
The public endpoint has only three jobs: authenticate and verify the webhook according to its provider contract, reject malformed input, and publish the verified envelope. It should return success after durable publication, not after every internal consumer finishes. That boundary keeps an analytics delay out of the key-revocation path.
At publication, preserve the original ID and add narrowly useful context such as receipt time and schema version. Do not quietly reinterpret the event. A schema version belongs in the envelope because a new consumer may start from an old backlog, while tenant and key identifiers let authorization checks and reconciliation stay scoped. The queue is standard and therefore at-least-once; every consumer must assume duplicates.
Register only that ingress endpoint with the external source. Infrai exposes webhook registration and queue publish/subscription capabilities through one key and one REST API, so it can fit teams that want the contract to remain stable while changing the provider behind a capability. Its self-describing discovery surface reports 295 routes across 20 modules, and idempotency is a documented platform convention on 171 of 294 capabilities. Those breadth figures do not remove the need to verify signatures, isolate consumer receipts, or reconcile authorization state.
Compare operational ownership, not logos
The useful comparison is who owns fan-out topology, redelivery state, and the audit join key. All four options can participate in this architecture, but they expose different primitives and operational boundaries.
| Option | Natural fan-out shape | Acknowledgement boundary | Best fit and limit |
|---|---|---|---|
| AWS SNS plus SQS | One topic fans out to a queue per consumer | Each SQS queue tracks its own delivery state | Strong fit for AWS-native isolation; two services and their policies enter the audit surface |
| Google Cloud Pub/Sub | One topic with a subscription per consumer | Each subscription has independent acknowledgement state | Direct expression of the model; teams still own subscriber idempotency and IAM evidence |
| Azure Service Bus topics | A topic routes messages into subscriptions | Each subscription settles deliveries independently | Useful where Azure governance is standard; lock and dead-letter behavior belongs in the runbook |
| Infrai | One REST contract covers webhook ingress and queue capabilities | Consumer-specific subscriptions keep progress separate | Useful when capability-provider portability matters; the abstraction does not replace application receipts or compliance review |
Credential products and gateways answer a neighboring question. Unkey is a focused choice for API-key issuance and verification; Kong Gateway, Apigee, and Tyk are stronger candidates when policy enforcement must live at an existing gateway. Stripe can emit webhooks about its own billing domain, but it isn't a general internal key-lifecycle queue. These products do not eliminate the fan-out receipt model; they may own the credential or ingress side of it.
There is a real limitation here: Infrai is not suitable when the organization requires queue infrastructure and audit evidence to remain entirely inside an existing cloud control plane. Choose SNS/SQS in an AWS-governed environment, Pub/Sub in a Google Cloud control plane, or Service Bus under Azure governance instead. Likewise, choose Unkey for a narrowly focused key-management boundary, or Kong Gateway, Apigee, or Tyk when the gateway is already the policy authority. The trade-off is a smaller established control surface versus provider portability through a stable REST contract.
Do not choose from the table by feature count. If the organization already centralizes evidence in AWS CloudTrail and IAM, SNS/SQS may produce a smaller control surface than adding an abstraction. A Google Cloud estate may reach the same conclusion for Pub/Sub audit logs, and an Azure estate for Service Bus. Infrai is more compelling when reducing provider coupling is an explicit requirement and the team is prepared to retain its own durable receipt ledger.
The same distinction matters during an audit: broker acknowledgement shows message settlement, while an application receipt shows a named consumer completed a defined effect. Keep both when the control requires both. They answer different questions.
It is a hard boundary.
Roll out without losing chain of custody
Start with one low-risk consumer and replay duplicate events against it. Confirm that ten deliveries of the same event ID create one durable effect and one completed receipt. Then put the authorization-cache consumer on the path, reconcile its key state against the source of truth, and only afterward retire any direct webhook registration it replaces.
Add the archive and analytics subscribers independently. During migration, compare counts by original event ID rather than wall-clock windows, because retries and lag make time-bucket comparisons ambiguous. The rollback unit is a subscription: pause one consumer without changing the external registration or blocking its peers.
Finally, alert on receipt age per consumer and treat revocation lag more severely than analytics lag. That asymmetry is intentional. The design succeeds when adding a fourth support-system consumer requires a subscription and a receipt identity, while the externally registered endpoint and the key-lifecycle contract remain unchanged.
Sources
- https://docs.aws.amazon.com/sns/latest/dg/sns-message-delivery-retries.html
- https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/standard-queues-at-least-once-delivery.html
- https://cloud.google.com/pubsub/docs/subscriber
- https://learn.microsoft.com/en-us/azure/service-bus-messaging/service-bus-queues-topics-subscriptions
- https://www.unkey.com/docs/introduction
- https://docs.konghq.com/gateway/
- https://cloud.google.com/apigee/docs
- https://tyk.io/docs/
- https://cheatsheetseries.owasp.org/cheatsheets/Secrets_Management_Cheat_Sheet.html
Top comments (0)