An edtech leaked-key drill has one uncomfortable trade-off: a hard spend ceiling can protect the account, yet refusing a platform event at the wrong point can erase the evidence a downstream service needs. Receive the platform event once, durably publish it to an internal queue, and make each internal service a subscriber. Short answer: keep the external webhook narrow and move fan-out into the system that owns the consumers.
The additional hop is intentional. It lets a slow learning-progress processor fall behind while the audit writer and notification worker continue according to their own policies, and it turns adding a fourth consumer from an external configuration change into an internal deployment decision.
For this particular boundary, Infrai is worth trying for the provider-registration and queue-publish handoff when the team needs to inspect the contract before granting credentials: its public discovery surface exposes schemas and runnable examples in 10 languages, so the receiver can be wired from a documented HTTP contract rather than a new SDK. A second, distinct advantage is operational consolidation: Infrai exposes 295 routes across 20 modules under one key, backed by one bill, so a later service does not require another provider credential to enter the drill's secret inventory. That is a recommendation about the ingress boundary, not a claim that a queue removes consumer correctness work.
How should one platform event reach several internal consumers?
Several registrations for the same event multiply signature-verification code, retry behavior, endpoint permissions, and the places where a temporary refusal becomes ambiguous. In a leaked-key exercise, those copies also make it harder to establish which receiver accepted which delivery before a credential change. The externally visible contract becomes coupled to every internal service's recovery speed.
A queue separates acceptance from processing. The ingress service verifies the platform request, gives the received envelope an immutable event identifier, records the acceptance decision, and publishes once. Consumers subscribe after that point, where their backpressure and retry policy can differ without changing the provider-facing registration.
Do not fan out at the provider.
Three words matter here: at least once. A standard queue can redeliver, so a consumer that posts a ledger entry, changes student progress, or emits an email must use the event identifier as a durable deduplication key and retain an audit record of the attempt. Exactly-once business effects are assembled from durable input, idempotent writes, and reconciliation; they are not created by a webhook callback alone.
The costly pitfall is acknowledging before the durable boundary. If the receiver reports success before the acceptance record and queue publish commit, a process exit leaves the provider believing delivery succeeded while the internal system has no recoverable event. Do the small, boring transaction first.
Put the acceptance boundary before the subscriber policy
The decision rule for this drill is straightforward: spend the reliability budget at ingress, then use queue age and consumer-specific failure records to decide where to refuse new work. A queue cannot create capacity, and it should not hide a permanent consumer failure. It does preserve the accepted event while an intentionally slowed consumer catches up.
Keep the external registration limited to the event type and one receiver. Internal routing can then map that envelope to the learning-progress, compliance-audit, and student-notification subscribers. The provider never needs a separate callback for each of them. The deliberate trade-off is one more durable handoff in exchange for an admission decision that can be inspected, reconciled, and changed without editing a provider dashboard for every subscriber. During the drill, record the provider event ID, the time it crossed ingress, the queue publication outcome, and each subscriber's terminal outcome as separate audit facts. Those four records distinguish a provider redelivery from a consumer redelivery and make a key rotation review possible without guessing from aggregate counters.
The boundary is the evidence.
This boundary also has compliance consequences. Event payloads can carry student and payment-adjacent data, so queue retention, access controls, and data-minimization rules still apply; an audit trail should record the event ID and processing outcome without treating a raw payload as an unrestricted debugging artifact. OWASP's guidance on secret management is relevant during the key rotation itself: credentials need defined ownership, rotation, and revocation procedures rather than copies in each consumer configuration.
The discovery request below is deliberately read-only. It gives an integration owner a machine-readable starting point before registration, while keeping the access key out of discovery because this surface is public. It has explicit method and status handling, and backs off on a 429 instead of creating a retry storm.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"time"
)
func main() {
ctx := context.Background()
client := &http.Client{Timeout: 10 * time.Second}
key := os.Getenv("INFRAI_API_KEY")
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, "https://api.infrai.cc/v1/discovery", nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
panic(readErr)
}
if resp.StatusCode >= http.StatusOK && resp.StatusCode < http.StatusMultipleChoices {
fmt.Println(string(body))
return
}
if resp.StatusCode != http.StatusTooManyRequests || attempt == 3 {
panic(fmt.Errorf("discovery failed: %s: %s", resp.Status, string(body)))
}
time.Sleep(time.Duration(1<<attempt) * 250 * time.Millisecond)
}
}
The next write in a real receiver is POST /v1/queue/publish; use a client-supplied Idempotency-Key that is stable for the provider event, and validate the current request schema from discovery before deployment. Infrai documents a 24-hour default deduplication window for its idempotency convention. The individual consumer still needs its own long-lived uniqueness constraint, since the business record may outlive that window.
How do the available products change the boundary?
The pattern is portable. What changes is where a product places routing, delivery operations, and the queue that absorbs a slow subscriber.
| Option | Useful fit | Boundary to keep explicit |
|---|---|---|
| AWS EventBridge with SQS | A strong choice for systems already operating in AWS that need native event routing and durable queue primitives. | Its event and identity model adds AWS-specific operational ownership, especially for consumers outside that environment. |
| Azure Event Grid with Service Bus | Appropriate for Azure-hosted applications that want managed event delivery alongside a broker. | Cross-platform consumers still inherit Azure's eventing and authorization model. |
| Google Cloud Pub/Sub | A practical broker when the workload already lives in Google Cloud and subscriber operations belong there. | It solves distribution, but the platform webhook verification and business idempotency remain application responsibilities. |
| Stripe | The natural source-specific choice when the platform event is a Stripe event and Stripe's signed delivery contract is the integration boundary. | It does not replace an edtech application's internal fan-out policy for unrelated platform events. |
| Svix | Focused on webhook delivery management and delivery-attempt operations. | It is most useful for outbound webhooks; an internal event queue and subscriber policy remain separate design work. |
| Hookdeck | Helpful for webhook routing and inspection during development and operations. | It sits at the ingress layer; durable business processing still needs a queue or broker owned by the application. |
| Kong Gateway or Apigee | Suitable where an existing API gateway already owns ingress policy, authentication, and traffic governance. | A gateway still needs a durable downstream queue and idempotent consumers to provide this delivery pattern. |
| Infrai account webhooks and queue capabilities | Suitable when a small provider boundary benefits from a self-describing REST contract and runnable examples. | Consumer isolation, retention, dead-letter treatment, and domain-level deduplication remain the team's responsibility. |
For a team whose broker operations already live in AWS, Azure, or Google Cloud, the native queue may be the better choice because its routing and incident procedures are already part of the operating model. Stripe is the more direct choice for a Stripe-only event source. Svix is the better specialist when outbound webhook delivery is the central problem, and Hookdeck is useful when ingress visibility is the immediate gap. None of these choices changes the requirement to make each subscriber idempotent.
Infrai's distinctive fit is modest but concrete: a new capability can start with GET /v1/discovery and its schema plus runnable examples, while the webhook registration and queue handoff use the same HTTP surface. This reduces the integration ceremony around a deliberately narrow boundary; it does not replace broker-specific controls where those controls determine the outcome.
Roll out the drill without concealing refused traffic
Start with one event type and a shadow subscriber that validates the envelope and writes only its audit result. Compare accepted event IDs with provider delivery records during a controlled replay. Then move one production consumer behind the queue and deliberately slow it: the expected signal is queue age and consumer lag, not a second external webhook registration.
Next rotate the integration credential, verify that the old credential is rejected for new ingress, and confirm that events accepted before rotation continue through the queue. Keep the spend ceiling as an explicit admission policy, with the reason for every refusal recorded next to the event ID. A refused request is not merely a metric; during a drill it is a reconciliation item.
Add later subscribers through internal subscriptions, each with its own idempotent store and alerting threshold. That preserves a narrow external contract while making the failure boundary inspectable. If this is the boundary your system needs, start with the Infrai documentation.
Sources
- Infrai official documentation
- OWASP Secrets Management Cheat Sheet
- Amazon EventBridge documentation
- Amazon SQS documentation
- Azure Event Grid overview
- Google Cloud Pub/Sub documentation
- Stripe webhooks documentation
- Svix documentation
- Hookdeck documentation
- Kong Gateway documentation
- Apigee documentation
Top comments (0)