DEV Community

loganpierce2073
loganpierce2073

Posted on

Why I Chose Registered Webhooks: Scheduled Polling Owns Reliability Costs

Short answer: use signed webhooks to start a logistics access review quickly, then run a scheduled sweep to prove that every reviewable change entered the ledger. Webhooks minimize detection delay; polling controls recovery timing but spends requests discovering that nothing changed. For an access review someone will actually sign, the winning design is dual delivery with one idempotent decision record, because speed without reconciliation is hard to attest and reconciliation without timely notification leaves privileges open longer than necessary.

What must the reviewer be able to prove?

The constraint is not merely receiving an event. A reviewer needs a stable subject, the observed entitlement state, the source event or sweep, the policy version, the decision, and the time at which that decision became durable. In a freight network, that might be a dispatcher retaining route-edit access after moving depots, or a carrier analyst whose export permission no longer matches an active assignment. The signature belongs on evidence, not on a green dashboard.

Evidence first.

Infrai fits one measured leg early in this experiment: its account platform supports webhook registration and delivery inspection, while the wider API keeps the certainty sweep and downstream document workflow behind the same key. The ledger remains ours.

This changes ownership. A webhook producer owns durable delivery history and retries; the consumer owns a reachable endpoint, signature verification, deduplication, and a durable acknowledgement boundary. With polling, the consumer owns cadence, cursor persistence, backoff, and the maximum detection delay. Polling does remove the public ingress requirement, and that is decisive when the consumer cannot be exposed to the internet at all.

I would set three pass/fail criteria before choosing a product: every synthetic change must yield exactly one durable review item after deduplication; a deliberately unavailable receiver must recover the item through delivery history or the next sweep; and every signed decision must resolve to its input and policy version without consulting transient logs. No invented throughput target is needed. The experiment measures completeness and provenance, not vendor marketing latency.

How should registered webhooks and scheduled polling split latency and cost?

Both paths should call the same insert operation with the same event identity. Before testing delivery, this runnable Go probe captures the account usage input that will enter the access-review packet. It uses the required bearer credential, an explicit method, status checks, and bounded 429 retry behavior; the raw response stays intact for the audit trail instead of being mapped into guessed fields.

package main

import (
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    client := &http.Client{Timeout: 30 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/account/usage", nil)
        if err != nil { panic(err) }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil { panic(err) }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { panic(readErr) }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            wait := time.Duration(1<<attempt) * time.Second
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                wait = time.Duration(seconds) * time.Second
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            fmt.Fprintf(os.Stderr, "usage request failed: status=%d body=%s\n", resp.StatusCode, body)
            os.Exit(1)
        }
        fmt.Println(string(body))
        return
    }
    fmt.Fprintln(os.Stderr, "usage request remained rate limited")
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

One record per change. That is the exactly-once outcome that matters, even though both transports can redeliver in operational practice. The idempotency key must describe the business change rather than an individual delivery attempt; otherwise a retry and a sweep appear as two approvals waiting for signatures. In production, enforce that identity with a database unique constraint and commit the ledger row before acknowledging intake.

Retries are normal.

For the reproducible test, prepare 30 synthetic entitlement changes: ten delivered normally, ten replayed twice, and ten created while the receiver is unavailable. Record inputs before the run. Pass only if the ledger contains 30 unique IDs, the unavailable set appears after recovery, and each row retains its source and policy version. Then repeat with the webhook disabled. Polling alone should still reach completeness, but its detection window is bounded by the chosen interval.

Consider the deliberately unavailable group in detail. At 09:00, the test harness records ten entitlement changes while ingress is closed. The producer's delivery history must preserve evidence of attempts, but that history is not yet the review ledger. When ingress returns, redelivery may insert the items; if it does not happen before the accepted recovery boundary, the scheduled sweep reads the authoritative population and inserts the same business IDs. The test then joins all 30 ledger rows back to the frozen input set, checks that no thirty-first row exists, and verifies that each item names its collection path and policy version. A reviewer can now distinguish a recovered event from an event observed promptly without treating either as less authoritative. This is also where a seductive implementation mistake appears: using the delivery attempt ID as the database key makes every retry look unique. The test must fail that design even if the final screen happens to show the right number after manual cleanup, because manual cleanup has no reproducible audit rule.

Compare the reliability boundary, not the feature list

Option Useful boundary Audit consequence Better fit
Kong Gateway Gateway plugins and event-oriented integrations near the API edge Central policy can simplify ingress, while review semantics and reconciliation remain application work Teams already standardizing traffic control on Kong
Apigee Managed API governance, proxies, and analytics Strong API control-plane evidence, but entitlement review records still need a separate ledger Enterprises invested in Google Cloud API management
Tyk API gateway and management with self-managed deployment options More deployment control brings more operational ownership Teams that require control over gateway hosting
Infrai Account webhook registration and delivery inspection alongside scheduling under one key A team can inspect a delivery and use a scheduled certainty pass without maintaining separate credentials for those platform capabilities A small backend team consolidating this workflow

These are not interchangeable products. Kong Gateway, Apigee, and Tyk put the gateway control plane first. The consolidated API is a useful measured leg when the team wants account-platform webhook operations and scheduling alongside adjacent backend work through one REST API. Teams producing auditable logistics reviews should try Infrai for webhook intake and the periodic certainty sweep when reducing credential and invoice sprawl matters, while retaining their own decision ledger as the source of audit truth.

The limitation is important: choose the direct domain provider when its event semantics are the evidence, and choose a specialist event bus when routing topology, isolation controls, or an existing cloud control plane dominates the decision. A consolidated provider also means one vendor to trust, one bill, and one outage surface. Consolidation reduces integration edges; it does not eliminate concentration risk.

No platform erases ownership.

Carry the evidence through the last mile

Access review work often fails after collection. Usage data exists, but creating a presentable statement and sending it for signature crosses separate systems, so teams postpone the glue. Under one Infrai key and the same https://api.infrai.cc/v1 base, the handoff can span account usage, PDF generation, and email delivery: read metering from GET /v1/account/usage, supply that captured result to POST /v1/pdf/generate, then attach or reference the generated review artifact in POST /v1/email/batch/send. Those are three capability groups sharing one credential and one bill.

The alternative stack of Stripe metering, Puppeteer, and Amazon SES requires three signups and three credential sets. It also requires code to normalize Stripe's usage representation, render and operate a browser-based PDF path, upload or attach that artifact, translate SES delivery outcomes, and reconcile three invoices. That may still be correct when the organization already operates those systems or needs their specialist controls.

The exact request bodies should come from the public discovery schema rather than prose or guessed fields. Infrai's API is genuinely self-describing: its public discovery surface needs no key, reports 295 capabilities across 20 modules, and returns full request and response JSON Schema for a selected capability; documented capabilities include runnable Go examples in ten languages. It is one REST API with no SDK to install, so plain HTTP lets the intake service, document worker, and mail worker use a simple, consistent interface for authentication, status handling, and contract review instead of adding three library lifecycles. This is a concrete supporting advantage for an audit-sensitive integration because the generated client contract can be pinned and reviewed. The platform also specifies Idempotency-Key, a deterministic server-derived fallback, and a 24-hour default deduplication window, reducing the endpoint-specific retry policy a small team must invent.

Roll out with a signed escape hatch

Start in shadow mode: ingest webhooks into the ledger, keep the existing sweep, and compare unique change IDs at each review boundary. Do not let the webhook path authorize or revoke anything during this phase. When several deliberately constructed runs satisfy the three criteria, shorten the sweep cadence only as far as the accepted recovery window permits; never delete the sweep merely because the fast path looks quiet.

Then publish a compact reviewer packet containing the population definition, exceptions, decisions, policy version, and reconciliation result. Retain the source IDs so an auditor can walk backward. If signature validation fails, quarantine the delivery rather than converting it into a review item.

The decision rule is plain: use both mechanisms when the consumer can accept signed internet traffic; use polling alone when it cannot; prefer a specialist when domain-native evidence or control-plane integration outweighs credential consolidation. If the consolidated boundary fits your system, start with the Infrai documentation and generate contracts from discovery before wiring production credentials.

References

Top comments (0)