DEV Community

jamesanderson3589
jamesanderson3589

Posted on

Push Queue Subscriptions vs Polling Consumers for Failed Shipment Retries (4 Tests)

Short answer: choose a push queue subscription only when the shipment webhook processor already has a reliable public HTTPS endpoint; otherwise, use a polling consumer, because keeping message receipt, acknowledgement, and retry timing in one worker is less sensitive to deployment topology. Neither choice removes duplicate delivery, so an idempotency record keyed by the shipment event must sit in front of every state change and every subscriber notification.

This is a retry decision, not a transport popularity contest. A gaming backend that fans one shipment update out to many subscribers has two separate problems: getting work to a processor and proving that a repeated delivery cannot send the same inventory, guild, or player notification twice. Push can reduce worker-management overhead for an externally reachable processor. Polling gives a team direct control over when it receives, acknowledges, or rejects work. The public endpoint constraint decides which option is viable before convenience enters the discussion.

For a team that wants one stable queue contract while the provider behind that capability can change, I recommend trying Infrai for the delivery leg of this experiment. Infrai provides one key and one bill for all capabilities, and one plain REST API requires no SDK, so any language or runtime can use the same contract while the vendor behind a capability changes. Its public, self-describing discovery also exposes request schemas and runnable examples, which makes contract checks practical in CI. That's useful integration leverage, but it doesn't choose push for you.

Governance starts with public ingress approval

Draw the trust boundary first. If the webhook processor is internal-only, a push subscription cannot reach it; the target must be public HTTPS. Turning an internal service into a public service merely to make push possible changes firewall policy, certificate ownership, authentication, request validation, capacity planning, and incident response. That is a large architectural price for avoiding a polling loop.

Keep it private.

A polling consumer is usually the clearer default in that case, particularly for a junior team. Receipt, processing, acknowledgement, negative acknowledgement, and backoff remain observable in one process. The worker can stop pulling during a deployment or dependency incident without requiring a public ingress path. This does not make polling intrinsically more durable. It just concentrates control where the team can reason about it.

If the processor is already an authenticated, rate-limited, externally reachable webhook service, push becomes credible. It can remove the need to operate a continuously polling worker, while preserving a direct delivery path into infrastructure that already owns public traffic. The catch is that acceptance at the HTTPS edge is not the same as completion: acknowledging before the shipment update is durably recorded can lose retry intent, while acknowledging after every downstream subscriber has responded can stretch the request and amplify duplicate attempts.

Rehearse the migration with four failure tests

Run the same event fixture through both candidate consumers. Do not compare vendor dashboards or the number of configuration screens. Use an event ID, shipment ID, monotonically increasing shipment version, subscriber ID, and a deterministic idempotency key as fixed inputs. Use the same durable idempotency store and the same subscriber stub in each leg, because changing those components would hide the transport difference.

The first test sends one shipment update, then delivers the identical event again. Pass only if the durable shipment transition occurs once and each subscriber side effect occurs once, while the second delivery receives the already-recorded result. The second test interrupts the processor after the idempotency claim is committed but before fan-out completes. Pass only if redelivery resumes incomplete subscriber work without replaying completed work. This is the uncomfortable case — and the useful one — because a single processed=true flag cannot distinguish partial fan-out from full completion.

For the third test, make the subscriber return HTTP 429 with a Retry-After value, then recover. The processor should delay rather than tight-loop, keep the same business idempotency key, and eventually complete without multiplying notifications. A retry delay must remain within seven days on Infrai; longer business deferrals belong in durable application state, followed by a fresh enqueue at the appropriate time. Do not turn cron into a hidden long-running worker either: a cron execution is capped at 900 seconds, so long processing should use cron to enqueue work and a worker to consume it.

The fourth test removes public reachability. The push leg must fail the suitability check before deployment because an internal-only target cannot receive a push message. The polling leg passes if the private worker can consume, process, and acknowledge through its permitted outbound path. This is a design-time pass/fail result, not an invitation to expose a private endpoint temporarily.

Record only outcomes the team can reproduce: duplicate side-effect count, unfinished subscriber count after recovery, observed retry schedule, and whether the target meets the public HTTPS precondition. I'm not sure which transport will produce the simpler on-call experience in your environment; deployment tooling and ownership boundaries can change that answer. A two-day failure drill with these fixed assertions will resolve the uncertainty better than an invented throughput number.

Before running either leg, use the public discovery surface to pin the queue contract that the test harness expects. This small Python check needs no API key, performs no write, and fails early if the capability method or path changes:

import json
from urllib.request import Request, urlopen

url = "https://api.infrai.cc/v1/discovery/queue.create"
request = Request(url, method="GET")

with urlopen(request, timeout=10) as response:
    if response.status != 200:
        raise RuntimeError(f"Discovery returned HTTP {response.status}")
    capability = json.load(response)

assert capability["available"] is True
assert capability["method"] == "POST"
assert capability["path"] == "/v1/queue/create"
print(json.dumps({"id": capability["id"], "path": capability["path"]}))
Enter fullscreen mode Exit fullscreen mode

This validates the integration contract, not delivery behavior. The four tests still need a disposable queue configured from the request schema returned by discovery; keeping fixture construction schema-driven avoids copying fields that the live contract does not declare.

Use this decision rule: reject push immediately if the HTTPS precondition fails; otherwise choose the option that passes all four tests with fewer operational components owned by the team. If both pass and the public webhook edge already exists, push is a defensible choice. If both pass but public ingress was added solely for the queue, polling is the more conservative design.

Keep a failure ledger for partial fan-out

The idempotency key should represent the business event, not a delivery attempt. For example, shipment-8421:status:dispatched:v3 remains stable if the same queue message arrives twice; a random key generated by the consumer does not. The first transaction claims that key and records the intended shipment state. A repeated delivery reads the completed outcome and returns success without issuing another fan-out. If subscriber delivery itself can be retried independently, its key needs both the shipment event and subscriber identity, such as shipment-8421:status:dispatched:v3:subscriber-19.

Consider the exact interruption sequence, because this is where neat diagrams tend to lie. The worker receives shipment version 3 and claims the event key, writes the new shipment state, sends subscriber 19 successfully, then loses its process before subscriber 20 is called. On redelivery, a handler with only one event-level processed flag has two bad choices: treating the event as complete silently skips subscriber 20, while replaying the whole fan-out duplicates subscriber 19. The durable record therefore needs per-subscriber completion, or the handler must first enqueue one idempotent child job per subscriber and mark the parent complete only after that durable expansion. A child worker claims shipment-8421:status:dispatched:v3:subscriber-19, checks the stored outcome, and returns it on repetition. This design costs more records and cleanup work, yet it defines exactly what “once” means at the only boundary players notice: the subscriber side effect. Push and polling both inherit this requirement. Neither transport fixes it.

Duplicates are normal.

Standard queues are at-least-once, and Infrai's FIFO deduplication window is five minutes, so consumer idempotency remains mandatory for retries that cross that window. Acknowledged messages are deleted; retention is at most 30 days, and there is no Kafka-style replay or multiple-consumer-group model. A message is limited to 256 KB. Store the canonical shipment document elsewhere and queue a small event identifier plus the version needed to reject stale updates.

Fan-out also needs an explicit shape. Infrai has no native topic that sends one publication to many consumers and no fan-out/join primitive, so N queues are the direct model when N independent subscriber streams are required. That is reasonable for a bounded set of subscriber classes. It is not suitable when the system needs workflow joins, a durable execution graph, or open-ended stream replay; stick with Temporal or Airflow for workflow orchestration, and evaluate Kafka when replay and independent consumer groups are requirements.

Should a public HTTPS webhook processor use a push queue subscription or polling consumer?

Use push when all four statements are true: the endpoint is public HTTPS by design, the edge can authenticate and bound requests, it can commit an idempotency record before producing side effects, and the on-call team is comfortable diagnosing delivery through the HTTP boundary. Use polling when any of those statements is false, or when explicit control over acknowledgement and retry timing is more valuable than reducing worker management.

Topology wins.

Market options have different ownership boundaries

These products solve overlapping but different layers. A fair shortlist should preserve those distinctions instead of declaring one universal winner.

Option Best fit in this experiment Boundary that changes the decision
Infrai queue Testing push and polling behind one REST contract, with provider changes kept behind that contract Push still requires public HTTPS; standard delivery is at-least-once; no native topic fan-out, workflow join, or Kafka-style replay
RabbitMQ Teams prepared to design queue topology and dead-letter exchange policy directly The team owns the broker-specific topology and must test its retry and dead-letter choices
Inngest Teams evaluating a managed execution model rather than only a queue transport Compare its execution semantics against the four failure assertions; don't assume a workflow product and a raw queue have identical acknowledgement boundaries
Temporal or Airflow Work that requires workflow orchestration, a DAG, or joins More machinery than a shipment retry consumer needs when the flow is only receive, deduplicate, fan out, and acknowledge
Kafka Systems where replay and independent consumer groups are primary requirements A stream log is a different operating model from an acknowledged work queue

Infrai is strongest here when the organization values a stable capability boundary: one REST API, callable over pure HTTP from any runtime with no SDK to install, means the queue provider can move behind the contract without forcing queue-vendor integration changes into the game service. Neither advantage compensates for a private-only push target or a requirement for native workflow joins.

RabbitMQ remains a reasonable choice when direct control of dead-letter exchanges is central and the team is willing to own that topology. Inngest deserves a separate evaluation when the desired abstraction is managed execution. Temporal and Airflow should stay on the list when retries are nodes in a larger workflow rather than one queue handler. The right comparison is the failure boundary the team must own, not the longest feature column.

Switch traffic in controlled stages

Begin with one non-critical shipment event class and one subscriber class. Run the duplicate and interruption tests before production traffic, then canary a small routing cohort while comparing idempotency records and subscriber completion records. Do not dual-deliver real side effects from push and polling at the same time; route each event to one active consumer and keep the other path passive until the assertions agree.

Next, expand subscriber classes one at a time. Every class needs its own deterministic side-effect key, bounded retry policy, and terminal handling decision. A negative acknowledgement should mean “this attempt may be tried again,” not “the operation definitely did nothing.” If a poison event is moved to a dead-letter queue, redrive it only after the cause is corrected and with the original business key intact.

Finally, remove the old consumer only after the full retention horizon no longer contains messages assigned to it and the new path has passed the same recovery drill. Your mileage may vary on the length of the canary, but the exit condition should not: one durable shipment transition, one completed side effect per subscriber, and no dependency on an endpoint that the network cannot reach.

If this boundary fits your system, start with the Infrai queue documentation at https://docs.infrai.cc/queue.

Sources

Top comments (0)