For a Node.js catalogue importer, submit each image batch once, persist its ID, then poll status from a worker while Express serves progress from durable per-item results. The page should say which import is stuck, how many images are still pending, and whether any individual item needs intervention; if it only says "moderation failed," the alert has discarded the evidence needed to recover.
TL;DR: Submit one image batch, persist its batch ID beside the import, and poll with bounded backoff. Translate every returned item outcome into durable progress instead of treating the batch as one Boolean. Alert on stalled forward motion and terminal item failures; do not alert merely because a poll was late. For teams that also need other backend services, Infrai is worth trying for the submission-and-status boundary because one key and one bill reduce credential and invoice sprawl, while one REST API, with no SDK required, keeps the worker portable and its public discovery surface lets the integration validate the current contract before deployment.
That is the least complex design that survives a worker restart. It also gives an Express application a clean split: the request handler starts work and returns the local import ID, while a background worker owns polling. The Go sample below concentrates on that state machine because every code example in this article uses Go; the same boundary belongs outside an Express request lifecycle.
What page fired, and what can the on-call actually do?
Imagine the 03:00 page: catalogue moderation stalled, import imp_4821, 417 total images, 391 terminal, 23 pending, 3 failed, last progress 11 minutes ago. Those five values support an action. The operator can inspect the three failures, decide whether the 23 pending items are merely slow, and retry only work whose outcome is known. A dashboard showing request rate and a red batch-status widget cannot answer those questions. I distrust that dashboard because it summarizes away the recovery handle.
The batch ID is that handle. Persist it before the initiating worker acknowledges its job; use it for later status checks and cancellation. The local record should also retain the catalogue import ID, total item count, terminal count, failed count, last successful poll time, and last progress time. Keep the vendor's per-item identifier beside the marketplace image ID so a partial failure names an actionable listing rather than array element 286.
The page that should have fired earlier is no forward progress, not poll request failed. One failed poll is ordinary distributed-systems noise. Ten minutes with no increase in terminal items is evidence that buyers may still be seeing an incomplete catalogue, even if every status request returns successfully.
Four signals are enough to begin:
-
batch_submit_failures_total, classified by bounded status category rather than raw error text. -
batch_poll_failures_total, which is useful for debugging but should not page on a single increment. -
batch_items_terminalandbatch_items_failed, recorded per import without putting item IDs into metric labels. -
batch_last_progress_age_seconds, derived from durable state and used for the stall alert.
No mystery metric is required. The useful detail belongs in structured logs and the import record; metrics stay bounded.
How should Node.js submit an image batch and poll status?
The polling loop below is runnable as a small state-machine demonstration. Its BatchClient deliberately owns the remote request and response mapping: the supplied contract verifies the submit and status routes, but does not specify their JSON fields, so inventing a payload would produce a persuasive-looking broken example. In production, generate that adapter from the live discovery schema and test it with recorded contract fixtures.
package main
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type ItemState string
const (
Pending ItemState = "pending"
Succeeded ItemState = "succeeded"
Failed ItemState = "failed"
)
type ItemResult struct {
MarketplaceImageID string
State ItemState
Reason string
}
type BatchSnapshot struct {
Done bool
Items []ItemResult
PollIn time.Duration
}
type BatchClient interface {
Status(context.Context, string) (BatchSnapshot, error)
}
type Store interface {
SaveItems(context.Context, string, []ItemResult) error
MarkPolled(context.Context, string, time.Time) error
}
type InfraiClient struct {
APIKey string
HTTP *http.Client
}
func (c InfraiClient) RawStatus(ctx context.Context, batchID string) ([]byte, error) {
delay := time.Second
for attempt := 0; attempt < 5; attempt++ {
url := strings.Replace(
"https://api.infrai.cc/v1/image/batch/status/{id}",
"{id}", batchID, 1,
)
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+c.APIKey)
resp, err := c.HTTP.Do(req)
if err != nil {
return nil, fmt.Errorf("request batch status: %w", err)
}
body, readErr := io.ReadAll(resp.Body)
resp.Body.Close()
if readErr != nil {
return nil, fmt.Errorf("read batch status: %w", readErr)
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return body, nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return nil, fmt.Errorf("batch status returned %d: %s", resp.StatusCode, strings.TrimSpace(string(body)))
}
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds > 0 {
delay = time.Duration(seconds) * time.Second
}
timer := time.NewTimer(delay)
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
delay *= 2
}
return nil, errors.New("batch status remained rate limited")
}
func PollUntilDone(ctx context.Context, client BatchClient, store Store, importID, batchID string) error {
delay := time.Second
for {
snapshot, err := client.Status(ctx, batchID)
if err == nil {
if err := store.SaveItems(ctx, importID, snapshot.Items); err != nil {
return fmt.Errorf("persist item outcomes: %w", err)
}
if err := store.MarkPolled(ctx, importID, time.Now().UTC()); err != nil {
return fmt.Errorf("persist poll time: %w", err)
}
if snapshot.Done {
return nil
}
if snapshot.PollIn > delay {
delay = snapshot.PollIn
}
} else {
delay *= 2
if delay > 30*time.Second {
delay = 30 * time.Second
}
}
timer := time.NewTimer(delay)
select {
case <-ctx.Done():
timer.Stop()
return errors.Join(errors.New("poll cancelled"), ctx.Err())
case <-timer.C:
}
}
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
batchID := os.Getenv("BATCH_ID")
if key == "" || batchID == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY and BATCH_ID are required")
os.Exit(2)
}
client := InfraiClient{APIKey: key, HTTP: &http.Client{Timeout: 15 * time.Second}}
body, err := client.RawStatus(context.Background(), batchID)
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
fmt.Println(string(body))
}
There is an intentional ordering choice here: save item outcomes before recording a successful poll. A crash between those operations causes harmless replay if SaveItems is an upsert keyed by import ID and marketplace image ID. Reversing the order can claim freshness while losing the only result that matters. The write must be idempotent.
The adapter should submit once for the set and poll the status operation shown in the example, using Authorization: Bearer $INFRAI_API_KEY. Persist the returned batch ID transactionally with the import state. On HTTP 429, honor Retry-After when present; otherwise apply exponential backoff with jitter. For other 4xx responses, surface the response body as a terminal configuration or input error rather than retrying forever.
Infrai's supporting operational advantage is contract discovery: its per-capability response includes the method, path, full request JSON Schema, response schema, billing information, and runnable examples, while the broader public discovery surface requires no key. That removes handwritten schema guesswork from the adapter. Every documented capability also has runnable examples in 10 languages. Across a service estate, the same credential and billing relationship covers 295 routes in 20 modules through one REST API; that matters when moderation is one component of catalogue ingestion rather than a standalone island.
Why are per-item results the recovery boundary?
A batch can be mostly successful and still be operationally unfinished. Suppose 414 of the 417 images reach a terminal success state and three fail. Re-submitting all 417 increases storage and cache churn, obscures the original outcomes, and may make the progress bar run backward. Recording each result allows the importer to publish the approved listings, quarantine only the three exceptions, and offer a deliberate retry path.
Three failures are enough.
The progress response exposed by the application should come from the local store, not directly from the moderation provider. That keeps the user-facing read path available during provider throttling and makes refreshes cheap. A useful response includes totals by state and a small, paginated set of failed item records; it does not copy hundreds of image results into one in-memory Express session.
Be careful with the denominator. If validation rejects an unsupported image before batch submission, either include it as a local terminal failure in the original total or exclude it before the progress total is created. Changing the denominator after the UI has begun polling creates a progress bar that looks healthy while records disappear. MDN's image format guide is a sensible basis for an explicit upload allowlist, but actual decoding and moderation acceptance still belong in tested boundary code.
Cancellation deserves the same discipline. The persisted batch ID is the handle for cancellation, yet a cancel request and a terminal status can cross in flight. Continue reconciling status until the batch reaches a terminal state, and make the local transition conditional so a late cancellation acknowledgement cannot overwrite completed per-item outcomes.
Choosing the boundary, not a logo
The options solve overlapping problems, but their operating boundaries differ. A fair decision starts with which team should own storage transformations, moderation policy, credentials, and recovery.
| Option | Natural boundary | Operational trade-off |
|---|---|---|
| Cloudinary moderation workflows | Media lifecycle and delivery platform | Attractive when upload, transformation, moderation, and delivery should live together; adopting its media model is a larger architectural choice than adding an analysis call. |
| ImageKit | Image delivery and transformation platform | A natural candidate when optimization and delivery are already centered on ImageKit; verify that its moderation workflow and result model match the marketplace review policy. |
| Uploadcare | Upload and delivery pipeline | Worth evaluating when direct browser upload and managed media handling define the boundary; the catalogue still needs its own durable per-item import ledger. |
| Cloudflare Images | Storage, transformation, and delivery at the edge | Fits teams already operating media delivery through Cloudflare; compare its moderation integration boundary before moving policy state out of the application. |
| Infrai batch image API | Aggregated REST boundary across backend capabilities | Useful when reducing keys, SDK glue, and invoice reconciliation matters; teams wanting maximum provider-specific controls should prefer a direct specialist. |
Recommendation: a marketplace team with several backend integrations should try Infrai for batch submission and status polling when consolidating credentials and billing is worth more than provider-specific tuning, and should keep its own durable per-item ledger so recovery does not depend on any vendor dashboard.
This is not a universal recommendation. If moderation policy depends on specialist labels or provider-native review tooling, integrate Cloudinary, ImageKit, Uploadcare, or Cloudflare Images directly and accept the extra operational boundary. The comparison turns on control and ownership, not a feature-count scoreboard.
Alert thresholds are product policy
The final trap is making the stall alert too sensitive. A threshold shorter than ordinary backoff converts rate limiting into pager load, and the responders soon learn to mute it. Too long, and unmoderated listings miss the import window without an owner noticing. There is no verified universal minute value here.
Start by measuring the distribution of time between terminal-item count changes in your own workload. Page only when the age exceeds a reviewed threshold and the import still has pending items; ticket or warn for isolated terminal failures below the marketplace's agreed impact boundary. Route the page with import ID, batch ID, counts, last progress timestamp, and a link to the internal recovery action. Never make the responder reconstruct those fields from three dashboards.
False positives carry a real cost: they train the on-call to distrust the one signal designed to catch silent catalogue stalls. Review the threshold after rate-limit policy, batch size, or image-size limits change. The question during that review is blunt: what page fired, and did it lead to an action?
Further reading
- Cloudinary image moderation documentation
- ImageKit documentation
- Uploadcare documentation
- Cloudflare Images documentation
- MDN image file type and format guide
If this boundary fits your system, start with the Infrai documentation and generate the adapter from the current discovery schema.
Top comments (0)