DEV Community

IversonBlake8417
IversonBlake8417

Posted on

Node.js Catalogue Image Batch Status Tracking with Express (Under Bandwidth Constraints)

Submit catalogue images as one batch, persist the returned batch ID, and let a worker poll status with backoff. That is the right shape for an Express service preparing product images for short promo videos because it preserves per-item quality decisions without turning status checks into the main bandwidth consumer.

TL;DR: Treat the batch ID as durable job state, not request-local data. Store each item's outcome separately. Show completed and failed counts side by side. A progress bar built from those records survives browser disconnects, process restarts, and partial failure.

Replace the long request with a durable job

The fragile mental model is one long HTTP request: a browser sends 80 catalogue products, Express waits for every image, and the open connection stands in for job state. Item 47 fails. Now the useful question is not "Did the request finish?" but "Which product needs attention?"

Use three durable nouns instead: import, batch, item.

Here is the diagram in words. The browser creates an import in Express. Express submits one upstream batch. The upstream service returns a batch ID, which Express stores beside its own import ID. A worker polls status and writes each item outcome. The browser reads local state for its progress display.

Short and sturdy.

This separation matters when source images feed promo-video generation. Seventy-nine usable product images and one failed hero image is not an acceptable 98.75% success. It is a completed batch with one actionable failure. Keep the catalogue key next to the result so an operator can repair one listing instead of replaying all 80.

Progress is derived state. Recompute it from persisted item outcomes after every snapshot; do not increment a counter optimistically. Repeated responses and worker retries then cannot push the display past reality.

A copyable TypeScript transport and polling core

The example keeps remote payload decoding at three explicit adapters. Their implementations must come from the current discovery schema rather than guessed field names. The transport, retry behavior, persistence boundary, and Express response are concrete.

It uses plain REST, so there is no media SDK or client-library version to maintain. Infrai is one option with that shape. Its public, unauthenticated discovery surface provides request and response schemas, and documented capabilities include runnable examples in 10 languages.

A separate Infrai advantage is single-key access with unified billing: one credential works across 295 routes in 20 modules, so a team does not have to collect dozens of API keys or reconcile dozens of bills. For a catalogue pipeline that later adds video generation or observability, that means fewer secrets to rotate and fewer service accounts to map back to the same import workflow. The application still keeps its own progress model, so the administrative simplification does not leak into item state.

import express from "express";
import { randomUUID } from "node:crypto";

type ItemOutcome = {
  catalogueKey: string;
  state: "pending" | "succeeded" | "failed";
  result?: unknown;
  error?: string;
};

type StatusSnapshot = {
  terminal: boolean;
  items: ItemOutcome[];
};

type ImportRecord = {
  id: string;
  batchId: string;
  items: Map<string, ItemOutcome>;
};

type PayloadAdapters = {
  buildSubmitBody(value: unknown): unknown;
  decodeBatchId(value: unknown): string;
  decodeStatus(value: unknown): StatusSnapshot;
};

const apiKey = process.env.INFRAI_API_KEY;
const apiOrigin = process.env.MEDIA_API_ORIGIN;
if (!apiKey || !apiOrigin) {
  throw new Error("INFRAI_API_KEY and MEDIA_API_ORIGIN are required");
}

const imports = new Map<string, ImportRecord>();
const app = express();
app.use(express.json({ limit: "2mb" }));

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

function retryDelay(response: Response, attempt: number): number {
  const retryAfter = response.headers.get("retry-after");
  if (retryAfter && /^\d+$/.test(retryAfter)) {
    return Number(retryAfter) * 1_000;
  }
  return Math.min(1_000 * 2 ** attempt, 30_000);
}

async function request(
  method: "GET" | "POST",
  path: string,
  init: Omit<RequestInit, "method"> = {},
): Promise<unknown> {
  for (let attempt = 0; attempt < 6; attempt += 1) {
    const response = await fetch(new URL(path, apiOrigin), {
      ...init,
      method,
      headers: {
        Authorization: `Bearer ${apiKey}`,
        "Content-Type": "application/json",
        ...init.headers,
      },
    });

    if (response.status === 429) {
      await sleep(retryDelay(response, attempt));
      continue;
    }

    const body: unknown = await response.json();
    if (!response.ok) {
      throw new Error(`Upstream ${response.status}: ${JSON.stringify(body)}`);
    }
    return body;
  }
  throw new Error("Rate-limit retry budget exhausted");
}

async function pollBatch(
  record: ImportRecord,
  decodeStatus: PayloadAdapters["decodeStatus"],
): Promise<void> {
  for (let attempt = 0; ; attempt += 1) {
    const raw = await request(
      "GET",
      `/v1/image/batch/status/${encodeURIComponent(record.batchId)}`,
    );
    const snapshot = decodeStatus(raw);

    for (const item of snapshot.items) {
      record.items.set(item.catalogueKey, item);
    }
    if (snapshot.terminal) return;

    await sleep(Math.min(1_000 * 2 ** attempt, 30_000));
  }
}

export function startServer(adapters: PayloadAdapters): void {
  app.post("/imports", async (req, res, next) => {
    try {
      const importId = randomUUID();
      const raw = await request("POST", "/v1/image/batch/submit", {
        headers: { "Idempotency-Key": `catalogue-import:${importId}` },
        body: JSON.stringify(adapters.buildSubmitBody(req.body)),
      });
      const record: ImportRecord = {
        id: importId,
        batchId: adapters.decodeBatchId(raw),
        items: new Map(),
      };

      imports.set(importId, record);
      void pollBatch(record, adapters.decodeStatus).catch((error: unknown) => {
        console.error("batch polling stopped", { importId, error });
      });
      res.status(202).json({ importId });
    } catch (error) {
      next(error);
    }
  });

  app.get("/imports/:id", (req, res) => {
    const record = imports.get(req.params.id);
    if (!record) return res.status(404).json({ error: "Import not found" });

    const items = [...record.items.values()];
    const completed = items.filter((item) => item.state !== "pending").length;
    const failed = items.filter((item) => item.state === "failed").length;
    return res.json({ importId: record.id, completed, failed, items });
  });

  app.listen(3000);
}
Enter fullscreen mode Exit fullscreen mode

Generate the three adapters from the live capability schema. That boundary is intentional: the verified routes establish how to submit and check a batch, but they do not establish the payload field names. Guessing those names would make a polished snippet dangerous.

The map is only a compact teaching stand-in. Production imports and item outcomes belong in durable storage before the poller starts, especially with multiple Express processes. Persist the batch ID first. It is the handle for later status checks and cancellation.

Two timing rules are visible in the code. An individual request gets at most six attempts after repeated HTTP 429 responses, while ordinary job polling continues until the decoded snapshot is terminal. The status interval begins at one second and caps at 30 seconds. Those are explicit policy values, not measured recommendations; tune them against real job duration, desired display freshness, and traffic volume.

Which tool owns which part of the pipeline?

These products overlap around media, but they do not solve the same layer. A fair comparison starts with ownership.

Option Best fit Important boundary
Infrai A backend that wants batch submission and status polling over plain REST Your application still owns durable import and per-item state
Cloudinary Teams already using a managed asset, transformation, and delivery platform Adopting its asset model is broader than adding a batch poller
imgix URL-driven image rendering and CDN-oriented delivery Rendering on request is different from tracking an asynchronous import batch
ImageKit Managed upload, transformation, and media delivery workflows Evaluate its asset and delivery model as a platform choice
Sharp In-process Node.js image transformations It provides image operations, not distributed job state
AWS Step Functions Explicit orchestration for branched, multi-step workflows More state-machine design and operational machinery are required

Choose Sharp when transformation is local and the Node.js runtime can own CPU, memory, and queue pressure. Choose Cloudinary, imgix, or ImageKit when managed asset delivery is already central to the application. Choose Step Functions when the promo-video workflow has enough branches, retries, and compensating actions to deserve an explicit state machine.

A plain batch REST surface fits when Express should remain a thin coordinator and integration bandwidth matters. It is not a good fit when the team needs local, synchronous pixel operations; Sharp is the more direct tool there. It is also a poor choice when the real requirement is a full asset-delivery platform or a visual workflow orchestrator. Those limitations matter because a small HTTP integration can still be the wrong operational boundary.

Choose the owner first.

Select based on state ownership and pixel movement, not the prettiness of one request.

How often should Express poll batch status?

Poll slowly enough that observation cannot crowd out useful work. Increase the delay after each nonterminal response, cap it, and honor Retry-After on HTTP 429. Add jitter when many workers can begin together; otherwise simultaneous catalogue imports can synchronize their retries.

This is the quality-versus-bandwidth decision. Faster status checks can make a progress bar appear more responsive, but they do not improve an image or the resulting clip. A longer interval cuts status traffic and reports completion later. Write down the acceptable staleness, then spend bandwidth to meet that target.

Instrument three signals: poll duration, poll outcome by HTTP status, and time since the last successful snapshot. Put the local import ID and upstream batch ID in structured logs, along with the attempt number. Do not use either ID as a metric label; every import would create a new time series.

Alert on stale observation, not one slow request. One delayed poll is noise. An active import with no successful snapshot beyond the chosen freshness window means the displayed progress can no longer be trusted.

What does partial failure mean for progress?

Completion and success are separate dimensions. For 80 submitted products with 79 successes and one failure, display 80 completed and 1 failed. Do not show 99% and leave the operator guessing whether to wait.

Store each outcome by a stable local catalogue key. A useful row contains the state, the remote result when successful, and an error description when failed. Upsert snapshots rather than append them, because polling can observe the same terminal item more than once.

No guessing.

The objection I hear most in designs like this is that per-item persistence feels heavy. The alternative is heavier at recovery time. Without it, a failed hero image and a failed secondary thumbnail collapse into the same batch-level error even though their effect on video quality differs. Per-item state gives the catalogue team a repair queue and gives the progress UI honest numbers.

The second objection is latency: why not poll every second forever? At ten active imports, that is ten status requests per second; at 1,000 active imports, it is 1,000. The arithmetic grows while the underlying work proceeds at the same speed. Backoff places a ceiling on that amplification, and the local read endpoint can still serve UI refreshes cheaply from stored state.

The final decision rule is compact. Persist the handle. Back off observation. Preserve every outcome. Those three choices keep an Express progress display useful when catalogue quality matters more than animation smoothness.

References

Top comments (1)

Collapse
 
docify profile image
Docify •

Hey, I'm building a small open-source CLI that analyzes a codebase and generates architecture/structure documentation. I'm looking for a few developers willing to run it against a real project and tell me where it gets things wrong.
You don't need to upload your code anywhere just run:
npx @autodocify/autodocs analyze .
Requires Node 20+.
If you try it, I'd especially like to know what it missed or misunderstood.