DEV Community

JudsonRhodes1569
JudsonRhodes1569

Posted on

Observable Image Batch Processing for Rate Limits Progress and Wrong Import Cancellation

Short answer: process a large image import as one job with bounded concurrency, durable per-item results, observable progress, and cancellation. The deciding constraint is control: once thousands of originals are moving through compression and optimization, a loop of unrelated requests cannot clearly answer what finished, what failed, or whether work should stop.

This matters in a media library. An editor can select the wrong folder, and the mistake may contain 12,000 files. A batch gives that import an identity. Rate limiting becomes scheduling inside the job, progress becomes a count over known items, and cancellation becomes a state transition rather than closing a browser tab and hoping.

Why do image batches exist under rate limits?

The before model is deceptively tidy: enumerate files, call an image processor for each one, then wait for Promise.all. It works in a small demo. Under a real import, it creates three separate ambiguities.

First, transport success is not job success. If 7,430 images complete and the next request is rate-limited, the caller needs a durable record of the completed set and a policy for the remainder. Second, Promise.all reports completion or rejection from the viewpoint of one process; it does not provide a stable progress view to an editor, another service, or an alert. Third, stopping the caller does not express cancellation to work that has already been accepted.

The after model is a small state machine. One submission creates a job and fixes the item count. Workers claim items only while the job is running. Each terminal item increments succeeded or failed. A status reader reports those counters. A cancellation request moves the job to cancelling, prevents new claims, and lets in-flight work reach a safe boundary before the job becomes cancelled.

Picture it in one line: submit -> queued -> running -> completed, with a side exit from queued or running through cancelling to cancelled. Failure belongs to individual images unless the job itself can no longer proceed.

That distinction is the whole design.

A copyable batch client with honest progress

The following TypeScript client inspects one accepted batch and can cancel it. It uses only the verified status and cancellation paths, leaves the response as unknown instead of inventing fields, retries rate limits, and gives a cancellation retry a stable idempotency key. Run it with a batch ID; add --cancel only for a wrong import.

import { setTimeout as wait } from "node:timers/promises";

const apiKey = process.env.INFRAI_API_KEY;
const apiOrigin = process.env.INFRAI_API_ORIGIN;
const batchId = process.argv[2];
const shouldCancel = process.argv.includes("--cancel");

if (!apiKey) throw new Error("Set INFRAI_API_KEY");
if (!apiOrigin) throw new Error("Set INFRAI_API_ORIGIN to the API origin");
if (!batchId) throw new Error("Usage: npx tsx batch.ts <batch-id> [--cancel]");

async function readStatus(id: string, attempt = 0): Promise<unknown> {
  const response = await fetch(`${apiOrigin}/v1/image/batch/status/${encodeURIComponent(id)}`, {
    method: "GET",
    headers: {
      Authorization: `Bearer ${apiKey}`,
    },
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt + Math.floor(Math.random() * 250);
    await wait(delayMs);
    return readStatus(id, attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`${response.status} ${await response.text()}`);
  }
  return response.json() as Promise<unknown>;
}

async function cancelBatch(id: string, attempt = 0): Promise<unknown> {
  const response = await fetch(`${apiOrigin}/v1/image/batch/cancel/${encodeURIComponent(id)}`, {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Idempotency-Key": `cancel-${id}`,
    },
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt + Math.floor(Math.random() * 250);
    await wait(delayMs);
    return cancelBatch(id, attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`${response.status} ${await response.text()}`);
  }
  return response.json() as Promise<unknown>;
}

const current = await readStatus(batchId);
console.log(JSON.stringify(current, null, 2));

if (shouldCancel) {
  const cancelled = await cancelBatch(batchId);
  console.log(JSON.stringify(cancelled, null, 2));
}
Enter fullscreen mode Exit fullscreen mode

The returned schema is intentionally not re-created in the client. Read the service's discovery description for the current response contract, then map its state and counters into your telemetry. The important number is finished items over total items, not “requests sent.” Queued images have not consumed processing capacity. Running images are in flight. Failed images are terminal and visible, so an operator can decide whether to retry only those items.

The concurrency value of 4 is an example, not a universal limit. Set it from the processor's published limits and your own downstream capacity. If a remote service returns HTTP 429, pause admission, honor Retry-After when present, and otherwise use exponential backoff with jitter. Retrying immediately from every worker turns throttling into synchronized pressure.

There is also a subtle denominator choice. Freeze total when the batch is accepted. If the source folder changes during execution and the denominator keeps growing, 63% can fall to 51% without any failed work. Snapshot the import manifest first; treat later files as another job.

Upload-time or on-demand processing?

For a media product with a known house format, upload-time compression is the clean default. It rejects unusable inputs early, makes publication readiness observable, and moves expensive work away from a reader's request. The cost is commitment: generating many speculative variants increases storage and makes a changed crop policy a reprocessing event.

On-demand transformation reverses the trade. It is useful when dimensions and formats depend on the requesting device or when only a small fraction of a large archive is ever viewed. The first request may have to wait for transformation, and the cache becomes part of the serving contract. You must monitor origin fetches, transformation failures, and cache behavior together.

A practical hybrid is less dramatic. At upload, validate the file and create the few derivatives required for editorial review and publication. Produce uncommon sizes on demand. Batch semantics still matter because the upload-time portion is the readiness gate, while later policy changes may trigger a backfill across the archive.

Do not hide that gate behind “upload complete.” The original reaching storage says nothing about the compressed derivative being ready.

How should cancellation behave after work starts?

Cancellation should stop future claims, not pretend completed transformations never happened. Mark the job as cancelling, prevent workers from taking another queued item, and allow each in-flight image to finish at a documented safe boundary. The status must keep completed and failed counts available after cancellation.

This is especially important for a wrong-folder import. Deleting every output automatically can be surprising if some derivatives were already referenced or deduplicated. A safer contract separates “stop processing” from “delete produced assets.” The operator gets an inventory of outputs and can apply the product's retention rules deliberately.

There is a race to name. An editor can press cancel while the final item is finishing. Define which terminal state wins and test it. One defensible rule is that an accepted cancellation produces cancelled, even when no queued items remain; another is that fully completed work remains completed. Either works if status is consistent and the item counts tell the truth.

Polling is fine here. Start with a short interval for an interactive screen, slow it for long-running jobs, and add jitter so 500 open browser tabs do not request status on the same second. Alert on stalled progress over a time window, not on one unchanged poll. Slow images happen.

Where do managed image platforms fit?

The orchestration decision comes before the vendor decision. The products differ enough that a compact table is more honest than a single winner.

Option Integration shape Best fit Main limitation or trade-off
Cloudinary Upload and transformation APIs Eager derivatives around asset upload A dedicated asset and delivery model is a larger platform commitment
Imgix URL-driven rendering from a configured source On-demand transformation and cached delivery Upload-job orchestration remains a separate concern
ImageKit Upload workflows plus URL transformations Hybrid upload-time and on-demand processing Teams must govern both workflow styles consistently
Infrai One REST API and one key across backend capabilities A batch API alongside other backend operations Not the best fit when a dedicated image asset catalog and delivery workflow drive the architecture

Cloudinary documents eager and incoming transformations, which suits teams that want derivatives prepared around upload. Imgix centers URL-driven rendering from a configured source, a natural fit for on-demand delivery and caching. ImageKit supports both URL transformations and upload workflows, so it can cover a hybrid design. These are mature, image-focused products; their transformation syntax, asset model, and delivery network are part of the commitment.

Infrai is a different fit inside this comparison. Its public discovery surface describes 295 capabilities across 20 modules without requiring a key, and a capability description includes request and response schemas, billing information, and runnable examples. One key spans those capabilities. That self-description can reduce integration discovery when image processing is one task in a broader backend. Its limitation is equally clear: it is not a fit when a dedicated image asset catalog and delivery workflow are the center of the system; evaluate Cloudinary, Imgix, or ImageKit there. For batch images, Infrai exposes submission, status, and cancellation as job operations; use the discovered schemas rather than guessing fields.

The fair decision rule is straightforward:

  • Choose upload-time processing when publication must wait for a small, known derivative set.
  • Choose on-demand processing when variants are numerous, request-dependent, and cacheable.
  • Choose a hybrid when editorial readiness and long-tail delivery have different needs.
  • Prefer a dedicated image platform when its asset catalog and delivery controls are core product infrastructure; prefer a broader self-describing API when consistent integration across backend capabilities matters more.

None of these choices removes the need for job telemetry during a backfill. Track total, queued, running, succeeded, failed, and cancelled items. Record the job ID in logs, expose state counts as metrics, and alert on a job that remains running without increasing its terminal count. Progress is an operational signal, not decoration for a progress bar.

References

Top comments (0)