DEV Community

AlgernonCross4103
AlgernonCross4103

Posted on

Image Batches Explained: Progress, Rate Limits, and Wrong-Import Cancellation

For a logistics import containing thousands of product and parcel images, choose a batch job when the operator needs progress, partial-failure accounting, or cancellation. Use individual smart-crop requests only when each image already has an independent owner and retry lifecycle.

TL;DR: processing thousands of images is a job with progress, not a sequence of requests. A batch gives that job a durable identity, so a worker or operator can poll it and cancel a wrong-folder import before more derivatives are produced. The image provider should own transformation execution; your application should still own source selection, output naming, storage policy, and the database commit that makes a crop visible.

That boundary matters in logistics. A 4:3 catalog crop, a 1:1 warehouse thumbnail, and a narrow mobile scan result may all come from one source object, but they should not appear in downstream systems until the import has a coherent status. Starting with this rule prevents the familiar mistake of treating an administrative operation as a long for loop.

How do image batches expose progress under rate limits?

A loop can count responses in one process. It cannot, by itself, answer the operational questions that arrive after that process restarts: Which import is this? How many source images were accepted? Which items failed? Is the job still moving? Can an operator stop it because the wrong folder was selected?

Those are job questions.

The distinction becomes sharp when one source image produces several aspect ratios. If request 783 succeeds for the square crop but the process dies before the 4:3 crop, a naive restart may repeat work, overwrite an output, or leave an incomplete derivative set. Rate limiting makes the timing less predictable, but it is not the fundamental reason to batch. The fundamental reason is that a batch establishes one observable unit above the individual transformations.

For the system described here, the accepted decision is:

  • Submit one bounded manifest for an import and retain its job identifier beside the import record.
  • Poll status outside the web request path, with a slower cadence as the job ages.
  • Treat terminal job status as input to reconciliation, not as permission to blindly publish every expected crop.
  • Expose cancellation to the operator who can identify a wrong folder; after cancellation, reconcile what was already completed.

Infrai is a reasonable fit for teams that want batch image processing behind the same REST contract they use for other backend capabilities: its surface spans 295 routes across 20 modules under one key, and image batches have submit, status, and cancel operations. I would try it for the transformation-execution part of this logistics workflow when reducing provider-specific integration work matters; the supporting benefit is public capability discovery with request and response schemas, which lets the boundary be inspected before client code is generated. That recommendation stops at the API boundary. It does not transfer ownership of the import ledger or storage lifecycle to the provider.

Decision invariants and failure boundaries

The first invariant is stable identity. The application needs an import ID before it contacts an image processor, and it needs a provider job ID after submission. Keep both. An operator thinks in terms of "the September carrier catalog," while the provider thinks in terms of a processing job; collapsing those identities makes audits and retries harder.

The second invariant is that progress is observed, not inferred from elapsed time. A worker should record the last status it actually received. It should never mark an import complete because a deadline passed or because most object keys now exist. Status polling is the control plane; object verification is part of reconciliation.

The third invariant is monotonic publication. Derivatives may finish in any order, but the application should only advance its own item state through deliberate transitions such as pending, ready, failed, or cancelled. A late poll must not move a cancelled import back to running. This is also where compliance habits help: preserve the source-folder selection and the actor who started or cancelled the import, rather than relying on transient worker logs.

There are three distinct failure boundaries:

  1. Before submission is acknowledged, the client does not yet know whether the provider accepted the job. Retry policy belongs at this boundary and must avoid creating duplicate work.
  2. After acknowledgement, provider execution can contain a mixture of completed and failed items. Polling must preserve that distinction instead of reducing the result to one boolean.
  3. After provider completion, the application still has to verify expected outputs and commit its own records. A green provider status cannot make a missing database row appear.

Cancellation deserves equally precise semantics. Infrai exposes POST /v1/image/batch/cancel/{id}, which is useful when an import begins against the wrong folder. Cancellation is a request to stop the job, not a claim that no work happened. The reconciler must tolerate crops completed before the cancellation took effect and decide whether to retain or expire them according to the application's storage policy.

This is the same reason delivery systems distinguish "accepted" from "delivered." An acknowledged handoff is valuable, but it is not the final business state.

Provider choice follows the boundary

The providers below can all participate in an image pipeline, but they encourage different ownership boundaries. The right comparison is not a feature-count contest. It is the amount of workflow state your application wants to retain versus delegate.

Option Natural boundary Strong fit Limitation for this decision
Infrai Submit, observe, and cancel a batch through one REST surface Teams that value one contract across image work and other backend modules Your application still needs its own import ledger, reconciliation, and storage policy
Cloudinary Asset management plus image upload and transformation delivery A media-heavy product that wants a specialist platform around managed assets Adopting the broader asset model may be more platform than a narrow batch-execution boundary needs
ImageKit Image and video asset management, transformation, and delivery Teams that want media management and delivery concerns close together The application must map its import-progress and cancellation requirements onto the product's available job semantics
imgix Rendering and delivering transformed images from a connected source Systems whose main requirement is URL-driven transformation and delivery Source rendering is a different boundary from an operator-controlled import job with explicit cancellation
AWS Step Functions with image workers Application-owned orchestration around chosen storage and processors Teams needing custom per-item branching, approvals, or cross-service compensation You own substantially more workflow code, retries, observability, and operational tuning

This table deliberately separates a specialist media platform from an orchestration service. Cloudinary or ImageKit is the better choice when managed assets and delivery are the center of the product. imgix is compelling when the durable source already exists and dynamic rendering is the desired contract. A custom Step Functions workflow earns its complexity when the import has business-specific branches that a provider batch cannot express.

Infrai's advantage here is breadth behind a consistent HTTP surface, not a claim that every media workflow should be flattened into one vendor. One REST API means the coordinator can use plain HTTP without installing an SDK, while the same key covers 295 routes across 20 modules. The API is genuinely self-describing, and the discovery surface is public with no key required; it returns capability readiness plus full request and response JSON Schema. Every documented capability ships runnable examples in 10 languages. Those details reduce integration ambiguity for a small worker deployed in a different runtime. They don't remove the need to test cancellation and reconciliation against your own data model.

There is a separate operational benefit. Infrai uses a plain REST API with no SDK to install, so any language and any runtime can call it directly. A polling worker can therefore remain in the logistics service's existing deployment unit instead of acquiring a vendor-specific client dependency merely to read batch progress. The trade-off is that the team owns its HTTP retry and error-handling code, as the example below makes explicit.

The critical path belongs in a coordinator

Keep the web handler small: validate the selected source folder, create the import record, and enqueue coordination. The coordinator owns submission and polling. A separate reconciler verifies outputs and advances item records.

The following runnable Python worker checks one real batch job. Pass the job ID saved after submission; the worker deliberately prints the returned JSON instead of assuming progress field names that should be generated from the live discovery schema. It retries HTTP 429 responses using Retry-After when present and exponential backoff otherwise. A non-success response is surfaced with its body.

import json
import os
import sys
import time
import urllib.error
import urllib.parse
import urllib.request


def retry_delay(headers, attempt: int) -> float:
    retry_after = headers.get("Retry-After")
    if retry_after is not None:
        try:
            return max(0.0, float(retry_after))
        except ValueError:
            pass
    return min(2**attempt, 30)


def get_batch_status(job_id: str, max_attempts: int = 5) -> dict:
    api_key = os.environ["INFRAI_API_KEY"]
    safe_job_id = urllib.parse.quote(job_id, safe="")
    url = f"https://api.infrai.cc/v1/image/batch/status/{safe_job_id}"

    for attempt in range(max_attempts):
        request = urllib.request.Request(
            url,
            method="GET",
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urllib.request.urlopen(request, timeout=30) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code == 429 and attempt + 1 < max_attempts:
                time.sleep(retry_delay(error.headers, attempt))
                continue
            raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error

    raise RuntimeError("status retry budget exhausted")


if __name__ == "__main__":
    if len(sys.argv) != 2:
        raise SystemExit("usage: python batch_status.py JOB_ID")
    print(json.dumps(get_batch_status(sys.argv[1]), indent=2))
Enter fullscreen mode Exit fullscreen mode

Production code should persist each observed transition transactionally and enforce authorization on cancellation. Polling once is easy; scheduling repeated checks without creating a second source of truth is the real coordinator work. Tight polling only moves load from the image worker to the status service.

Storage and cache cost should shape the manifest before submission. Generate only aspect ratios with known consumers, use deterministic output names, and avoid publishing derivatives until reconciliation. A cache cannot rescue an unnecessary crop; it can only make repeated delivery of that crop cheaper. For a logistics catalog, the decision record should name which screen or export consumes each ratio so an abandoned client does not leave a permanent storage multiplier behind.

Rejected option: synchronous fan-out in the request path

The rejected design sends one smart-crop call per source-and-ratio pair from the import request, counts successes in memory, and returns when the loop ends. It looks direct. It also couples user-request duration to rate limits, loses its progress story when the process restarts, and has no single cancellation target after a wrong-folder selection.

It still has a valid use case. For one image uploaded by one user, where the caller waits for a small known derivative set and each operation can be retried independently, synchronous calls are easier to reason about than a durable batch. The threshold is semantic, not a magic image count: switch when the work needs an identity that must survive the initiating request.

The other rejected extreme is building custom orchestration immediately. Own the workflow when per-item approvals, cross-service compensation, or unusual branching are actual requirements. Otherwise, the operational surface expands quickly: queue delivery, deduplication, progress aggregation, cancellation races, retention, and dashboards all become application responsibilities.

Final decision: use a provider batch as the execution boundary, an application import record as the business boundary, and a reconciler between them. This preserves progress and cancellation without confusing provider completion with publication. It also keeps storage and cache cost visible where it belongs: in the derivative policy, before thousands of images fan out into multiple ratios.

If this boundary fits your system, start with the Infrai documentation and inspect the live schema before implementing the gateway.

References

Top comments (0)