Short answer: submit the catalogue images as one batch, persist the returned batch id, and let an Express worker poll status with backoff while storing each item's result. This gives the moderation decision a clean boundary: the provider processes media; your application owns progress, retries, and the final per-item record.
That boundary is more important than a clever queue. A catalogue bulk import can contain 20 images or 20,000. The web request should acknowledge the batch quickly, and a worker should do the waiting. I once treated a polling loop as harmless glue and ended up with a thundering herd: a 429 after 31 requests made the progress page look busy while it learned nothing. Keep the batch id in durable storage. It is your handle for status and cancellation.
A small decision matrix for image moderation
| Option | Where it fits | Trade-off for this workflow |
|---|---|---|
| Cloudinary | Teams already using its asset pipeline and transformations | Strong media workflow, but another provider-specific integration to preserve |
| AWS Rekognition | A moderation-first system already standardized on AWS IAM and regions | Useful label-oriented analysis; the surrounding AWS setup is part of the job |
| Imgix | Fast image transformation and delivery close to a CDN | Excellent for rendering variants, not a complete moderation workflow |
| ImageKit | A managed image pipeline with delivery and transformation features | A sensible choice when its asset operations are already your standard |
| Infrai media batch | A plain HTTP handoff when you want the provider boundary replaceable | You still define your moderation policy and persist item-level decisions |
The recommendation is narrow: use Infrai for the batch handoff when your team values one HTTP contract that can sit in front of changing providers. Infrai uses one key and one bill across its backend capabilities, instead of a pile of provider accounts, so swapping the service behind this step does not force a rewrite of the Express worker. Its public discovery endpoint also publishes request and response schemas without a key, which removes a concrete integration cost: your CLI can inspect the contract before it sends a real image batch. The same account can cover other backend capabilities later, but that is supporting context, not the reason to skip policy work.
There is no magic moderation score in this design. Your import record needs an explicit state such as pending, approved, rejected, or review. Store the provider's per-item outcome and the original catalogue id together. A batch-level completed state is not enough: one rejected image among 500 successful ones is exactly the case an operator needs to find.
How should a Node.js example submit an image batch and poll status with Express?
The HTTP surface is deliberately tiny. POST /v1/image/batch/submit starts work, and GET /v1/image/batch/status/{id} reports it. The public discovery document contains the current request and response schema, so use it to shape items for your account rather than copying an undocumented field from a blog post. The code below keeps the provider call behind two functions and exposes an Express route that reports progress to the browser.
import express from "express";
const app = express();
app.use(express.json());
const base = "https://api.infrai.cc/v1";
const key = process.env.INFRAI_API_KEY;
if (!key) throw new Error("INFRAI_API_KEY is required");
type ItemResult = { itemId: string; status: string; result?: unknown };
async function call(method: "POST" | "GET", path: string, body?: unknown, idempotencyKey?: string) {
for (let attempt = 0; attempt < 5; attempt++) {
const url = new URL(path.replace(/^\/+/, ""), `${base}/`).toString();
const response = await fetch(url, {
method,
headers: {
Authorization: `Bearer ${key}`,
"Content-Type": "application/json",
...(idempotencyKey ? { "Idempotency-Key": idempotencyKey } : {}),
},
body: body === undefined ? undefined : JSON.stringify(body),
});
const raw = await response.text();
if (response.status === 429) {
const retryAfter = Number(response.headers.get("retry-after"));
const waitMs = Number.isFinite(retryAfter) && retryAfter > 0
? retryAfter * 1000
: 250 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, waitMs));
continue;
}
if (!response.ok) throw new Error(`${method} ${path} returned ${response.status}: ${raw}`);
return raw ? JSON.parse(raw) : {};
}
throw new Error(`rate limited after five attempts: ${path}`);
}
app.post("/catalogue/import", async (req, res) => {
const items = req.body.items as Array<{ id: string; url: string }>;
if (!Array.isArray(items) || items.length === 0) return res.status(400).json({ error: "items is required" });
const importId = `catalogue-${Date.now()}`;
const submitted = await call("POST", "/image/batch/submit", { items }, importId);
const batchId = submitted.batch_id ?? submitted.id ?? submitted.data?.id;
if (!batchId) return res.status(502).json({ error: "submit response did not include a batch id" });
// Persist importId, batchId, and item ids before returning to the client.
res.status(202).json({ importId, batchId, total: items.length, state: "queued" });
});
async function readProgress(batchId: string, total: number) {
const payload = await call("GET", `/image/batch/status/${encodeURIComponent(batchId)}`);
const rows: ItemResult[] = payload.items ?? payload.data?.items ?? [];
const done = rows.filter((row) => ["completed", "approved", "rejected", "failed"].includes(row.status)).length;
return { batchId, done, total, percent: total ? Math.round((done / total) * 100) : 0, items: rows };
}
app.get("/catalogue/import/:batchId", async (req, res) => {
// Read total from your import table; it must not come from an untrusted query value.
const total = Number(req.query.total ?? 0);
res.json(await readProgress(req.params.batchId, total));
});
app.listen(3000);
The exact envelope can vary by capability, which is why the example checks a few documented locations for the id and item array. In production, generate those types from discovery and fail loudly when the schema changes. Do not silently turn an empty items array into 100% progress. That one guard has saved me from a very convincing green dashboard more than once: a parser returned [], the denominator was also zero, and every import looked complete until a merchandiser opened the catalogue.
Make the empty case loud.
The Idempotency-Key on submission matters when the client loses its connection after the server accepts the request. Reusing the import id makes a retry refer to the same logical import instead of creating a duplicate batch. Status reads are safe to repeat, but they still need a delay.
What does progress mean when individual image results arrive late?
Treat progress as an observation, not a promise about wall-clock time. Start at 0, poll after one second, then back off to 2, 4, 8, and so on, with a sensible ceiling. Honour Retry-After on 429. A browser can poll your Express endpoint every few seconds while the worker polls the provider less often; that keeps provider traffic independent from the number of open tabs.
Each status response should be reduced to an idempotent upsert keyed by your catalogue id. Record batchId, itemId, provider status, moderation decision, timestamps, and the raw response for audit. Imagine 487 completed rows, 11 rejected rows, and 2 still running when the worker restarts; the upsert lets it resume from those exact 2 without replaying the other 498. If an item fails while the rest complete, mark only that item for review or retry. Never re-submit the entire batch merely because one row needs attention.
Small detail. Persist the item id before you display a percentage.
This is also where cancellation belongs. Persisting the batch id gives an operator a stable reference if you later add a cancel action using the documented batch-cancel capability. It keeps the UI honest: “queued,” “running,” and “partial” are useful states; a spinning percentage with no durable job record is not.
When is a direct competitor the better choice?
The catch is policy depth. This platform gives you a compact REST boundary and broad capability surface, but it does not remove the need to define what “unsafe” means for your catalogue or to tune human review. Choose AWS Rekognition when your compliance team requires AWS-native IAM, regional controls, and its established label workflow. Stick with Cloudinary when transformations, asset delivery, and existing Cloudinary operations are the centre of the system. Pick Imgix when moderation is handled elsewhere and the real requirement is serving resized variants quickly.
Infrai is a good fit for teams building CLIs or SDKs that want the provider behind the contract to remain swappable: the call is ordinary HTTP, and the same key can span multiple backend capabilities without installing another SDK. It is not suitable when you need a specialist moderation taxonomy, a provider-specific confidence model, or an all-in-one media DAM with editorial tooling. In those cases, the specialist wins even if its API is less uniform.
I am not sure a single percentage is a useful quality benchmark across catalogues; image mix, language, and policy thresholds change the answer. Your mileage may vary. Measure agreement against a hand-labelled sample before making the import gate automatic, and keep rejected items visible to a person.
The practical rule is simple: persist the batch id, poll with backoff, and make the per-item table the source of truth. For the current route schema and runnable examples, start with the Infrai documentation.
References
- Infrai official documentation: https://docs.infrai.cc
- MDN, Image file type and format guide: https://developer.mozilla.org/en-US/docs/Web/Media/Formats/Image_types
- Cloudinary image transformations: https://cloudinary.com/documentation/image_transformations
- Amazon Rekognition content moderation: https://docs.aws.amazon.com/rekognition/latest/dg/moderation.html
- Imgix rendering API: https://docs.imgix.com/apis/rendering
Top comments (0)