DEV Community

xanderblack5716
xanderblack5716

Posted on

Bulk Process Product Catalogue Images with a Node.js Batch API: Progress and Retention

Short answer: for a property-management catalogue with thousands of product photos, submit an idempotent batch, persist one small progress record per image, and retain the original only until the derived image has passed review. The expensive decision is usually retention and cache policy, not the HTTP call that starts the job.

Start with the bytes on the bill

Background removal creates at least two objects: the uploaded original and a transparent derivative. A thumbnail cache, retry copy, and audit metadata can multiply that footprint. If a catalogue has 12,000 images averaging 4 MB, the originals alone are about 48 GB before derivatives and replicas. That is the number I put beside the retention proposal.

The useful unit is not “one API request.” It is an image version over a number of days. A simple estimate is:

stored_gigabytes = image_count * average_bytes * retained_versions / 1,000,000,000

Keep the original for a defined review window, then move it to a colder tier or delete it according to the property team’s policy. Keep the transparent output longer because it is what the listing system serves. A cache should have its own expiry; it is a performance copy, not a second archive.

That distinction changes the design. A failed thumbnail regeneration should not force you to keep every intermediate file forever. It should create a replayable manifest entry with a reason, checksum, and last attempt timestamp.

One short sentence helps here: bytes are inventory.

How should a Node.js API batch track progress for thousands of images?

Treat submission, work, and observation as separate operations. The submitter sends stable image identifiers and a requested transformation. A queue worker claims a bounded slice. A progress reader reports counts from durable state, rather than inferring progress from HTTP connections that may have already closed.

The API contract can remain ordinary HTTP. The paths below are deliberately vendor-neutral placeholders; map them to the interface your provider documents.

curl -X POST https://api.example.test/batches \
  -H 'content-type: application/json' \
  -d '{"catalogue":"spring-2026","items":[{"id":"unit-0042-front","source":"s3://photos/unit-0042-front.jpg","operation":"remove-background"}],"idempotency_key":"spring-2026-v3"}'

curl https://api.example.test/batches/batch_7f2/items?limit=100
Enter fullscreen mode Exit fullscreen mode

In Node.js, the producer should stream a manifest instead of materializing thousands of image payloads in memory. Store batch_id, item_id, source checksum, operation version, state, attempt count, and output key. A unique constraint on (batch_id, item_id, operation_version) makes a retry safe. If the same item is submitted twice, the worker can acknowledge the existing row and avoid a second derivative.

Progress is a snapshot, not a promise. Report queued, running, succeeded, failed, and cancelled counts, plus updated_at. The denominator should be the number of accepted manifest rows, not the number the caller hoped to send. This prevents a dashboard from claiming 100% while a parser silently skipped malformed records.

I once assumed a single percentage would be enough. It was not. An operator needed to know that 9,870 files were complete, 80 were waiting on a retry, and 50 had been rejected for an unsupported format. Those states imply different actions and different storage decisions.

Queue boundaries, retries, and cache keys

A batch endpoint should enqueue references, not multi-megabyte binaries. Workers fetch the source, validate its media type and dimensions, run the transformation, write the derivative, and atomically mark the item complete. A lease or visibility timeout prevents a crashed worker from holding work indefinitely.

The failure mode worth simulating is a worker dying after the derivative upload but before the progress row is committed. On restart, an at-least-once queue delivers the item again. The worker checks the content-addressed output key and the operation version, verifies that the object is complete, and then records success without creating a second file. If the object is partial or the checksum differs, it writes a new temporary key and replaces the pointer only after the upload has been verified. That sequence matters for a catalogue import because a deployment can interrupt hundreds of workers at once; without it, the retry wave creates duplicate storage, misleading progress, and a cache full of mutually inconsistent versions. The database row and object pointer are separate systems, so the design must make either order recoverable rather than pretending the write is one transaction.

Retry only failures that are plausibly transient. Use exponential backoff with a cap and a maximum attempt count. A decode error, an absent object, or a policy rejection is deterministic; retrying it increases traffic and telemetry without changing the outcome. Record a machine-readable failure class so an operator can fix the manifest rather than press “retry all.”

Cache identity must include the source checksum and operation version. A key such as catalogue/unit-0042/front.png is unsafe when a photographer uploads a replacement under the same filename. Prefer a content-addressed component, for example derivatives/<sha256>/<operation-v3>.png, and put a short-lived listing alias in front of it.

The catch is that aggressive deduplication can hide a legitimate reprocess. When the segmentation model or background policy changes, increment the operation version. Keep the old derivative only for the review window, then apply the same retention rule as any other superseded version.

What should be kept when a property photo job fails?

Keep enough evidence to replay one item, not enough to reconstruct every byte forever. The durable record needs the source location, checksum, transformation version, timestamps, and a bounded error detail. A sampled request log can carry latency and status, while a separate counter tracks totals for the whole batch.

Telemetry is storage too. High-cardinality labels such as full object keys, tenant names, and arbitrary error strings make metrics expensive and difficult to query. Use a low-cardinality operation, format, result, and failure_class; put the item identifier in a trace or short-lived structured log with an expiry. I am not sure any team gets this balance right on the first pass, so review cardinality after the first real catalogue and adjust the schema.

A practical policy might retain item metadata for 30 days, originals for the team’s documented review period, and cache entries until they have been cold for a defined number of days. Those durations are policy choices, not universal constants. If legal or insurer requirements demand longer evidence, that requirement wins and the estimate must be recalculated.

A decision rule for the implementation

Choose the smallest architecture that can answer three questions at any moment: which images are done, which need attention, and how many bytes will remain next month? For a small team, one relational progress table, an object store with lifecycle rules, and a worker pool are often enough. Add a separate workflow engine when you need human approvals, fan-out across several transformations, or long-running compensation steps.

Do not use a synchronous request for a catalogue-sized job. It couples client timeouts to image duration and gives operators no durable checkpoint. Do not make the cache the source of truth. It can be evicted and rebuilt from the derivative record.

Stick with a simpler per-image queue when batches are short, formats are predictable, and the review window is measured in hours. Use a manifest plus resumable workers when imports arrive in waves, files are large, or the property team needs to pause a rollout. The trade-off is operational surface area: durable progress costs database rows and migration work, but it prevents a second full scan after every interruption.

The success metric is not maximum throughput. It is a catalogue whose state is explainable, whose failed items are replayable, and whose retained bytes have an owner.

References

Further reading

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

The upload-before-progress-commit recovery case is a useful detail. I'd test one more boundary: the source URL stays the same but its bytes change after the manifest checksum is recorded. Does the worker compare the fetched bytes with that accepted checksum before transforming, or could it record success for a different image version? The (batch_id, item_id, operation_version) constraint also needs an explicit conflict rule when a resubmission uses that tuple with a different source checksum. Rejecting that mismatch would keep idempotency from silently accepting a changed request.