Short answer: for property-management thumbnail batches, submit one immutable work unit per source image, apply bounded backoff only to retryable failures, and treat user cancellation as a durable intent that workers check before every derivative. Process the usual aspect ratios at upload when listings need predictable previews; defer on-demand variants when the requested crop set is genuinely unknown.
The useful boundary is not the queue vendor. It is the contract between an upload request, a derivative, and the person who can change their mind. A leasing agent may upload 180 apartment photos, then cancel after spotting the wrong floor plan. If cancellation only changes a browser spinner, the queue will still spend CPU and publish stale assets.
The decision record: where should thumbnail work happen?
The system I would approve has four invariants: an idempotency key is stable for a source revision and crop profile; a derivative is published only after its bytes and metadata pass validation; retries have a deadline and a reason; cancellation is recorded in the same store workers consult. These are storage rules, not UI preferences.
Hard stop.
No shortcuts.
| Choice | Strength | Cost or boundary |
|---|---|---|
| Process standard ratios at upload | Fast listing pages, predictable cache warming, simple user feedback | Upload latency and burst capacity rise; a large batch can delay the request path |
| Process every ratio on demand | Work is proportional to actual reads and new ratios are easy to add | First viewer pays the wait; duplicate requests need coalescing |
| Hybrid: enqueue a small baseline, defer the rest | Keeps common previews warm while preserving flexibility | Two priority classes and two observability paths must agree on cancellation |
For a property catalog, the hybrid is usually the least surprising: generate 1:1, 4:3, and 16:9 thumbnails after the original is committed, then create an unusual crop only when a page asks for it. Keep the original immutable. A crop request points to a revision, never to a mutable filename.
I am not sure the same split fits a photo marketplace with thousands of unknown display formats; your mileage may vary. Measure cache hit rate and time-to-first-thumbnail before moving more work into the upload path.
How do submission, backoff, and cancellation interact in a thumbnail batch?
Submission should return a batch identifier quickly, with each item assigned a deterministic key such as sha256(source_revision + profile + encoder_version). The API can accept the batch while a worker later claims individual items. That makes a timeout safe to retry: the second submission collides with the same key instead of producing a second JPEG.
Backoff belongs to the item, not the entire batch. A temporary object-store timeout deserves a delayed retry; an unsupported media format does not. Use capped exponential delay with jitter, and stop at an absolute deadline so a poisoned item cannot occupy a queue forever.
Cancellation has three states: requested, observed, and finalized. The client asks for cancellation, the service persists that request, and a worker observes it before decoding, before encoding, and before publishing. A running encoder may finish one small unit; the publish step still checks the cancellation marker. That distinction prevents a late success event from resurrecting a thumbnail the user intentionally removed.
Here is the critical path in Python. It is deliberately boring, because the interesting part is the state transition, not a framework wrapper.
from dataclasses import dataclass
from random import uniform
@dataclass
class Item:
key: str
attempts: int = 0
deadline: float = 0.0
RETRYABLE = {"object_timeout", "rate_limited", "temporary_decode"}
def next_delay(attempts: int, cap: float = 300.0) -> float:
base = min(cap, 2 ** attempts)
return uniform(base * 0.5, base * 1.5)
def run_item(item: Item, store, clock):
if store.cancel_requested(item.key):
store.finalize(item.key, "cancelled")
return
try:
source = store.read_source(item.key)
if store.cancel_requested(item.key):
store.finalize(item.key, "cancelled")
return
derivative = make_crop(source)
if store.cancel_requested(item.key):
store.finalize(item.key, "cancelled")
return
store.publish_if_absent(item.key, derivative)
store.finalize(item.key, "complete")
except Exception as exc:
reason = classify(exc)
item.attempts += 1
if reason in RETRYABLE and clock.now() + next_delay(item.attempts) < item.deadline:
store.reschedule(item.key, clock.now() + next_delay(item.attempts), reason)
else:
store.finalize(item.key, "failed", reason=reason)
The sample omits image-library details on purpose. Validate dimensions, color profile, and encoded byte limits at the boundary; MDN's media format guidance is a useful reminder that a file extension is not a trustworthy content declaration.
Which failure boundaries deserve a visible contract?
A batch can partially succeed. Report item-level states rather than pretending the batch is atomic: queued, running, complete, cancelled, and failed with a stable reason. A UI can then offer “retry failed” without resubmitting completed work. In practice, the race is easy to miss: the agent clicks cancel while item 37 is between encoding and its database commit, a second worker receives the same item after a visibility timeout, and a webhook arrives out of order. Persisting a monotonic state version lets the publish transaction reject an older completion, while the UI can still show the last accepted state and explain which derivatives were already durable.
The storage layer needs conditional publish semantics. If two workers race, only one should create the derivative record; both may safely report the same completed key. Garbage collection should retain the source revision until all baseline derivatives and audit records reach a terminal state. Deleting first creates an unrecoverable hole.
Observe four timings: queue wait, decode, encode, and publish. Also count cancellation lag (request timestamp to observed timestamp), retry reasons, and duplicate-submit collisions. A single p95 for the whole batch hides the exact place users feel pain.
The catch is operational complexity. The hybrid path is not suitable when the team cannot operate two priority classes or when legal retention rules require every derivative to be regenerated synchronously. Stick with upload-time processing for a small, fixed profile set; choose on-demand generation when profile demand is sparse and a slower first view is acceptable.
I once treated a cancelled batch as a front-end concern and later found completed files arriving after the agent had removed the listing. The fix was not a faster queue. It was making cancellation durable and checking it at the publish boundary.
Top comments (0)