DEV Community

NoahHayes7250
NoahHayes7250

Posted on

Event Galleries — Node.js Queues for Batch Processing, Status, and Cancellation

Short answer: for an event gallery, use a durable Node.js queue when you need predictable batch throughput; use serverless workers when bursts are rare and cancellation is mostly a user-facing state change. In both cases, make status a stored state machine, and treat cancellation as a cooperative request, not a promise that already-running image work vanishes.

Ship the contract first.

Keep it boring.

Measure twice.

That decision is about quality versus bandwidth, not about picking a fashionable runtime. A marketplace gallery can receive 8,000 original photos after a weekend event. Serving every original makes the first scroll slow and the CDN expensive. Compressing too aggressively produces blocky faces and angry photographers. My revenue-per-hour test is simple: can the pipeline ship weekly, stay observable, and outsource the undifferentiated image plumbing?

The decision matrix for a busy gallery

Option Best fit Strength Trade-off
Durable Node.js queue + workers Steady batches, strict ordering, operator control Backpressure, retries, and explicit cancellation checkpoints You own worker capacity and deployment
Serverless image workers Spiky traffic, small operations team, stateless transforms Automatic burst capacity and little host management Runtime limits, cold starts, and less control over long jobs

Choose the queue when a batch must finish within a known window or when operators need to pause intake. Choose serverless when a gallery is quiet most of the week and a short burst is worth paying for in exchange for less infrastructure. Neither option fixes a vague contract.

The contract should name the output before code exists: JPEG or WebP quality range, maximum pixel dimensions, an allowed byte budget, and the fallback format for clients that cannot decode a newer codec. The browser's Accept header and the image element's srcset let delivery vary without creating duplicate originals. MDN's format guide is a useful reference for codec and container support, but it is not a performance benchmark.

How should batch processing, status tracking, and cancellation work together?

Model one gallery import as a job with child items. The parent records queued, running, cancelling, cancelled, completed, or failed; each image records its own state and attempt count. A parent is completed only when every child is terminal. This prevents the familiar lie where a progress bar reaches 100% while five derivatives are still missing.

Keep transitions server-side and idempotent. A repeated POST /jobs/{id}/cancel should return the same cancellation intent, not create another signal. Workers check that intent before downloading, after decoding, before writing a derivative, and before acknowledging the queue message. A check after the write matters: otherwise a cancelled job can still publish a thumbnail that a later request assumes is current.

Here is a small TypeScript shape for that boundary. It uses generic interfaces so the queue can be swapped without changing gallery code.

type JobState =
  | 'queued'
  | 'running'
  | 'cancelling'
  | 'cancelled'
  | 'completed'
  | 'failed';

type ImageItem = {
  id: string;
  sourceKey: string;
  state: 'queued' | 'running' | 'completed' | 'cancelled' | 'failed';
  attempts: number;
};

interface JobStore {
  get(id: string): Promise<{ state: JobState; cancelRequested: boolean }>;
  requestCancel(id: string): Promise<void>;
  markItem(id: string, state: ImageItem['state']): Promise<void>;
}

async function processItem(
  jobId: string,
  item: ImageItem,
  store: JobStore,
  transform: (key: string) => Promise<string>,
): Promise<void> {
  const before = await store.get(jobId);
  if (before.cancelRequested || before.state === 'cancelling') {
    await store.markItem(item.id, 'cancelled');
    return;
  }

  await store.markItem(item.id, 'running');
  const derivativeKey = await transform(item.sourceKey);

  const after = await store.get(jobId);
  if (after.cancelRequested || after.state === 'cancelling') {
    // Do not publish work after cancellation; clean-up can be asynchronous.
    await store.markItem(item.id, 'cancelled');
    return;
  }

  await publishDerivative(derivativeKey);
  await store.markItem(item.id, 'completed');
}

async function publishDerivative(key: string): Promise<void> {
  // Write to immutable storage, then update the gallery index in one idempotent step.
}
Enter fullscreen mode Exit fullscreen mode

The empty publishDerivative body is intentional pseudocode, not a claim about a particular storage API. The production version needs a unique (jobId, itemId, variant) key and an upsert. If a worker dies after storage but before the database update, a retry should discover the existing object and finish the index update instead of producing a second charge or a second visible thumbnail.

Where image quality and bandwidth actually collide

A single quality number is a poor policy. A 4,000-pixel portrait and a dark indoor group shot fail differently at the same setting. Use a small matrix: constrain the long edge, encode a modern format when the client advertises support, keep a JPEG fallback, and sample results at phone-sized display dimensions. Measure bytes per delivered view, decode time, and a human quality score on faces and text.

The useful threshold is product-specific. Start with a 1,600-pixel long edge for listing cards, then compare a 2,400-pixel detail variant. Store the original separately so you can re-encode when the threshold changes. Do not overwrite it with a lossy derivative.

A 12 MB source becoming a 280 KB card image is a bandwidth win only if the card remains legible. Your mileage may vary by camera, lighting, and the proportion of text-heavy images; I am not sure a universal quality target exists, and a ten-image sample from each event is a better calibration than a global promise.

Operating the pipeline after launch

Status tracking needs more than a percentage. Expose processed, total, bytesRead, bytesWritten, startedAt, updatedAt, and a short failure reason. Keep the raw error in logs with a correlation ID. A client can then distinguish “waiting for capacity” from “three files rejected,” while an operator can retry only the failed children.

Use a monotonic progress rule: processed never decreases, even if a retry is scheduled. Emit heartbeats for long transforms so a 90-second decode is not mistaken for a dead worker. Set a timeout around network reads and image decoding; cancellation should be an AbortSignal where the library supports it, with a database flag as the durable fallback.

Test the ugly paths deliberately. Kill a worker between the object write and the index update. Send the same message twice. Cancel at each checkpoint. Feed a truncated file and an image with an extreme pixel count. The expected result is a terminal child state and a retry decision, not a stuck spinner.

A weekly shipping cadence favors a boring rollout: shadow the new encoder on a small event, compare byte and quality metrics, then raise the percentage. Keep a feature flag for codec selection. Outsource the undifferentiated parts, but keep ownership of the state machine and the acceptance tests; those are where gallery trust is won.

When is the runner-up the better choice?

A queue is the wrong fit when your platform cannot keep workers patched, when jobs routinely exceed its memory envelope, or when the gallery has only a few bursts per year and idle capacity would sit unused. In those cases, serverless workers can be the pragmatic boundary, provided each invocation handles a bounded item and the parent job remains in durable storage.

Serverless is not suitable when you need a long, coordinated batch with operator pause controls, deterministic worker placement, or access to specialized native codecs that the runtime does not package. Stick with a durable queue when those controls are part of the event team's contract.

The compromise is common: enqueue one item per image, run a short transform, and let either platform consume the same job schema. That keeps cancellation semantics and quality tests portable while the execution layer changes with traffic.

The decision is therefore operational. Pick the boundary that lets a one-person team explain every state, replay every failure, and keep image quality inside the promise made to photographers.

Sources

Top comments (0)