Short answer: Use a server-owned state machine, bounded polling, and acknowledged cancellation so a media batch console never treats “request accepted” as “safe to serve.”
State first.
For an edtech image-compression console, moderation coverage should decide what the UI shows, while a small, explicit state machine decides when it stops polling. Mixing those concerns is how a cancelled resize gets displayed as successful, or how a duplicate delivery triggers a second moderation decision. The practical design is a server-owned terminal state, a client poller with a deadline, and cancellation that is recorded as an intent rather than treated as an instant deletion.
I have been paged for missed jobs and duplicate deliveries. The page rarely started with a dramatic outage. It started with a button that looked finished.
The incident lesson: “done” is a contract
Imagine a teacher uploads 240 lesson thumbnails. The console submits a batch, workers fetch the originals, compress them, and run a moderation check before the CDN URL is released. A browser tab closes after 18 seconds. When it opens again, it asks for the batch status.
The dangerous shortcut is to infer status from the last response or from a local timer. A 202 response means the request was accepted; it does not mean every item is safe to serve. The browser needs a durable batch record with an enum such as queued, running, cancelling, succeeded, failed, or cancelled. Only the last three are terminal. The worker, not the browser, owns that transition.
| State | UI meaning | Poll? |
|---|---|---|
queued / running
|
Work is pending or active | Yes |
cancelling |
Stop requested; acknowledgement pending | Yes |
succeeded / failed / cancelled
|
Final result is recorded | No |
This distinction paid for itself in review. A retry can safely ask for the same batch ID, while a refresh can rebuild the view from durable state. A second click can be rejected or coalesced by an idempotency key, and a worker retry can check that key before publishing an output. The alternative is a long chain of compensating guesses: the browser assumes completion, a queue redelivers, moderation sees the same asset twice, and an operator has to reconstruct the timeline from access logs. Those are boring properties. Boring is what you want at 02:00.
How should a batch operations UI handle polling, cancellation, and terminal states?
Polling is a control loop, not a setInterval sprinkled beside a progress bar. Start with a short delay after submission, then use bounded backoff and add jitter so a class of open tabs does not wake the API together. Stop on a terminal state, on an explicit deadline, or when the user navigates away and no background refresh is required. On a timeout, show “status unknown” and keep the batch recoverable; do not invent failure.
Cancellation needs two timestamps: when the user requested it and when the service acknowledged it. The UI can become less optimistic immediately by moving to cancelling, but it must continue polling until the server reports cancelled or another terminal result. Work already handed to an encoder may finish; the invariant is that its output is not published unless the moderation and commit rules allow it.
Here is the shape I use in Go for the decision point. It is deliberately independent of a queue vendor.
package batch
type State string
const (
Queued State = "queued"
Running State = "running"
Cancelling State = "cancelling"
Succeeded State = "succeeded"
Failed State = "failed"
Cancelled State = "cancelled"
)
func Terminal(s State) bool {
return s == Succeeded || s == Failed || s == Cancelled
}
// ApplyCancel records intent; the worker later commits the terminal result.
func ApplyCancel(s State) State {
if Terminal(s) || s == Cancelling {
return s
}
return Cancelling
}
The storage update must be conditional. For example, a worker completing at the same moment as a cancel request should win or lose according to one documented rule, enforced by an atomic compare-and-swap or transaction. “Last HTTP request wins” is not a rule; network timing is not business policy.
Do not let the spinner become your source of truth.
Compression and moderation are separate gates
The media pipeline has two different questions: did the bytes get transformed, and may those bytes be served? A successful JPEG or WebP encode cannot stand in for moderation coverage. Keep those facts in separate item fields, then derive the batch state from item outcomes. A batch with 237 approved images and 3 moderation rejections is not the same as a batch with three encoder errors.
For each item, retain the original media type, output type, dimensions, byte count, moderation decision, and a stable operation ID. The UI can then explain “3 blocked by policy” without exposing an internal retry as a second asset. MDN's media format guidance is useful here because browser support and codec choice are separate from workflow state; format support should be tested at the delivery boundary, not guessed from the worker's success flag.
The catch is that a single all-or-nothing terminal state is unsuitable when teachers need partial progress and policy review. Use an aggregate such as succeeded_with_rejections only if product and compliance agree on its meaning. Otherwise, keep the batch terminal as succeeded with item-level dispositions, and make the publish step require every required gate. Stick with a simpler all-or-nothing model when downstream consumers cannot handle partial results.
A poller that survives refreshes and slow workers
The poller should consume a snapshot, not mutate a guessed counter. A compact response might include state, updated_at, total, completed, and a list of item errors. Counters are hints for display; the state enum is the contract. If updated_at stops advancing, emit a stale-status signal and widen the interval rather than hammering the service.
In practice I set a maximum poll window per foreground view and persist the batch ID in the route. I am not sure every product needs background notifications, but the recovery path is clear: a later visit can fetch the same ID, and an operator can reconcile it from the event log. That is safer than making the browser the only witness.
Instrument transitions, not just request latency. Useful dimensions include batch size, terminal state, cancellation latency, moderation rejection count, and the age of the oldest non-terminal batch. Alert on age and missing transitions. A spike in 409 responses from duplicate cancel requests is usually a client contract problem, not a worker capacity problem.
What to test before shipping the console
Test the transition table with property-style cases: cancellation before work starts, during encoding, after moderation, and after success. Repeat every command with the same idempotency key. Refresh between every response. Advance a fake clock through backoff, deadline, and jitter. Then run browser tests against a deliberately slow worker so the UI spends time in running and cancelling, not only in the happy path.
The operational runbook should answer three questions quickly: which state is authoritative, which event is missing, and whether serving is gated on moderation. Include the batch ID in logs and support links. Never ask an operator to infer ownership from a screenshot.
This design is not suitable when the job is guaranteed to finish synchronously in a single request and there is no cancellation or policy gate; a normal request/response flow is simpler there. For large media sets, flaky networks, or human moderation, the explicit state machine earns its complexity.
Top comments (0)