Short answer: give every promo-video job its own cancellation token, check it between expensive stages, and record each state transition in an append-only audit stream. Cancel cooperatively rather than assuming that closing an HTTP request stops work. For a marketplace pipeline, compress source images against a measured quality budget before serving them to the renderer; cancel when the job exceeds its explicit byte, attempt, or elapsed-time budget.
That rule keeps three concerns separate: the API accepts intent, a worker owns execution, and the audit trail explains what happened. It also makes cancellation testable in a notebook before the same state machine moves into production.
No polling trick can rescue a vague state model.
How should a running generation job actually stop?
A cancellation endpoint should record intent and return quickly. The worker observes that intent at checkpoints, stops scheduling new work, releases temporary resources, and records a terminal state. The useful distinction is between cancel_requested and cancelled: the first means the request was accepted; the second means execution reached a safe stopping point.
Do not infer cancellation from a disconnected browser or an expired request. A client can retry, a proxy can close an idle connection, and a worker may still be processing a frame after the API process has returned. A durable job record is the shared contract.
Use a small state machine. queued can move to running or cancelled; running can move to cancel_requested, succeeded, or failed; cancel_requested can move to cancelled, succeeded, or failed. The success branch matters because a final encode may finish just before the worker observes the token. Preserve that race in the audit record instead of rewriting history.
This design has a limitation: cooperative cancellation is not instantaneous. A long, indivisible encoder call can delay acknowledgment until control returns to the worker. If hard termination time is more important than graceful cleanup, isolate that stage in a process or disposable worker with a strict external deadline, accepting extra orchestration and the possibility of abandoned temporary output. If output integrity matters more, keep the cooperative boundary and expose the delay honestly in status and metrics.
The data flow is plain: an API creates a job with budgets and an idempotency key, a queue hands its identifier to a worker, and the worker fetches marketplace images, normalizes and compresses them, renders the promo sequence, and checks cancellation around every expensive boundary. Status reads come from the job store. Audit events come from a separate append-only table or stream keyed by the same job identifier.
A minimal cooperative implementation
This Python example keeps storage in memory so the transitions stay visible. The interface maps to a web framework, durable database, and worker queue; those substitutions should not change the cancellation contract.
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime, timezone
from threading import Event, Lock, Thread
from time import monotonic, sleep
from typing import Any
from uuid import uuid4
def now_iso() -> str:
return datetime.now(timezone.utc).isoformat()
@dataclass
class Job:
job_id: str
max_seconds: float
max_source_bytes: int
state: str = "queued"
source_bytes: int = 0
cancel_token: Event = field(default_factory=Event)
events: list[dict[str, Any]] = field(default_factory=list)
class JobStore:
def __init__(self) -> None:
self._jobs: dict[str, Job] = {}
self._lock = Lock()
def create(self, max_seconds: float, max_source_bytes: int) -> Job:
job = Job(str(uuid4()), max_seconds, max_source_bytes)
with self._lock:
self._jobs[job.job_id] = job
self._record(job, "job.created", {})
return job
def transition(self, job: Job, state: str, detail: dict[str, Any]) -> None:
with self._lock:
job.state = state
self._record(job, f"job.{state}", detail)
def request_cancel(self, job_id: str, actor: str, reason: str) -> str:
with self._lock:
job = self._jobs[job_id]
if job.state in {"succeeded", "failed", "cancelled"}:
self._record(job, "cancel.noop", {"actor": actor, "reason": reason})
return job.state
job.cancel_token.set()
job.state = "cancel_requested"
self._record(
job,
"job.cancel_requested",
{"actor": actor, "reason": reason},
)
return job.state
@staticmethod
def _record(job: Job, event: str, detail: dict[str, Any]) -> None:
job.events.append({"at": now_iso(), "event": event, "detail": detail})
def run_job(store: JobStore, job: Job, image_sizes: list[int]) -> None:
started = monotonic()
store.transition(job, "running", {})
for index, source_bytes in enumerate(image_sizes):
if job.cancel_token.is_set():
store.transition(
job, "cancelled", {"checkpoint": "before_image", "index": index}
)
return
job.source_bytes += source_bytes
if job.source_bytes > job.max_source_bytes:
job.cancel_token.set()
store.transition(
job, "cancel_requested", {"reason": "source_byte_budget"}
)
continue
# Replace this pause with decode, resize, encode, and durable writes.
sleep(0.02)
if monotonic() - started > job.max_seconds:
job.cancel_token.set()
store.transition(
job, "cancel_requested", {"reason": "elapsed_time_budget"}
)
if job.cancel_token.is_set():
store.transition(job, "cancelled", {"checkpoint": "before_render"})
return
# Pass the same token into a renderer that supports cooperative cancellation.
sleep(0.02)
store.transition(job, "succeeded", {"source_bytes": job.source_bytes})
store = JobStore()
job = store.create(max_seconds=2.0, max_source_bytes=8_000_000)
worker = Thread(
target=run_job,
args=(store, job, [1_500_000, 2_200_000, 5_000_000]),
)
worker.start()
worker.join()
print(job.state)
print(job.events)
The HTTP layer needs only three operations: create a job, read its status, and request cancellation. Keep each transition inside one atomic database transaction. Require an authenticated actor and a reason for cancellation, and make repeated requests return the current state instead of enqueueing duplicate cleanup.
That simplicity is deliberate.
The example uses bytes and elapsed time because they are stable policy inputs. It does not estimate currency. Price tables change, while input volume, attempts, render duration, and output bytes remain useful audit dimensions regardless of the execution provider.
Put image quality before rendering cost
Promo generation often begins with seller-uploaded images that vary in dimensions, format, transparency, and compression. Normalize them before they enter the video renderer. Otherwise, two visually equivalent jobs can consume very different bandwidth and processing time, making a cancellation budget difficult to interpret.
MDN's image format guide describes broad browser support and characteristics for common web image formats, including JPEG, PNG, WebP, and AVIF. Format selection is a compatibility and content decision, not a universal leaderboard. A photograph, a logo with transparency, and a screenshot with sharp text do not reward the same encoding choice.
Start with an eval set drawn from the marketplace's real image classes, but do not invent one acceptable quality number. For every candidate encoding, retain the source identifier, dimensions, encoded bytes, format, encoder settings, and the result of a visual or task-specific quality check. The decision becomes explicit: choose the smallest output that passes both the quality gate and the delivery compatibility policy.
Small first.
A 200 KB candidate that fails legibility is not an optimization. A 900 KB candidate that is visually indistinguishable from a 500 KB candidate wastes transfer. The useful frontier sits between those outcomes, and an eval harness should expose it per content class. This is the same discipline used for prompt changes: pin a representative corpus, define acceptance before tuning, and promote only configurations that pass regression checks.
There is no free format switch. More aggressive compression reduces delivered bytes but can erase texture, soften listing text, or add encode time; preserving every source pixel protects detail but pushes more data through fetching, storage, and rendering. Precomputing several variants improves delivery choices at the cost of more storage and invalidation work. Encoding on demand avoids unused variants, yet it shifts latency and compute into the request path. The right trade-off follows the workload: precompute stable, frequently reused catalog assets; use bounded on-demand work for rare sizes; and retain the original only where policy requires future reprocessing.
Record the selected asset manifest before rendering. If cancellation arrives later, the audit trail can distinguish time spent fetching images, encoding assets, generating frames, and packaging the final video. That breakdown is more actionable than a lone cancelled label.
Budgets, retries, and race conditions
Cancellation controls resource use only when the worker checks it frequently enough. Put checkpoints before downloads, after each image encode, between scene or frame batches, before the final render, and before publishing output. Never place one halfway through a non-atomic file replacement or database update.
Retries need their own budget. Track attempts on the logical job and require each worker lease to carry a unique execution identifier. An expired lease may allow another worker to resume, but a late result from the old execution must not overwrite the winner. Conditional updates against the expected state and execution identifier make the rule enforceable.
Three tests reveal most design errors. First, request cancellation twice and verify there is one terminal cleanup path. Second, cancel at every checkpoint while injecting a slow stage; each run should terminate with an attributable event and no published partial output. Third, race cancellation against successful completion. Either terminal result can be valid under the declared transition rules, but the stored sequence must explain which commit won.
The race is unavoidable.
I would also evaluate policy changes offline before rollout: replay recorded job metadata through proposed byte and duration limits, count which jobs would be stopped, and inspect the affected image classes. This is an eval, not a forecast of exact spend. It prevents a tidy global threshold from disproportionately cancelling listings with detail-heavy product photography.
Make the audit trail answer real questions
Logs help operators investigate, but the audit model should be queryable without reconstructing truth from free-form messages. Each event needs a job identifier, execution identifier, timestamp, previous and next state, actor or policy source, reason code, and bounded metadata such as cumulative source bytes. Avoid storing image payloads, prompts, credentials, or unrestricted user text in the event.
The record should answer: Who requested cancellation? Which checkpoint observed it? Which budget triggered automatically? Was any output published? How many attempts ran? Those questions serve incident review, support, and cost attribution without coupling the system to a particular queue or rendering engine.
Audit storage has a cost too. Per-frame events can overwhelm the signal with volume, while one event per job hides the stage that consumed the budget. Record state changes and coarse stage boundaries by default, then sample detailed timing into observability data with a defined retention window. This separation limits audit growth without sacrificing the durable facts needed to explain a cancellation.
Operationally, alert on trends rather than treating every user cancellation as an error. Watch cancellation latency from request to terminal acknowledgment, jobs stuck in cancel_requested, repeated attempts, and orphaned temporary objects. Compare these signals by pipeline stage and input class. A rise after an encoder-policy change deserves investigation even when aggregate throughput looks healthy.
Before deployment, walk one job through success, explicit cancellation, automatic budget cancellation, worker loss, retry, and the completion race. Verify authorization at the API boundary, atomic transitions in storage, cancellation checks in every expensive loop, idempotent cleanup, and retention rules for artifacts and audit events. Then run the same cases with realistic marketplace image fixtures. The implementation is ready when every outcome is bounded, reproducible, and explainable.
Top comments (0)