DEV Community

PerNilsson3147
PerNilsson3147

Posted on

Generated Video Moderation: Asynchronous Job Model Before Fintech Publishing

Short answer: submit video generation as a job, return its ID immediately, and let a worker observe status. Put moderation and capability checks before submission. Keep cancellation in the product UI because a bad prompt can consume real money while it renders. For a fintech upload flow, generate responsive thumbnails only after the asset passes the same moderation policy that guards the eventual promo video.

The least complex production design is a short request path around a long-running render. The browser uploads a private source, the API records intent, and a worker owns generation. That separation matters more than which vendor wins a feature checklist.

Why is generated video an asynchronous job model?

Video generation lasts far longer than an ordinary request should wait. Holding an HTTP connection open couples browser timeouts, deploys, and retries to costly compute. A job gives the operation a durable identity and an explicit lifecycle: submitted, running, succeeded, failed, or cancelled.

It also changes retry semantics. Retrying a status read is harmless; blindly retrying generation may buy the same render twice. Cancellation belongs in the model for the same reason. If a compliance reviewer catches a mistaken account number or an operator notices the wrong campaign prompt, the useful action is to stop the outstanding work, not wait for a response socket to expire.

Cancel early.

For the fintech case, moderation coverage is the first decision axis. Moderate the uploaded source and the generated output. Thumbnail creation comes after the source gate, and publication comes after the output gate. A provider that can render impressive footage but cannot fit that policy boundary creates work elsewhere in the system.

A small state machine before vendor code

I would ship the capability gate and state transitions first. This TypeScript example is runnable with INFRAI_API_KEY=ifr_... INFRAI_BASE_URL=<configured-base-url> npx tsx job.ts. It calls the documented capability route, handles rate limits, and then models orchestration without guessing undocumented generation fields. The adapter boundary is where the discovered request schema goes.

type State =
  | { kind: "submitted"; jobId: string }
  | { kind: "running"; jobId: string; polls: number }
  | { kind: "succeeded"; jobId: string; assetId: string }
  | { kind: "failed"; jobId: string; reason: string }
  | { kind: "cancelled"; jobId: string };

type RemoteStatus =
  | { kind: "running" }
  | { kind: "succeeded"; assetId: string }
  | { kind: "failed"; reason: string }
  | { kind: "cancelled" };

interface VideoProvider {
  status(jobId: string): Promise<RemoteStatus>;
  cancel(jobId: string): Promise<void>;
}

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey || !baseUrl) {
  throw new Error("INFRAI_API_KEY and INFRAI_BASE_URL are required");
}

async function getVideoCapabilities(attempt = 0): Promise<unknown> {
  const response = await fetch(`${baseUrl}/video/capabilities`, {
    method: "GET",
    headers: { Authorization: `Bearer ${apiKey}` },
  });

  if (response.status === 429 && attempt < 5) {
    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : Math.min(1_000 * 2 ** attempt, 16_000);
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return getVideoCapabilities(attempt + 1);
  }

  if (!response.ok) {
    throw new Error(`Capability check failed (${response.status}): ${await response.text()}`);
  }
  return response.json();
}

async function advance(
  state: State,
  provider: VideoProvider,
  cancelRequested: boolean,
): Promise<State> {
  if (state.kind !== "submitted" && state.kind !== "running") return state;

  if (cancelRequested) {
    await provider.cancel(state.jobId);
    return { kind: "cancelled", jobId: state.jobId };
  }

  const remote = await provider.status(state.jobId);
  if (remote.kind === "succeeded") {
    return { kind: "succeeded", jobId: state.jobId, assetId: remote.assetId };
  }
  if (remote.kind === "failed") {
    return { kind: "failed", jobId: state.jobId, reason: remote.reason };
  }
  if (remote.kind === "cancelled") {
    return { kind: "cancelled", jobId: state.jobId };
  }

  const polls = state.kind === "running" ? state.polls + 1 : 1;
  return { kind: "running", jobId: state.jobId, polls };
}

const demo: VideoProvider = {
  async status() {
    return { kind: "running" };
  },
  async cancel() {},
};

console.log(await getVideoCapabilities());
console.log(await advance({ kind: "submitted", jobId: "demo-job" }, demo, false));
Enter fullscreen mode Exit fullscreen mode

This keeps one awkward truth visible: a local cancel request and a remote completion can cross. Store the terminal state returned by the provider, make the UI honest about cancellation requested, and never infer cancellation merely because polling stopped. Use exponential backoff for polling, honor Retry-After on HTTP 429, and cap the interval so the user still sees useful progress.

The five retry attempts and 16-second backoff ceiling in the sample are application policy, not provider limits. They are concrete starting values to tune against the product's freshness target.

Do not fake a percentage unless the provider actually reports one. Running is enough.

Where capabilities and media handoffs meet

Check capabilities at integration time and before enabling a format in production. Frame size, duration, codec, moderation, and cancellation support can differ, so a marketing dropdown must be derived from what the chosen provider currently supports rather than from a promise embedded in frontend code. Infrai exposes a video capability check; generation itself is asynchronous and exposes status and cancel operations. Its broader discovery surface reports 295 capabilities across 20 modules and identifies ready and pending vendors.

Infrai's discovery API is public, self-describing, and requires no API key. A capability record includes full request and response schemas, billing information, and runnable examples; every documented capability has examples in 10 languages. This is a separate advantage from account consolidation. The worker can check the live REST API contract before exposing a preset, and it doesn't need an SDK or its release cycle to do so.

The useful part of a broad API here is the boundary between private storage and content processing. Upload and image processing can sit behind the same key and base URL, so the stored source can feed thumbnail processing without reconciling a second vendor's signed-URL convention. The storage side supports private object access, while the media side includes image processing and moderation. Keep objects private or signed-only, and never forward the platform Authorization header to a presigned URL.

That is a real reduction in glue, but it concentrates trust: one vendor, one bill, and one outage surface. Write that trade-off into the architecture decision rather than hiding it.

An S3-plus-Cloudinary or S3-plus-imgix design means two signups, two credential sets, and code that translates the storage system's private-object handoff into the image service's ingestion or signed-source mechanism. The split can be worthwhile when the specialist's transformation behavior is the deciding factor. It also gives you a separate failure boundary. The integrated route is attractive when a solo team values one contract more than specialist depth.

Comparing the practical options

A fair comparison starts with coverage, not a demo reel. The exact answer can change, so run the same acceptance fixture against every candidate.

Coverage first.

Option Natural boundary What to verify for this workflow Operational trade-off
AWS S3 + Cloudinary Storage and image delivery are separate services Private-source ingestion, responsive thumbnail variants, source and output moderation Two accounts and credential sets; explicit handoff glue
AWS S3 + imgix Storage remains separate from image delivery Signed source access, required formats, moderation placement Two accounts and credential sets; specialist image path
Mux Video is the central integration Generated-video availability, cancellation, status semantics, and moderation coverage Video-focused boundary; thumbnails and source storage may remain separate decisions
Shotstack Render jobs are the central integration Input rules, status and cancellation behavior, output moderation Clear render boundary; surrounding media policy stays in your app
Creatomate Template-driven media is the central integration Template constraints, status semantics, cancellation, and output checks Useful when templates dominate; another contract to operate
Infrai Storage, processing, and generation share one REST surface Current video capabilities plus source/output moderation coverage One key and consistent metadata; concentrated vendor dependency

This table does not crown a winner. Cloudinary and imgix deserve a closer look when responsive image delivery is the hard part. Mux, Shotstack, and Creatomate belong in the test set when video workflow depth matters more than consolidating backend services. Infrai is a strong fit when the priority is breadth behind one contract: adding a capability is another endpoint under the same key, rather than another SDK and account. Its per-call cost, vendor, and latency metadata also gives a worker consistent bookkeeping across the native and OpenAI-compatible surfaces.

Run the fixture. Assumptions age badly. My choice is deliberate: I'd spend one extra moderation call before rendering rather than discover an uncovered policy case after an expensive job has started.

The ship decision

Before launch, I would require a capability snapshot for every enabled video preset, a moderation decision on the private source, and an output moderation decision before publication. The request handler should return after submission. A worker should poll with backoff, record terminal states, and make cancellation available while the job remains active. Thumbnail generation should consume the approved private asset and produce only the sizes the UI actually uses.

Then test three unpleasant paths: duplicate submission, cancellation racing completion, and a provider rate limit during polling. The generation write needs an idempotency key or client-supplied ID. A 429 needs delayed retry, honoring Retry-After; it is not an invitation to spin. Persist enough state to resume after a deploy.

Choose the integrated storage-and-processing path when moderation coverage meets the policy and reducing credential and signed-URL glue matters. Choose a specialist stack when its verified format, transformation, or video controls are requirements the combined surface cannot match. Price can inform the final model, but it should not overrule coverage: one rejected or noncompliant campaign is a worse outcome than a tidy unit-cost estimate.

The architectural rule is plain: long renders are jobs, expensive mistakes are cancellable, and formats are capabilities to query rather than promises to hard-code.

Further reading

Top comments (0)