DEV Community

zhenyu xu
zhenyu xu

Posted on

A callback-first recovery pattern for AI video jobs on Vercel

An AI video job should keep running when the user closes the tab. That makes the browser a useful progress display, but a poor owner of the job lifecycle.

In H3 Max, the video tool I operate, we use a callback-first flow with a delayed recovery check. This is a short implementation note about that design, not a benchmark or a claim that it eliminates background costs.

Separate completion from recovery

The normal path lets the model provider's callback save the result. The callback does not create a Workflow just to save a completed video. If a recovery Workflow already exists, the callback can wake it.

When a job is created, a delayed queue message is also registered with a 60-second delay. That message carries a job ID. When delivered, the consumer reads the current database row rather than assuming the original work still needs attention.

The consumer returns immediately if the job has been deleted or has reached a terminal state. If the next check is not due, or another worker still holds a lease, it defers the check. Only unresolved, eligible work is handed to a recovery Workflow.

What the delay does — and does not — mean

Sixty seconds is an initial recovery checkpoint, not a promise that every model call finishes within a minute. The consumer may defer again based on the job's current state. The queue still incurs work even when the callback succeeds first; what it avoids is starting a recovery Workflow for every successful job.

This also makes failure handling explicit. A failed Workflow handoff throws so the queue can redeliver the message. A duplicate delivery must check the current state again. Provider completion, database persistence and UI progress are separate concerns.

Checks worth keeping

  • A callback completes the job before the delayed message arrives.
  • The callback never arrives, so recovery must take over.
  • A message is delivered more than once.
  • A job is deleted while background work is pending.
  • The Workflow handoff fails temporarily.

This pattern trades constant polling for an event-driven normal path plus a durable recovery path. Its cost depends on callback reliability, job duration and retry behavior, so I would measure those before choosing a delay for a different application.

Disclosure: I operate H3 Max. This article was prepared with AI assistance using the project's current implementation.

Top comments (0)