DEV Community

Cover image for How to Build a Reliable AI Video Generation Pipeline
陳小聰
陳小聰

Posted on

How to Build a Reliable AI Video Generation Pipeline

I work on Bulin47AI, the project used as a case study below. This draft was prepared with AI assistance.

An AI video generation pipeline has an awkward failure mode: your server sends a request, the connection times out, and you cannot tell whether the provider started the job. Retrying may create a second billable generation. Telling the user it failed may be equally wrong.

This design problem shaped [Bulin47AI](https://bulin47ai.net/), our small studio built around one prepared video scene. A user supplies one portrait and chooses a 5, 12, or 23-second cut. The model attempts to replace the central performer; our application is designed to restore the source audio and deliver a private MP4. The model can still miss the target or change background details, so the result needs review.

The useful lesson is the job design around that model call. Here is a pattern you can adapt to other asynchronous media APIs.

1. Define a narrow generation contract

Before sending a request, decide exactly what the user approved. For a fixed-source workflow, that means more than a prompt:

  • A versioned source clip and its hash.
  • The selected cut, starting point, output resolution, and aspect ratio.
  • The portrait file after server-side validation.
  • A quote version and the credit amount shown to the user.
  • A stable job ID and the account that owns it.

Our creation route checks the selected source hash and recalculates the quote on the server. If either changed since the user opened the form, it asks for another review instead of silently applying the new settings. A generic prompt-based product may have different fields, but the rule is the same: persist an immutable description of the accepted request.

ts:
「
// Simplified model, not a copy of the production schema.
type JobStatus =
| "preparing"
| "uploading"
| "submitting"
| "submission_unknown"
| "queued"
| "in_progress"
| "processing"
| "completed"
| "failed";

type GenerationJob = {
id: string;
ownerId: string;
sourceHash: string;
sourceVersion: string;
quoteVersion: string;
reservedCredits: number;
status: JobStatus;
providerRequestId?: string;
};
」

This record gives the user something concrete to return to if their browser closes or the request stalls.

2. Reserve credits before the billable boundary

An asynchronous job should have a clear order of operations:

  1. Validate the account, portrait, source selection, consent, and quote.
  2. Create the job and reserve its credits atomically.
  3. Save the inputs in private storage.
  4. Persist a submitting state before calling the provider.
  5. Submit using a stable idempotency key tied to the job ID.
  6. Save the provider's request ID as soon as it is returned.

The fourth step matters. If a worker crashes after the provider receives the request, a later worker must not interpret the job as a fresh request. The saved state tells it that submission may already have happened.

Do not treat an HTTP timeout as proof of rejection. Only a definitive provider rejection should move directly to a failed state with an automatic credit return. For an ambiguous outcome, hold the job for reconciliation.

ts「
await saveJob({ ...job, status: "submitting" });

try {
const receipt = await provider.submit(input, {
idempotencyKey: job.id,
});
await saveJob({ ...job, status: "queued", providerRequestId: receipt.id });
} catch (error) {
if (isDefinitiveRejection(error)) {
await failAndReturnCredits(job);
} else {
await saveJob({ ...job, status: "submission_unknown" });
}
}
」

This is illustrative code. Its safety depends on the provider's idempotency behavior and on your own atomic credit and job operations. Test those properties rather than assuming a header solves duplicate billing.

3. Give uncertainty its own state

Many job systems model only pending, completed, and failed. That leaves no honest answer when submission is uncertain.

Our workflow uses submission_unknown for that case. It does not resubmit during status polling. The user sees that the request needs review, and the reserved credits stay held until the outcome is known. That is less convenient than an automatic retry, but it avoids promising a refund while a chargeable generation may still be running.

A useful state path is:

preparing -> uploading -> submitting -> queued -> in_progress
                                             -> processing -> completed
                         \-> submission_unknown
Enter fullscreen mode Exit fullscreen mode

In a production system, reconciliation should check any provider receipt, webhook, dashboard record, or idempotency lookup available. Record the evidence used to settle the job. If your provider cannot answer whether a timed-out call was accepted, make that limitation visible to operators and users.

4. Treat the returned video as untrusted input

“Provider says completed” is not the same as “the user has a playable result.” Downloading, processing, storing, and serving the output are separate steps that can each fail.

For our fixed scene, we check the generated video's dimensions and duration against the approved source. We allow only small timing drift, then map the generated video stream to the original source audio. After processing, we probe the final MP4 again for expected codecs, size, dimensions, and duration. Only then does the job become completed.

The distinction also improves error handling. A model failure and a temporary download or storage failure need different recovery paths. If the model already ran, a storage retry should not start another generation.

Keep the source portrait and finished file private. A result route should verify the signed-in user owns the job before returning media, including byte-range playback and downloads. A hard-to-guess URL alone is not a substitute for an ownership check.

5. Show the same states to users and operators

Users need a durable result page that can answer three questions: Did my request start? What is happening now? Were my credits used or returned? Operators need the corresponding job ID, provider request ID, timestamps, and error category.

Avoid a progress bar that implies precision your provider does not supply. Clear states such as “uploading,” “generation in progress,” “processing,” and “needs review” are more useful than a fabricated percentage. If a user retries, make it a deliberate new job with a new quote, not a hidden second call behind a refresh button.

Frequently asked questions

Is an idempotency key enough to prevent duplicate charges?

Only if the provider documents and honors its behavior for your request type. Persist the local submission state first, retain the key, and test timeout and replay scenarios.

When should credits be returned?

Return them after a definitive failure under your billing rules. An ambiguous submission needs reconciliation before you classify it as failed.

Why validate the output if the provider reports success?

The returned file may have the wrong duration, framing, audio, or codec, or may fail during download. Validate the final file that the user will actually play.

Conclusion

A reliable AI video generation pipeline is mostly about preserving the truth of a request across slow, billable, and uncertain steps. Start with a versioned input contract, persist the submission boundary, represent unknown outcomes explicitly, and validate the final media before marking a job complete. Those safeguards matter even when the creative workflow is as narrow as a single scene.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to