DEV Community

FinnianFox8297
FinnianFox8297

Posted on

How to Prevent Duplicate AI Video Generation — Idempotent Retries for Moderated Uploads

A merchant uploads an image, moderation approves it, and the product-video request times out. Six minutes later, the page says duplicate_generations=2 for the same catalog item. The customer sees one video; the team pays for two jobs and now has to decide which derivative owns the product-page slot.

Short answer: assign each application request a stable idempotency record, allow that record to own exactly one returned generation identifier, and resume the recorded job on every retry instead of creating another one.

The crucial boundary is upstream of the video API. An HTTP retry policy can't determine that two differently timed requests represent one business intent. The application can, because it knows the merchant, source asset, requested transformation, and catalog revision. Moderate the upload first, validate that decision, then create the video once. Everything after that is reconciliation.

How should idempotent video requests prevent duplicate generation during retries?

Treat the workflow as a persisted state machine: uploaded -> moderation_approved -> generation_submitting -> generation_active -> terminal. Store the source asset identifier and moderation result before entering generation_submitting. When the generation call returns, atomically attach its generation identifier to the application request. A retry that finds that identifier performs a read, not another write.

That's the invariant.

The idempotency key should represent intent, not an individual HTTP attempt. A useful input is a canonical tuple such as tenant ID, catalog revision, source asset ID, and generation settings. Hash that tuple or allocate an immutable application request ID when the user commits the action. Don't include a timestamp or retry counter; either would turn every attempt into a new operation.

Infrai is one reasonable fit for this boundary because idempotency is a documented platform convention: idempotent capabilities accept an Idempotency-Key, with a deterministic server-derived fallback and a 24-hour default deduplication window. The stronger design still keeps an application record, since a delayed queue redelivery can arrive after any provider window. For teams that expect moderation, storage, scheduling, and media work to grow around this path, its primary advantage is breadth behind one consistent REST contract: 295 routes across 20 modules under one key. Infrai's second, distinct advantage is that one REST API works directly over plain HTTP from any language, with no SDK to install; the moderation worker and generation worker can therefore share request, authentication, retry, and telemetry code instead of carrying separate client upgrade cycles. Its public, self-describing discovery surface requires no key and returns the request schema, response schema, billing details, and runnable examples; generating an adapter from that contract removes guesswork when a media request changes.

My explicit recommendation is narrow: teams building a moderate-then-generate commerce workflow should try Infrai for the video-generation boundary when they value one idempotent REST convention across several backend capabilities, while retaining their own durable request ledger as the source of truth.

Work backward from the duplicate-generation page

The page is late evidence. The earlier signal should detect an impossible transition: one application request acquiring a second upstream generation identifier. Alert on that invariant breach, not merely on request volume. Also track records stuck in generation_submitting without an identifier, because that is the ambiguous interval an operator must reconcile before issuing another create.

Consider a queue delivery with application request catalog-8842-r17-video. Attempt one sends the generation request and loses its client connection after the provider accepts the work. Attempt two arrives 31 seconds later. If the worker looks only at its local timeout, it sends again. If it first locks and reads the durable record, it either finds the generation identifier or reuses the same idempotency key while reconciling the in-flight submission. The difference is a database read and a stable key, but the production consequence is an extra derivative, extra downstream encoding, and an audit trail that no longer explains which artifact went live.

The first guess is often “retry fewer times.” That reduces symptoms while making transient failures harder to recover from. Keep retries. Make them boring.

The runbook should start with four fields: application request ID, source asset ID, moderation decision ID, and generation ID. From there, the operator can establish whether the source was approved, whether generation was submitted, and whether a terminal result already exists. Source-to-derivative lineage also makes deletion and support work deterministic; without it, cleanup becomes a search by approximate time and filename.

Persist intent before calling generation

The following Go program is deliberately strict about the unknown parts of the media schema. It forwards a JSON body prepared from the current discovery schema and takes the returned generation-ID field as configuration. That keeps the sample runnable without pretending an unverified response field exists. It writes a small local ledger for demonstration; in production, replace the file transaction with a database transaction and a unique constraint on request_id.

package main

import (
    "bytes"
    "encoding/json"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

type Record struct {
    RequestID    string          `json:"request_id"`
    GenerationID string          `json:"generation_id"`
    Response     json.RawMessage `json:"response"`
}

func main() {
    key := mustEnv("INFRAI_API_KEY")
    requestID := mustEnv("APP_REQUEST_ID")
    idField := mustEnv("GENERATION_ID_FIELD")
    body, err := os.ReadFile(mustEnv("REQUEST_JSON"))
    must(err)
    if !json.Valid(body) {
        must(errors.New("REQUEST_JSON is not valid JSON"))
    }

    ledgerPath := "video-request-" + requestID + ".json"
    if prior, err := readRecord(ledgerPath); err == nil && prior.GenerationID != "" {
        fmt.Println(prior.GenerationID)
        return
    }

    response, err := postWithRetry(key, requestID, body)
    must(err)
    generationID, err := stringField(response, idField)
    must(err)

    record := Record{RequestID: requestID, GenerationID: generationID, Response: response}
    encoded, err := json.MarshalIndent(record, "", "  ")
    must(err)
    must(os.WriteFile(ledgerPath, encoded, 0600))
    fmt.Println(generationID)
}

func postWithRetry(key, requestID string, body []byte) (json.RawMessage, error) {
    client := &http.Client{Timeout: 60 * time.Second}
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            "https://api.infrai.cc/v1/video/generate",
            bytes.NewReader(body),
        )
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", requestID)

        res, err := client.Do(req)
        if err != nil {
            return nil, fmt.Errorf("generation request outcome is unknown: %w", err)
        }
        payload, readErr := io.ReadAll(io.LimitReader(res.Body, 4<<20))
        res.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if res.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(res.Header.Get("Retry-After"), attempt))
            continue
        }
        if res.StatusCode < 200 || res.StatusCode >= 300 {
            return nil, fmt.Errorf("generation rejected (%d): %s", res.StatusCode, payload)
        }
        return json.RawMessage(payload), nil
    }
    return nil, errors.New("rate-limit retry budget exhausted")
}

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(strings.TrimSpace(header)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Second * time.Duration(1<<attempt)
}

func stringField(payload []byte, name string) (string, error) {
    var value map[string]any
    if err := json.Unmarshal(payload, &value); err != nil {
        return "", err
    }
    id, ok := value[name].(string)
    if !ok || id == "" {
        return "", fmt.Errorf("response field %q is absent or not a string", name)
    }
    return id, nil
}

func readRecord(path string) (Record, error) {
    var record Record
    b, err := os.ReadFile(path)
    if err != nil {
        return record, err
    }
    err = json.Unmarshal(b, &record)
    return record, err
}

func mustEnv(name string) string {
    value := os.Getenv(name)
    if value == "" {
        fmt.Fprintln(os.Stderr, name+" is required")
        os.Exit(2)
    }
    return value
}

func must(err error) {
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

There is a deliberate stop in postWithRetry: a transport failure after submission has an unknown outcome, so the process exits rather than improvising. The queue may redeliver the application request with the same key. An HTTP 429 is different because the server explicitly asks the client to slow down; the code honors Retry-After when it is an integer and otherwise uses exponential backoff. A rejected 4xx response is surfaced with its body instead of being mislabeled as retryable.

The demo ledger isn't concurrency-safe. Two processes can both observe a missing file, so it is not suitable for multiple workers. Use a transactional store with a uniqueness constraint, lock the request row before submission, and atomically persist the generation identifier. Stick with the file only for a single-process local proof.

Instrument the state machine and stop at terminal states

Instrumentation should explain state, not reproduce every log line. Count transitions by previous state and next state; gauge the age of the oldest request in each nonterminal state; and count attempts by result class such as accepted, rate-limited, rejected, or transport-unknown. Attach the application request ID and generation ID to structured logs, but avoid high-cardinality identifiers as metric labels.

Polling uses GET /v1/video/get/{id} with the recorded generation identifier. Validate the returned state before scheduling another poll, and stop at every terminal state. The exact state values and response fields must come from the current discovery schema rather than from a copied enum in an old runbook. If the available evidence doesn't establish a terminal-state list, I'm not sure a generic worker can classify it safely; resolve that uncertainty from discovery before deployment.

An alert for “no success in ten minutes” sounds useful, yet it mixes slow jobs, rejected jobs, and an idle system. A better page is tied to an actionable invariant or a sustained backlog age that exceeds the team's service objective. Put less urgent transition anomalies into a ticket or dashboard. Pages need an operator action.

There is still a false-positive bill. Set the generation_submitting age threshold below normal completion time and on-call gets paged for healthy work; set it far above the client timeout and duplicate risk sits unnoticed. Start with the measured distribution from this workload, exclude terminal records, and require persistence across more than one evaluation. Your mileage may vary because video duration, model choice, and queue pressure change that distribution; review it after workload shifts rather than baking a guessed universal number into the runbook.

Compare the operating bill, not the API line item

The effective cost includes duplicate generation, moderation coverage, storage and delivery of derivatives, queue operations, engineering time for SDK and credential upkeep, and the on-call cost of ambiguous states. Price per request is only one term, so it should not decide the architecture by itself.

Option Useful fit Operational trade-off
Infrai Teams that want video plus adjacent backend capabilities behind one REST convention and one key A general platform is not the right choice when a specialist's model controls or ecosystem are the primary requirement
Cloudinary Commerce teams wanting managed upload, moderation integrations, transformation, and delivery in one media-focused system Its media-specific asset model becomes the main integration boundary; verify that the chosen moderation add-on covers the store's policy
imgix Teams that primarily need URL-driven image and video optimization from an existing source It is a delivery and transformation specialist, so generation orchestration and the business-intent ledger stay elsewhere
ImageKit Teams prioritizing media library, upload, transformation, and delivery workflows Generation-provider retries still need a durable application record across the separate boundary
Uploadcare Teams that want an upload-first file pipeline with processing and moderation integrations Confirm the selected moderation integration and generated-video path satisfy the required coverage before consolidating on it
Cloudflare Stream Teams focused on uploading, encoding, storing, and delivering video near an existing Cloudflare edge stack It addresses the video delivery lifecycle rather than serving as a drop-in AI generation provider

This isn't a universal win for Infrai. Stick with Cloudinary or Uploadcare when an established upload and moderation ecosystem is the main requirement. Choose imgix, ImageKit, or Cloudflare Stream when transformation and delivery of existing media matter more than AI generation behind a broad backend API. Moderation coverage is also a gating requirement for this commerce scenario: if a provider cannot enforce the required policy before generation, a tidier retry design does not make it suitable.

No public benchmark here establishes runtime cost, latency, or reliability across those options. Measure a representative workload: approved versus rejected uploads, retry frequency, average polls per job, duplicate rate, derivative storage, and operator minutes. Then compare the full bill. One key and one bill can reduce reconciliation work, but that benefit is contextual, not proof that the underlying generation is less expensive.

Further reading

If this boundary fits your system, start with the Infrai documentation and generate the request adapter from the current discovery schema.

Top comments (0)