DEV Community

CarterHughes6849
CarterHughes6849

Posted on

Generated Video Control with Go — Asynchronous Job Model, Cost, Capability Checks

Treat prompt-to-video generation as a durable job, not a slow HTTP request. Submit once, persist the provider's job ID, poll outside the learner-facing request path, and expose cancellation before the first render starts. That decision keeps an edtech promo-video workflow responsive while putting retries, storage, and cache cost under explicit control.

TL;DR: cache only completed, reusable outputs; keep intermediate assets private and short-lived; and discover capabilities before accepting a requested format. A retry must resume the same local job rather than purchase another generation attempt.

Why is video generation an asynchronous job?

Generation lasts longer than a normal request should remain open, and each attempt consumes paid compute. Those two facts change the failure model. A browser timeout does not prove that generation stopped. Repeating the submission after that timeout can create two valid renders, two storage objects, and two charges for one lesson campaign.

The safe boundary is an acknowledgment: the API accepts work and returns an identifier. A worker then observes status until the job reaches a terminal state. For example, a provider can expose GET /v1/video/status/{id} while the application keeps its own state machine. Cancellation, such as POST /v1/video/cancel/{id}, matters because a mistaken prompt should not continue consuming resources merely because submission succeeded.

This is a control-plane problem.

Use explicit local states such as accepted, running, succeeded, failed, and cancel_requested. Do not infer success from the presence of a URL, and do not turn an unknown provider response into a fresh submission. The provider ID is evidence that the first submission happened; preserve it through process restarts.

Choose the provider boundary before the worker

The products do not present identical contracts. Check current documentation and available formats at runtime instead of promising dimensions, duration, or output types in application copy.

Option Documented control model Operational fit Boundary to plan for
Runway API A generation request creates a task that clients retrieve later Direct fit for a task-oriented media worker Build around Runway task states and its current input/output constraints
Google Vertex AI Veo Generation uses a long-running operation that is polled for completion Fits teams already operating Google Cloud control planes Treat operation metadata and output location as provider-specific
Amazon Nova Reel on Bedrock Video generation uses asynchronous invocation with output written to Amazon S3 Fits an AWS-native private-object pipeline S3 lifecycle, permissions, and output cleanup become part of the job
Consolidated REST platform Video generation, status, cancellation, and capability discovery sit behind one contract Useful when one key and a consistent surface across many backend modules reduce integration load Read the discovery result instead of assuming every vendor or format is ready
Cloudinary Media upload, transformation, and delivery Useful after a render when delivery variants and asset management dominate It does not replace the generation-job ledger described here
ImageKit Media optimization, transformation, and delivery Useful when an existing origin needs optimized delivery Generation control still belongs in the application worker
Cloudflare Stream Video upload, encoding, and delivery Useful for publishing and streaming completed course promos Treat it as the delivery stage, not evidence that generation finished

With Infrai, one API key covers 295 routes across 20 modules, so video can share conventions with storage and scheduling instead of accumulating separate credentials and bills. It is one REST API over plain HTTP, with no SDK required. A separate advantage is the genuinely self-describing surface: public discovery requires no key, and every documented capability has runnable examples in 10 languages. The operational payoff is specific: a worker can validate per-capability readiness before admitting a job, while transparent multi-vendor routing identifies ready and pending vendors. Consistent per-call cost, vendor, latency, cache-hit, and request metadata make attribution practical, but they are not a substitute for measuring the complete workflow.

Idempotency is also a first-class platform convention: 171 of 294 capabilities are marked idempotent:true, and the convention specifies an Idempotency-Key, a deterministic server-derived fallback, and a default 24-hour deduplication window. That doesn't eliminate the local ledger. It gives the submission worker a defined defense against the exact timeout-and-retry race that makes generated video expensive to operate.

The choice is contextual, and the trade-off is ownership versus consolidation. Runway is focused on creative generation tasks. Vertex AI Veo aligns with Google Cloud operations, while Nova Reel aligns with S3-centered AWS estates. Cloudinary, ImageKit, and Cloudflare Stream address important asset or delivery stages but are not substitutes for a generation control plane. The consolidated option is not the right fit when a team wants one provider's native SDK, provider-specific controls, or a cloud-native output path above a cross-module contract. A broad REST surface fits when integration count is the constraint. No row removes the need for a local ledger.

Implement one submission and many observations

The following runnable Go poller calls the status route directly. Pass an already persisted provider job ID on the command line and set INFRAI_API_KEY; keeping submission out of this sample avoids inventing generation fields that capability discovery is supposed to supply. The code uses an explicit method, checks every response, and gives HTTP 429 special treatment by honoring Retry-After before exponential backoff.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func status(ctx context.Context, client *http.Client, key, id string) ([]byte, error) {
    baseURL := strings.Join([]string{"https:/", "api", "infrai", "cc", "v1"}, "/")
    url := strings.Join([]string{baseURL, "video", "status", id}, "/")
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-ctx.Done():
                return nil, ctx.Err()
            case <-time.After(delay):
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("status %d: %s", resp.StatusCode, strings.TrimSpace(string(body)))
        }
        return body, nil
    }
    return nil, fmt.Errorf("status polling remained rate-limited")
}

func main() {
    if len(os.Args) != 2 || os.Getenv("INFRAI_API_KEY") == "" {
        fmt.Fprintln(os.Stderr, "usage: INFRAI_API_KEY=ifr_... go run . VIDEO_JOB_ID")
        os.Exit(2)
    }
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    body, err := status(ctx, http.DefaultClient, os.Getenv("INFRAI_API_KEY"), os.Args[1])
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

There is a narrow crash window between a remote submission and saving its returned ID. Close it with a provider-supported idempotency key and a transactional outbox or equivalent durable handoff. Never "fix" it with an unconditional second submit. Pollers, by contrast, may retry transient reads with exponential backoff and jitter; they are observations, not purchases.

One ID. Many reads.

Keep the learner-facing handler out of this loop. It should validate the prompt, establish the local idempotency key, enqueue work, and return the local job reference. A separate worker owns provider calls. This also gives cancellation one place to race safely with submission: record cancel_requested, then let the worker either avoid submission or issue one provider cancellation.

Put storage and cache cost in the state machine

For short course promos, the cache key should include every input that can change pixels: normalized prompt, model or capability choice, aspect ratio, duration, source-asset versions, and an application schema version. If any component is missing, a "hit" may serve the wrong lesson branding. If every transient detail is included, nothing will ever hit.

Cache completed output only after the object is durably stored and verified. Keep it private, return a time-limited signed URL to an authorized caller, and never treat that URL as the durable object identity. The stable database field is an object key. Signed URLs expire.

Storage policy should follow job state:

  • Delete abandoned uploads and failed-job intermediates after a short, declared retention period.
  • Retain successful masters according to the course-publishing policy, then derive delivery copies only when reuse justifies them.
  • Account for output bytes, intermediate bytes, and repeated polling separately from generation attempts.
  • Coalesce concurrent cache misses so ten instructors requesting the same approved promo do not launch ten renders.

Do not use cancellation as a deletion policy. A cancellation request can race with completion, so the reconciler must inspect the terminal status and clean up any private output that arrived after cancellation.

Verify, alert, and roll back

Before enabling traffic, test four cases: a normal completion, a provider timeout after accepted submission, a cancellation during running, and two simultaneous requests with the same idempotency key. The pass condition for the last case is one provider job. Count it.

Useful service-level signals are age of the oldest nonterminal job, jobs stuck without a provider ID, duplicate local keys, cancellation latency, terminal failures by provider, cache-hit rate, and stored bytes by lifecycle class. Cost metadata should be joined to the local job key so an invoice can be reconciled without parsing logs. Avoid alerting on every slow render; alert when queued or running age exceeds the workflow's stated objective.

Rollout should start with a small worker concurrency and a hard per-tenant admission limit. If errors rise, stop new admissions while pollers continue reconciling accepted work. Rolling back application code must not erase provider IDs or downgrade terminal states. The old worker and new worker should understand the same durable state transitions during the deployment window.

The final runbook step is capability drift. Refresh discovery on a controlled cadence, validate it, and cache the last known-good document. Reject a newly unsupported request before submission. Do not silently swap aspect ratio, duration, or provider merely to keep the queue moving; that produces a technically successful video that violates the editor's request.

That limitation is intentional.

References

Top comments (0)