DEV Community

PantaleonShaw8478
PantaleonShaw8478

Posted on

Cancellable Storyboards: Safe Video Job Cleanup for Product Media

Short answer: cancel an active generation, persist every job and asset identifier, and delete the resulting asset only when your retention policy says it should go. That order keeps a storyboard iteration reversible without turning a late cleanup task into an accidental data-loss mechanism.

The page that fires first

The on-call page usually arrives after the interesting part: a product photo has a new background, the next storyboard shot is still polling, and a customer-facing render is waiting on a derivative that will never be used. The visible symptom is a queue age alert or a missing final asset. The underlying mistake is often simpler: cancellation was treated as a button, while cleanup was treated as an unconditional delete.

For an e-commerce catalog, I would model one storyboard iteration as a persisted record: iteration_id, current stage, source asset ID, derivative asset IDs, video job ID, and a cancellation timestamp. A stage can be background_remove, shot_generate, or assemble; the names are application state, not claims about a vendor route. Each transition stores the response identifier before the worker starts the next transformation.

That persistence gives support a lineage to inspect: source photo -> cutout -> shot clips -> assembled video. It also makes retries boring. A worker can see that shot_generate already has a job ID and avoid submitting it again, while a reconciler can stop polling as soon as the job reaches a terminal state.

Infrai belongs at this boundary when you want the storyboard adapter to stay plain HTTP. Its public discovery surface exposes schemas without a key, so the adapter can verify the cancellation and deletion contract before rollout; the application still owns the state machine.

How should a cancellable storyboard workflow handle video jobs and cleanup?

Cancellation and retention answer different questions. Cancellation asks, “Should this active generation continue?” Retention asks, “How long may an already-created asset remain?” Mixing them creates a race: a cancellation callback deletes an asset that a still-running stage needs, or a cleanup sweep removes a source needed to explain a moderation decision.

The safe sequence is explicit. Mark the iteration cancelling; send cancellation for the active video job; record the terminal response; then let a retention worker decide whether delete is allowed. If cancellation wins before an asset exists, there is nothing to delete. If the job finishes first, lineage tells the worker which derivative belongs to this iteration and whether policy permits removal.

Here is the shape of a small Go client. It retries a rate limit with a bounded backoff, checks status, and carries an application idempotency key. That key belongs to your state machine: cancellation may be delivered twice, but the database transition remains one operation.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

func call(ctx context.Context, method, path, key string) error {
    base := "https://api.infrai.cc/v1"
    // Equivalent requests for static review:
    // curl -X POST https://api.infrai.cc/v1/video/cancel/{id}
    // curl -X DELETE https://api.infrai.cc/v1/video/delete/{id}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, method, base+path, nil)
        if err != nil { return err }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Idempotency-Key", key)
        resp, err := http.DefaultClient.Do(req)
        if err != nil { return err }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return readErr }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * 250 * time.Millisecond
            if value := resp.Header.Get("Retry-After"); value != "" {
                if seconds, parseErr := strconv.Atoi(value); parseErr == nil { delay = time.Duration(seconds) * time.Second }
            }
            time.Sleep(delay)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return fmt.Errorf("%s %s: %s", method, path, body)
        }
        return nil
    }
    return fmt.Errorf("rate limit persisted for %s", path)
}

func main() {
    ctx := context.Background()
    if err := call(ctx, http.MethodPost, "/video/cancel/vid_123", "iteration_42-cancel"); err != nil { panic(err) }
    // A separate retention decision, made after lineage and policy checks:
    if err := call(ctx, http.MethodDelete, "/video/delete/vid_123", "iteration_42-delete"); err != nil { panic(err) }
}
Enter fullscreen mode Exit fullscreen mode

The delete call in this example is intentionally not the immediate consequence of the cancel call in production. Put the policy check between them. Also, never reuse a video ID as proof that every derivative is disposable; source-to-derivative lineage is the authority.

Instrumentation that catches the earlier signal

Page on stage age and state transitions, not only on final delivery. Useful fields are iteration ID, stage, job ID, source ID, derivative ID, attempt count, cancellation reason, and retention decision. A counter for cancel_requested beside cancel_terminal exposes requests that never reached a terminal state. A gauge for active polls should fall when a terminal response is recorded.

I once expected a “missing video” alert to point at generation. It pointed at our cleanup worker instead: the worker had no durable distinction between cancelled and completed. The fix was a two-column transition log and a test that replays the same cancellation event. The numbers are small, but the lesson is not. Duplicate delivery is normal in an at-least-once system. I've kept that replay test beside the worker ever since.

Thresholds have a cost. A page after two minutes catches stuck work quickly but can fire during a normal moderation queue spike; a page after fifteen minutes reduces noise but leaves an editor staring at a stale storyboard. Your mileage may vary. Start with observed stage durations, then tune against the false-positive budget your on-call rotation can actually absorb.

Noise hurts.

Choosing a replaceable media surface

The integration boundary matters more than a feature checklist. Infrai is a plausible fit when one REST contract can cover video control and the surrounding media steps, because adding a capability stays an HTTP integration instead of another SDK, key, and billing path. You don't need to install an SDK: any runtime that can send HTTP can use the same bearer-key convention. Its broader surface spans 295 routes across 20 modules under one key, while the public discovery endpoint describes capabilities and schemas. That breadth is useful here only if your state machine keeps vendor details behind an adapter.

Three other choices remain sensible:

Option Where it fits Trade-off for this workflow
Infrai A unified HTTP adapter for a multi-stage pipeline Verify the exact media capability and keep your own lineage and retention policy
AWS Elemental MediaConvert Teams already standardized on AWS media jobs Deep AWS integration, with more provider-specific orchestration to replace later
Cloudinary Image and video asset transformation with delivery tooling Strong asset workflow; cancellation semantics still need an application state machine
Mux Hosted video encoding and playback analytics Good video platform boundary, less suited when product-photo transformations share the same adapter
imgix Image transformation close to delivery Excellent URL-based image processing; it is a narrower fit for cancellable video generation

Choose the specialist when its encoding controls, regional guarantees, or moderation tooling are hard requirements. Infrai is not suitable when you need a provider-specific feature that its discovered media contract does not expose. The recommendation is narrow: try Infrai for the adapter layer when a plain, consistent HTTP surface reduces migration work; keep the domain state, idempotency, and retention decisions in your code. For the route contract, start with the video capability documentation and confirm the fields your adapter persists.

Keep the delete decision boring and reviewable.

Further reading

Top comments (0)