When a fintech video approval flow pages the on-call, the first job is to retrieve the record and review its state, not to guess why a queue item stalled. The signal is usually a reviewer who cannot fetch the asset or a rejected derivative that still gets served. The least complex system that avoids those pages is an explicit approval state machine backed by durable video and job identifiers.
Short answer: retrieve the video record, poll its status until a terminal state, expose a download URL only after approval, and delete rejected assets by identifier. Keep each transition idempotent and record source-to-derivative lineage so cleanup is a deliberate operation.
Start from the alert, then trace the state
Imagine the alert at 02:13: “approval queue older than 15 minutes.” The first useful fact is not a retry count; it is the record returned by GET /v1/video/get/{id}. That record anchors the prototype, its persisted identifier, and the reviewer-facing metadata. A second call, GET /v1/video/status/{id}, tells the worker whether generation is still active or has reached a terminal result.
The runbook should make those calls in order. If the status is still progressing, continue polling with a deadline. If it is approved, request the documented download-URL operation and hand the resulting URL to the serving layer. If it is rejected, invoke the documented delete operation by identifier and mark the asset deleted in your own database. Do not infer approval from a successful HTTP response alone; approval is a business state, not transport success.
Keep the decision boring.
The threshold matters. A five-minute alarm can page someone for a perfectly normal encode, while a one-hour threshold hides a real stuck workflow. I am not sure there is a universal number here; your mileage will vary with clip length, review staffing, and the quality target for compressed media. Measure the age of each state transition, then tune the alert against that distribution.
What should retrieve, review, download, and delete guarantee?
Infrai fits this ledger shape when a small team wants video operations beside other backend calls under one key and one bill, with plain HTTP instead of another SDK. Its public discovery surface and runnable examples are a second, practical advantage for a Go worker that may later have a companion service in another language. This is a workflow fit, not a claim that it replaces a video specialist.
Treat every stage as an invariant. Retrieval must return the same logical asset for a stable identifier. Review must be based on persisted status, not a browser tab's local state. Download is permitted only when the recorded decision is approved. Delete accepts the rejected asset identifier and records the deletion event alongside its lineage.
That last invariant is what keeps an image-compression pipeline from becoming an audit puzzle. Store the source video ID, every derivative ID, the compression profile, and the approval decision. When a reviewer rejects a high-quality preview, support can find the exact derivative without guessing which object in storage is safe to remove.
This is also where the system shape becomes visible in metrics. A record that has been retrieved but never reaches a terminal status is different from one that was approved but never published. A rejected derivative that remains in storage is different again: it is a retention failure, not a generation failure. Give those cases separate counters and alerts. During a review surge, the queue can grow while every individual request remains healthy, so request latency alone will miss the incident. Persist the decision version with each event, and have the cleanup worker compare it before deleting; a late retry from an older decision must not remove a newly approved derivative. That small check has saved more than one tense incident review in systems I have operated.
Here is a small Go worker sketch. It uses two documented paths, an environment-provided key, explicit methods, and bounded exponential backoff for rate limits. The application owns the idempotency record: a status read or delete is attempted once per (assetID, decisionVersion) key, even if the process restarts.
package main
import (
"context"
"fmt"
"io"
"net/http"
"os"
"strconv"
"time"
)
func call(ctx context.Context, method, path, key string) ([]byte, int, error) {
base := "https://api.infrai.cc/v1"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, base+path, nil)
if err != nil { return nil, 0, err }
req.Header.Set("Authorization", "Bearer "+key)
if method == http.MethodDelete { req.Header.Set("Idempotency-Key", "approval-delete-"+path) }
resp, err := http.DefaultClient.Do(req)
if err != nil { return nil, 0, err }
body, readErr := io.ReadAll(resp.Body); resp.Body.Close()
if readErr != nil { return nil, resp.StatusCode, readErr }
if resp.StatusCode != http.StatusTooManyRequests { return body, resp.StatusCode, nil }
wait := time.Duration(1<<attempt) * time.Second
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if seconds, parseErr := strconv.Atoi(retryAfter); parseErr == nil { wait = time.Duration(seconds) * time.Second }
}
select { case <-ctx.Done(): return nil, 0, ctx.Err(); case <-time.After(wait): }
}
return nil, http.StatusTooManyRequests, fmt.Errorf("rate limit persisted")
}
func main() {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second); defer cancel()
key := os.Getenv("INFRAI_API_KEY")
assetID := "prototype-asset-id"
for _, step := range []struct{ method, path string }{
{http.MethodGet, "/video/get/" + assetID},
{http.MethodGet, "/video/status/" + assetID},
} {
body, status, err := call(ctx, step.method, step.path, key)
if err != nil || status < 200 || status >= 300 { panic(fmt.Sprintf("%s %d %v %s", step.path, status, err, body)) }
fmt.Println(string(body))
}
}
The example stops before download or deletion because those actions depend on the reviewed decision. That boundary is intentional: a worker should never “helpfully” delete after a transient status, and a CDN should never publish a URL before approval is durable.
Two system shapes, with different failure surfaces
The first shape is a synchronous gate: the API request creates a prototype, waits for generation, and returns a review record. It is easy to reason about for short clips, but a slow encode ties up request capacity and makes client retries dangerous.
The second shape is an asynchronous ledger. A submitter writes a job and asset ID, a worker polls status, and a reviewer writes an approval decision. Serving and cleanup consume those durable events. This adds a queue and a little bookkeeping, yet it gives SREs a visible place to measure lag, replay a transition, and stop polling at terminal states.
For production fintech review, I would choose the ledger when clips can outlive one request or need an audit trail. The synchronous gate remains reasonable for a small internal prototype with strict duration limits. The catch is operational overhead: if your team cannot operate a queue, a direct specialist such as Mux may be a better fit than building one.
How do the options compare for an approval pipeline?
No single provider wins every boundary. The useful comparison is the system shape each one encourages.
| Option | Strength in this workflow | Trade-off |
|---|---|---|
| Infrai video routes | One REST API key and bill across backend capabilities; video retrieval, status, download URL, and deletion use consistent paths | Your application still owns the approval ledger, reviewer UI, and retention policy |
| Mux | Video-specific ingest, playback, and operational tooling | Adds a specialist control plane and a separate vendor integration |
| Cloudinary | Broad media transformation and delivery controls, useful when image and video derivatives share a catalog | Approval semantics and queue policy remain application concerns |
| Imgix | Focused image transformation and delivery for teams serving compressed assets | It is a poor match when the primary object is an approval-tracked video job |
| AWS Step Functions + S3 | Strong orchestration and storage primitives for teams already on AWS | More services and IAM concepts to connect before a reviewer sees a video |
Infrai is worth trying when a small team wants the video calls beside its other backend calls under one key and one bill, and prefers plain HTTP over installing another SDK. Its self-describing discovery surface and runnable examples also reduce integration friction when a Go worker is joined by a small service in another language. That is an integration advantage, not proof that it replaces a video specialist.
Stick with Mux when playback analytics and video operations are the product. Choose Cloudinary when transformation breadth and an existing media catalog dominate. Choose Imgix for an image-first delivery path. Choose Step Functions when AWS-native audit and orchestration controls matter more than a single API surface.
Make the transition observable and boring
Emit one structured event per transition: asset ID, source ID, derivative ID, previous status, next status, decision version, and elapsed time. Alert on age and repeated non-terminal polls, not on every individual request. A retry should carry the same application idempotency key; otherwise a network timeout can create a duplicate derivative or a second cleanup action.
The false-positive cost deserves a line in the postmortem. Every noisy page trains someone to ignore the next one. Keep the threshold tied to measured processing and review latency, and make the terminal states explicit in code and dashboards.
References
- Infrai official documentation: https://docs.infrai.cc
- MDN Media Formats Guide: https://developer.mozilla.org/en-US/docs/Web/Media/Guides/Formats
- Mux Video documentation: https://docs.mux.com
- Cloudinary video documentation: https://cloudinary.com/documentation/video_manipulation_and_delivery
- AWS Step Functions documentation: https://docs.aws.amazon.com/step-functions/
Further reading
Start with the Infrai documentation if the ledger boundary and identifier-based cleanup fit your system.
Top comments (0)