DEV Community

knoxblackwood2375
knoxblackwood2375

Posted on

How to Ship Go Video Delivery with Verified Download URL State

How to Ship Go Video Delivery with Verified Download URL State

A download center for generated clips has one operational constraint that changes the design: a URL must never imply that bytes are ready before the encoder has committed them. Short answer: model delivery as an explicit state machine, issue a short-lived URL only from ready, and make every transition observable and idempotent.

I learned this while reviewing a healthtech workflow that produced several aspect ratios for the same patient-education clip. The UI polled a job record, saw status: complete, and immediately rendered a download link. The encoder had finished its database update, but the object-store commit was still in flight. A handful of users received a link that returned an empty object. The HTTP request was valid; the ordering was not.

That distinction matters more than the choice of storage provider.

The incident lesson: completion is not the same as availability

The invariant is simple: publish a download URL only after the media object is durable and its metadata has passed validation. A worker can report progress, but only a verifier should grant the ready state. For generated video, that verifier should check the object length, checksum, MIME type, and the expected renditions (for example, 16:9, 1:1, and 9:16) before the download center can expose anything.

The state machine I use is deliberately boring:

queued -> rendering -> verifying -> ready

Failures go to failed with a reason that operators can act on. A retry returns to rendering; it never jumps from failed straight to ready. This gives the SLO a measurable boundary: the delivery SLO starts when verification succeeds, not when a renderer emits its last log line.

A status endpoint should return a stable representation, while the URL endpoint enforces the gate again. Clients can cache status for a few seconds, but authorization must not rely on a stale browser response.

How should status-gated download URLs work in Go?

Keep the policy in one handler so every client follows the same rule. The example below uses an in-memory repository to make the decision path executable; production code should put the same interface over a transactional database and object storage.

package main

import (
    "encoding/json"
    "net/http"
    "strings"
    "time"
)

type Clip struct {
    ID       string `json:"id"`
    State    string `json:"state"`
    Object   string `json:"object"`
    Checksum string `json:"checksum"`
    Bytes    int64  `json:"bytes"`
}

type ClipStore interface {
    Get(id string) (Clip, bool)
    SignedURL(object string, expiry time.Duration) (string, error)
}

func downloadURL(store ClipStore) http.HandlerFunc {
    return func(w http.ResponseWriter, r *http.Request) {
        id := strings.TrimPrefix(r.URL.Path, "/clips/")
        clip, ok := store.Get(id)
        if !ok {
            http.Error(w, "clip not found", http.StatusNotFound)
            return
        }
        if clip.State != "ready" || clip.Bytes == 0 || clip.Checksum == "" {
            http.Error(w, "clip is not ready", http.StatusConflict)
            return
        }
        url, err := store.SignedURL(clip.Object, 10*time.Minute)
        if err != nil {
            http.Error(w, "download unavailable", http.StatusServiceUnavailable)
            return
        }
        w.Header().Set("Content-Type", "application/json")
        json.NewEncoder(w).Encode(map[string]string{"url": url, "checksum": clip.Checksum})
    }
}
Enter fullscreen mode Exit fullscreen mode

The 409 is intentional: the caller can poll status and try again without treating a not-yet-ready clip as a permanent missing resource. The URL is scoped to one object, expires quickly, and is generated after the second state check. In a real service, bind the clip ID to the authenticated patient-organization context before this handler runs; a signed URL is not an authorization substitute.

Make verification and retries observable

A queue retry can create two render attempts, so the transition into verifying needs an idempotency key such as (clip_id, rendition, render_version). The verifier records the observed byte count and checksum in the same transaction that changes state. If the worker crashes after the object upload but before that transaction commits, the next verifier run inspects the existing object and safely completes the transition.

Track at least these measures: time from rendering to ready, percentage of jobs stuck in verifying, URL issuance denials by state, and download starts that fail before the first byte. Alert on an SLO burn rate, not on a single slow render. A useful starting budget is a 99.5% monthly success rate for ready clips, with a separate latency target for URL issuance; your mileage may vary because clinical review windows and clip duration change the workload.

Log identifiers, not patient data. Include clip ID, rendition, state transition, attempt number, and checksum algorithm. Do not put signed URLs in logs, traces, or analytics events.

Quality, bandwidth, and format decisions

Healthtech users may download on a clinic network, a phone, or an embedded portal. Quality versus bandwidth is therefore a product constraint with an SRE cost. Generate only the aspect ratios your UI actually requests, retain a mezzanine object for reprocessing, and attach an explicit content type such as video/mp4 to each rendition. The Media Formats guide explains why browser support differs across containers and codecs; test the target browsers instead of assuming that an extension describes the payload.

A small decision table keeps the trade-off visible:

Choice Helps Costs or boundary
Pre-render every ratio Predictable download latency More storage and encode minutes; wasteful for rarely used views
Render on demand Lower idle storage First-download latency and queue pressure; needs capacity headroom
Long-lived URLs Fewer refresh calls Wider exposure window and harder revocation
Ten-minute URLs Limits accidental sharing Clients must refresh after expiry

The catch is that this pattern is not suitable when users need byte-range playback from a public CDN with no session context; use a media delivery design that supports authenticated manifests and range requests, and keep the state gate at manifest publication. Stick with a simpler static export when clips are immutable, non-sensitive, and produced in a batch window where operators can verify the whole set before release.

Test the gate before you tune the encoder

Most regressions appear at the boundary, so test transitions rather than only happy-path rendering. A table-driven Go test can assert that rendering, verifying, zero-byte, and checksum-missing records never produce a URL, while a verified record does. Add a concurrency test that calls the URL handler during a transition and confirms that no response leaks the object key before ready.

Run a canary with synthetic clips whose durations and aspect ratios match production. Compare first-byte latency, rebuffer rate, and egress volume against the quality target. If bandwidth rises, lower bitrate or add a rendition; do not weaken the status gate.

Capacity planning belongs in the same runbook. Estimate peak concurrent encodes, multiply by the measured CPU minutes for the longest clip, and reserve queue capacity for retries rather than average traffic. During a rollout, cap new renditions behind a feature flag and watch the verifying backlog for two SLO windows. If it grows while object storage remains healthy, roll back the profile change; if verification latency is flat but egress spikes, keep the gate and revisit bitrate policy with product owners.

I am not sure a single global bitrate policy will survive every clinical program. That uncertainty is a reason to keep encoding profiles configurable and versioned, not a reason to hide the trade-off in a handler.

Sources

Top comments (0)