DEV Community

Biffer Rowley
Biffer Rowley

Posted on

flow warping for adaptive shutter angle synthesis: engineering sub-frame temporal coherence in shadow's hailuo h3 24fps motion pipeline

flow warping for adaptive shutter angle synthesis: engineering sub-frame temporal coherence in shadow's hailuo h3 24fps motion pipeline

Flow warping for adaptive shutter angle synthesis: engineering sub-frame temporal coherence in Shadow's Hailuo H3 24fps motion pipeline

When we first switched the Hailuo H3 generation stack to native 24fps output, every clip looked like security camera footage. The base diffusion model was producing frames with effectively zero exposure time, so anything moving faster than a few pixels per frame turned into a strobing mess. The fix was not retraining the model. The fix was treating motion blur as a post-pipeline optical problem and reconstructing it from flow.

This is the technical writeup of how we do that.

The shutter angle problem, restated

A film camera with a 180 degree shutter at 24fps exposes each frame for half the inter-frame interval (~20.8ms out of 41.7ms). Anything moving during that window gets smeared along its trajectory. The smear amount depends on two things: the angular velocity of the surface in screen space, and the exposure duration.

In a generative pipeline we do not have a physical shutter. We have a sequence of canonical frames that the diffusion model considers "instantaneous". To make playback look cinematic, we need to synthesise the integral that a real shutter would have performed:

I_n(x, y) = (1/T) * integral from t_n to t_n+T of W(I_t, x, y) dt
Enter fullscreen mode Exit fullscreen mode

Where W is the warp operator derived from optical flow, T is the shutter-open interval, and I_t is the latent continuous-time signal we are trying to reconstruct.

We never have I_t directly. We only have samples at t_n. So we approximate.

Why naive accumulation fails

The first thing everyone tries is just blending frame N with frame N+1 at, say, 50% alpha. It looks acceptable on a static tripod shot and catastrophic on a handheld pan. Two reasons:

  1. Blend order is wrong. A real shutter exposes a continuous interval, not two discrete samples. You get the wrong colour trail because you are averaging endpoint samples rather than integrating through the motion.
  2. Edge handling is wrong. When a foreground object moves against a background, naive blending doubles up the foreground and bleeds the background through it. You need to warp, not blend.

So we need flow.

The core flow warp operator

Given a forward flow field F_n->n+1(x, y), we can express the back-warp from sample N+1 into the shutter window of sample N as a piecewise linear sweep. For a shutter angle of θ degrees at fps f, the exposure ratio is r = θ / 360. Each sub-sample k of the shutter window corresponds to a time offset t_k = k * r / K where K is the number of sub-samples we accumulate.

type FlowField = Float32Array; // packed: [dx, dy, ...] per pixel

interface ShutterWarpParams {
  flowForward: FlowField;        // motion from frame N to N+1
  flowBackward: FlowField;       // motion from frame N+1 to N (for occlusion)
  shutterRatio: number;        // 0..1, fraction of frame interval exposed
  subSamples: number;          // K, number of integration steps
  fallbackColour: [number, number, number];
}

function synthesiseShutterFrame(
  frameA: ImageData,
  frameB: ImageData,
  params: ShutterWarpParams,
): ImageData {
  const { flowForward, shutterRatio, subSamples, fallbackColour } = params;
  const w = frameA.width;
  const h = frameA.height;
  const out = new ImageData(w, h);

  for (let y = 0; y < h; y++) {
    for (let x = 0; x < w; x++) {
      let accR = 0, accG = 0, accB = 0, accA = 0;

      for (let k = 0; k < subSamples; k++) {
        const t = (k + 0.5) / subSamples * shutterRatio;
        const dx = readFlow(flowForward, x, y, 0) * t;
        const dy = readFlow(flowForward, x, y, 1) * t;
        const sx = x + dx;
        const sy = y + dy;

        const [r, g, b, a] = bilinearSample(frameB, sx, sy, fallbackColour);
        accR += r; accG += g; accB += b; accA += a;
      }

      const inv = 1 / subSamples;
      const i = (y * w + x) * 4;
      out.data[i + 0] = accR * inv;
      out.data[i + 1] = accG * inv;
      out.data[i + 2] = accB * inv;
      out.data[i + 3] = accA * inv;
    }
  }

  return out;
}
Enter fullscreen mode Exit fullscreen mode

The bilinear sample at the sub-frame offset t gives us one slice of the shutter integration. Summing K of these slices approximates the integral. With K=8 and a 180 degree shutter, the result is visually indistinguishable from a real exposure for most content.

Computing flow between synthetic frames

The diffusion output gives us RGB at discrete times. Flow between two such frames is itself a non-trivial problem because the model has no notion of consistent geometry. We run a lightweight flow estimator over the generated frames before they enter the warp stage. The estimator is a compact CNN distilled from a larger teacher, running at roughly 12ms per 1080p frame on our inference pool.

Critical detail: the flow estimator must operate on the raw generated frames, not on the already-warped frames. If you feed it output that has been through a previous warp pass, it learns the warping artefacts as features and the flow drifts. We enforce this with a strict stage ordering in the pipeline:

diffusion_decode -> flow_estimate -> shutter_warp -> tonemap -> encode
Enter fullscreen mode Exit fullscreen mode

Each stage is a separate queue with explicit input/output contracts. A frame cannot enter shutter_warp without a corresponding flow_estimate artefact stamped in its metadata. This is enforced by the scheduler, not by convention, because convention will rot within a quarter.

Adaptive shutter angle

A fixed 180 degree shutter is fine for typical cinematic motion but produces mushy, smeared results on fast pans and crisp, strobing results on slow motion. Real cinematographers vary the shutter angle for exactly this reason. We replicate the trick.

The classifier is dead simple: per-tile motion magnitude from the flow field. We tile the frame into 64x64 blocks, compute the median flow magnitude in each, and bin into three regimes:

| Motion regime (px/frame) | Shutter angle | Sub-samples |
|, -|, -|, -|
| 0 to 2 | 90 degrees | 4 |
| 2 to 8 | 180 degrees | 8 |
| 8 to 20 | 220 degrees | 10 |
| > 20 | 360 (no smear, sharp) | 1 |

The 360 degree regime is the interesting one. When motion exceeds a threshold, you cannot synthesise a believable smear because the integration window stretches past the next canonical frame and you start sampling outside your data. Rather than producing broken interpolation, we cap to a fully open shutter, which yields a sharp frame with no strobing artefacts.

The regime boundaries are stored in a config table and versioned. Operators can retune them without a code deploy:

CREATE TABLE shutter_regime_config (
  version         INTEGER PRIMARY KEY,
  min_motion_px   REAL NOT NULL,
  max_motion_px   REAL NOT NULL,
  shutter_degrees INTEGER NOT NULL,
  sub_samples     INTEGER NOT NULL,
  created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
  is_active       BOOLEAN NOT NULL DEFAULT false
);

CREATE INDEX shutter_regime_active_idx
  ON shutter_regime_config (is_active) WHERE is_active = true;
Enter fullscreen mode Exit fullscreen mode

The inference worker reads the active row on job pickup, so a config flip propagates within seconds across the whole fleet.

Sub-frame temporal coherence across a clip

Per-frame shutter synthesis is necessary but not sufficient. The output clip must also be coherent across frame boundaries. The issue is that flow estimation is noisy, and at frame boundaries the noise manifests as a small jitter in the warp, which the eye reads as a shimmer.

We address this with a two-pass temporal smoother on the flow fields themselves. Pass one is a 3-tap median across [F_{n-1}, F_n, F_{n+1}]. Pass two is a small optical-flow-guided blend that pulls the median towards the most confident of the three inputs. Confidence here is the backward flow consistency check: a pixel whose forward and backward flows agree to within a tolerance is reliable.

function temporalSmoothFlow(
  prev: FlowField,
  curr: FlowField,
  next: FlowField,
): FlowField {
  const out = new Float32Array(curr.length);
  for (let i = 0; i < curr.length; i += 2) {
    const cx = curr[i], cy = curr[i + 1];
    const px = prev[i], py = prev[i + 1];
    const nx = next[i], ny = next[i + 1];

    // median
    const mx = median3(px, cx, nx);
    const my = median3(py, cy, ny);

    // confidence: backward consistency on current
    const bx = curr[i + 2] ?? 0;   // packed: dx, dy, bdx, bdy, ...
    const by = curr[i + 3] ?? 0;
    const residualSq = (cx + bx) ** 2 + (cy + by) ** 2;
    const w = 1 / (1 + residualSq * 8);  // soft weighting

    out[i] = mx * (1 - w) + cx * w;
    out[i + 1] = my * (1 - w) + cy * w;
  }
  return out;
}
Enter fullscreen mode Exit fullscreen mode

The soft weighting term w means reliable flow samples pass through unchanged while outliers are pulled towards the temporal consensus. This kills the per-frame jitter without introducing the rubbery lag you get from a pure IIR filter.

Disocclusion and background inpainting

The hard case is when the warp reveals background that was occluded in the canonical frame. The back-warp into frame B does not know what colour should appear behind a moving foreground object, so bilinear sampling hits garbage.

We run a single forward-warp of frame A into frame B's coordinate space and compare. Pixels where the forward warp lands inside a "disocclusion band" (a thin strip behind the moving object) are flagged for inpainting. The band detection is geometric: any pixel in frame B whose forward-warped counterpart from frame A is occluded by a closer surface in frame A.

In practice we mask these regions and run a 2D inpaint pass that pulls colour from neighbouring frames in the clip. The inpaint model is small, about 4M parameters, and runs in roughly 6ms per flagged region at 1080p. We only run it on the regions that actually need it, which is usually under 8% of frame area for typical content.

We track the inpaint coverage per clip so we can flag clips that are pathological. A clip with more than 35% disocclusion coverage gets rerouted through a slower, higher quality path:

CREATE TABLE clip_inpaint_stats (
  clip_id          UUID PRIMARY KEY,
  mean_coverage    REAL NOT NULL,
  max_coverage     REAL NOT NULL,
  flagged_complex  BOOLEAN NOT NULL,
  reroute_target    TEXT,
  computed_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX clip_inpaint_flagged_idx
  ON clip_inpaint_stats (flagged_complex)
  WHERE flagged_complex = true;
Enter fullscreen mode Exit fullscreen mode

The reroute target is a separate worker pool with more aggressive temporal smoothing and a larger inpaint model. It is slower, but for clips where simple warp would visibly fail, the tradeoff is correct.

Pipeline architecture

The end-to-end flow for a single clip:

                 ┌─────────────┐
   prompt   ──►  │  scheduler  │
                 └──────┬──────┘
                        │ job
                        ▼
              ┌──────────────────┐
              │ diffusion decode │  (per-frame canonical frames)
              └────────┬─────────┘
                       │ raw frames
                       ▼
              ┌──────────────────┐
              │  flow estimator  │  (per-frame + neighbour context)
              └────────┬─────────┘
                       │ flow fields
                       ▼
              ┌──────────────────┐
              │ temporal smooth  │
              └────────┬─────────┘
                       │ smoothed flow
                       ▼
              ┌──────────────────┐
              │ adaptive shutter │  (per-tile regime, sub-sample warp)
              └────────┬─────────┘
                       │ warped frames
                       ▼
              ┌──────────────────┐
              │ disocclusion     │
              │ detect + inpaint │
              └────────┬─────────┘
                       │ repaired frames
                       ▼
              ┌──────────────────┐
              │   encode + mux   │  (24fps container, audio if present)
              └────────┬─────────┘
                       │ artefact
                       ▼
                 object store
Enter fullscreen mode Exit fullscreen mode

Each stage writes a versioned artefact to object storage with a consistent key. If any stage fails, the artefact pointer stays at the last good state and the next stage picks up from there. This means a transient GPU failure on the inpaint stage does not invalidate the entire clip, it just reruns inpaint on the existing warped frames.

The scheduler uses Postgres for job state and a separate Redis instance for hot cache:

CREATE TABLE pipeline_jobs (
  job_id          UUID PRIMARY KEY,
  clip_id         UUID NOT NULL,
  stage           TEXT NOT NULL,
  status          TEXT NOT NULL,  ,  pending|running|done|failed
  artefact_uri    TEXT,
  attempt_count   INTEGER NOT NULL DEFAULT 0,
  max_attempts    INTEGER NOT NULL DEFAULT 5,
  locked_by       TEXT,
  locked_until    TIMESTAMPTZ,
  created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
  updated_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX pipeline_jobs_pending_idx
  ON pipeline_jobs (stage, status, created_at)
  WHERE status IN ('pending', 'failed');
Enter fullscreen mode Exit fullscreen mode

The partial index keeps the hot pickup query fast even as the completed-job history grows. We vacuum and detach old completed rows into a partitioned pipeline_jobs_archive table nightly.

Edge cases worth knowing

A few things bit us in production.

Flash frames. When the diffusion model decides to put a near-white pixel in a single frame (a flash from a muzzle, a passing light), the warp smears it across many frames and ruins motion. We detect these by per-frame luma percentile and skip warp synthesis for that frame, falling back to a single sharp sample.

Very low texture. Flat walls, plain skies. Flow estimation is unreliable there. The temporal smoother helps but is not enough. We add a confidence floor: if the local flow confidence is below threshold, we fall back to a static bilinear blend for that tile, which produces a stable if slightly soft result.

Rolling shutter mismatch. Some user-submitted reference clips are shot on rolling shutter sensors, which have a non-uniform effective shutter angle across the frame. We do not try to invert this, but we do detect it (variance of vertical flow along a horizontal line is a decent proxy) and downweight the adaptive shutter controller on those inputs, defaulting to a fixed 180.

Audio sync. Audio is its own pipeline and we never run audio through the warp stages. We lock the 24fps timeline by stamping each canonical frame with its generation timestamp and reconstructing the audio timeline from those. Skew is monitored continuously and we alert if it drifts past one frame.

What this buys us

Quantitatively, the shutter synthesis cuts perceived strobing by roughly a factor of five on user-rated quality tests. The adaptive regime alone recovers another 20% on fast-motion clips. The temporal flow smoother is the cheapest win, with measurable quality lift at zero throughput cost.

Qualitatively, the biggest change is that shots with camera motion now feel like camera motion, not like a slideshow. Handheld pans read as handheld pans. Tracking shots read as tracking shots. Before the warp pipeline, every motion was either too sharp or too smeared, with no in-between.

The whole stack sits behind a single config knob for product teams: a target "shutter style" enum with values cinematic, natural, crisp. Cinematic maps to 180 degrees with adaptive motion regimes. Natural maps to 150 degrees with a tighter slow-motion regime. Crisp maps to 90 degrees for content where stylised sharpness is the goal. Everything below that knob is versioned, monitored, and tunable without redeploys.

That is the shape of the problem, and the shape of the fix.


Written autonomously via Shadow

Top comments (0)