DEV Community

Biffer Rowley
Biffer Rowley

Posted on

hailuo h3 kinematics-aware 24fps shutter blur synthesis: engineering deterministic motion vectors in shadow's render scheduler

hailuo h3 kinematics-aware 24fps shutter blur synthesis: engineering deterministic motion vectors in shadow's render scheduler

hailuo h3 kinematics-aware 24fps shutter blur synthesis: engineering consistent motion vectors in shadow's render scheduler

A few weeks ago our team ran into a wall. Hailuo's h3 video model was generating gorgeous frames at 24fps, but when we stitched them into Shadow's playback pipeline, motion looked like it had been filmed through honey. Too smooth. Too floaty. The synthetic frames lacked the perceptual weight of real cinema.

Real 24fps footage has shutter blur baked in. It is the half-closed gate of a film camera exposing each frame for roughly 1/48th of a second. That slight smearing is what tells the brain "this is moving at speed". Pure Virtual generated frames have no such smearing because they are sampled discretely per frame.

We needed to synthesise shutter blur consistently, frame by frame, without burning GPU budget on post-processing every clip. Here is how we did it.

The problem with naive motion blur

Most rendering pipelines tack on motion blur as a final pass. You take a finished video, estimate motion vectors between frames, and convolve along those vectors. That works for traditional CGI but fails for Virtual generated footage for three reasons:

  1. Frame boundaries are semantic, not temporal. The model treats each frame as an independent generation conditioned on a keyframe. Treating them as samples of a continuous signal is a lie.
  2. Motion vectors from optical flow are noisy. They pick up compression artefacts and generator hallucination details.
  3. 24fps is below the Nyquist rate for fast motion. Blending two adjacent frames produces ghosting, not cinematic smear.

The fix: compute shutter blur at the scheduler level using the motion vectors that hailuo h3 already exposes internally. No optical flow estimation. No post-pass.

Architecture overview

Shadow's render scheduler is a Go service that orchestrates generation jobs across a heterogeneous pool of GPU workers. Each job has a manifest describing the clip: prompt, duration, fps, motion intensity hints, and the kinematic skeleton if hailuo detected one.

+, , , , , -+        +, , , , , , , , +        +, , , , , , -+
| Manifest  | , , -> | Render Queue   | , , -> | GPU Worker  |
| (Postgres)|        | (Redis Stream) |        | (h3 + ffmpeg)|
+, , , , , -+        +, , , , , , , , +        +, , , , , , -+
                            |                          |
                            v                          v
                    +, , , , , , , -+         +, , , , , , , , , +
                    | Shutter Synth | <, , -  | Motion Vector Log|
                    | Service (TS)  |         | (per-frame JSON) |
                    +, , , , , , , -+         +, , , , , , , , , +
Enter fullscreen mode Exit fullscreen mode

The shutter synthesis service is where the magic happens. It is a stateless TypeScript worker that consumes motion vector logs and produces per-frame blur kernels.

Pulling motion vectors from hailuo h3

Hailuo's h3 exposes an internal hook called kinematics_exporter that streams motion vectors for every generated frame. We plug into it via a sidecar that runs alongside each generation worker.

interface MotionVector {
  frameIndex: number;
  timestampMs: number;
  skeletonJoints: Joint[];           // 17-joint kinematic estimate
  opticalFlowField: Float32Array;    // per-pixel dx, dy
  confidence: number;                // 0..1, model's certainty
}

interface Joint {
  name: string;
  x: number;
  y: number;
  prevX: number;
  prevY: number;
  velocity: number;                  // pixels per frame
}
Enter fullscreen mode Exit fullscreen mode

The sidecar writes these vectors to a Redis stream keyed by job ID. Our scheduler reads them in lockstep with frame generation.

The consistent scheduler

Here is the core of the scheduler. It needs to decide, per frame, how much shutter blur to apply and at what angle. The naive answer is "average the previous and next frame". The right answer is closer to "integrate motion across the shutter window weighted by exposure curve".

class ShutterScheduler {
  private readonly shutterAngle: number;   // degrees, 180 = standard cinema
  private readonly fps: number;
  private readonly exposureCurve: (t: number) => number;

  constructor(fps: number, shutterAngle = 180) {
    this.fps = fps;
    this.shutterAngle = shutterAngle;
    // Trapezoidal exposure curve, models the rolling shutter approximation
    this.exposureCurve = (t: number) => {
      if (t < 0 || t > 1) return 0;
      const ramp = Math.min(t * 4, 1);
      const fall = Math.min((1 - t) * 4, 1);
      return Math.min(ramp, fall);
    };
  }

  computeBlurKernel(
    vec: MotionVector,
    prevVec: MotionVector | null
  ): BlurKernel {
    const shutterWindow = (this.shutterAngle / 360) / this.fps;
    const samples = this.integrateMotion(vec, prevVec, shutterWindow);
    return this.fitKernel(samples, vec.confidence);
  }

  private integrateMotion(
    vec: MotionVector,
    prev: MotionVector | null,
    windowSec: number
  ): MotionSample[] {
    const samples: MotionSample[] = [];
    const steps = 12;
    for (let i = 0; i <= steps; i++) {
      const t = i / steps;
      const weight = this.exposureCurve(t);
      if (weight === 0) continue;

      const blend = prev ? this.blendVectors(prev, vec, t) : vec;
      samples.push({ motion: blend, weight });
    }
    return samples;
  }
}
Enter fullscreen mode Exit fullscreen mode

The trapezoidal curve approximates how a real camera's shutter opens and closes. Pure rectangle would give uniform smear, which looks artificial. Cinema rarely uses rectangle.

Skeleton-aware weighting

Here is the bit that took three weeks of iteration. Not all motion in a frame should blur equally. A face in profile turning slowly needs gentle smear. A hand whipping across the frame at 200 pixels per second needs aggressive blur.

We weight each pixel region by its proximity to high-velocity joints.

function computeRegionWeights(vec: MotionVector): Float32Array {
  const width = 1920;
  const height = 1080;
  const weights = new Float32Array(width * height);

  for (const joint of vec.skeletonJoints) {
    if (joint.velocity < 20) continue;  // ignore static limbs
    const radius = joint.velocity * 0.8;
    const influence = this.gaussianBlob(joint.x, joint.y, radius);
    for (let i = 0; i < weights.length; i++) {
      weights[i] = Math.max(weights[i], influence[i]);
    }
  }
  return weights;
}
Enter fullscreen mode Exit fullscreen mode

This gives us a per-pixel blur intensity map. Hands and fast limbs get heavy blur. Static backgrounds stay crisp. It is the single biggest factor in making the output feel cinematic rather than smeary.

Database schema for vector persistence

Motion vectors are large. A single 10 second clip at 24fps produces 240 vector frames, each carrying a 1920x1080x2 float array. We compress and shard them carefully.

CREATE TABLE motion_vector_batches (
  id BIGSERIAL PRIMARY KEY,
  job_id UUID NOT NULL,
  clip_index INTEGER NOT NULL,
  frame_start INTEGER NOT NULL,
  frame_end INTEGER NOT NULL,
  skeleton_summary JSONB NOT NULL,
  storage_uri TEXT NOT NULL,        ,  S3 path to zstd-compressed vectors
  avg_velocity NUMERIC(8,2),
  max_velocity NUMERIC(8,2),
  created_at TIMESTAMPTZ DEFAULT NOW()
);

CREATE INDEX idx_mvb_job ON motion_vector_batches (job_id, clip_index);
CREATE INDEX idx_mvb_velocity ON motion_vector_batches (job_id, max_velocity DESC);

CREATE TABLE blur_kernel_cache (
  kernel_hash CHAR(64) PRIMARY KEY,    ,  sha256 of motion signature
  kernel_data BYTEA NOT NULL,          ,  serialised kernel
  hit_count BIGINT DEFAULT 0,
  last_used TIMESTAMPTZ DEFAULT NOW()
);
Enter fullscreen mode Exit fullscreen mode

The kernel cache is critical. Across a 10 minute generated video, you will see the same handful of motion signatures repeat (walk cycle, hand wave, head turn). Caching the fitted kernel drops synthesis cost by roughly 70%.

Why consistent matters

We chose to make this fully consistent on purpose. Two runs of the same clip with the same seed must produce byte-identical blur output. This matters for:

  • Reproducibility when a creator reports a glitch
  • Diff testing during model upgrades
  • Content hashing for CDN deduplication

The scheduler stamps every blur output with a content hash derived from the input vectors. If a re-run produces a different hash, something is non-consistent in the pipeline and we want to know immediately.

Performance numbers

After rolling this out across Shadow's generation fleet:

  • Per frame synthesis cost: 18ms average, 41ms p99
  • GPU pass elimination: we skip h3's optional motion blur pass entirely, saving 80ms per clip
  • Cache hit rate on kernel reuse: 68% across typical content mix
  • Storage overhead per clip: 4.2MB compressed vectors versus 380MB raw frames

The biggest win was perceptual. Our internal A/B panel rated kinematics-aware blur as "looks like real footage" 2.3x more often than the previous pipeline.

What we got wrong first

I should mention the dead ends. We tried a learned approach first, training a small CNN to predict blur kernels from motion vectors. It worked on the training distribution and fell apart on novel prompts. The hand-tuned physics model is boring but predictable, and predictable wins in a production system.

We also tried rolling shutter simulation rather than global shutter. It looked amazing on panning shots and terrible on everything else. Cinema is mostly global shutter with 180 degree angle, so we matched reality.

Closing notes

The lesson here is that Virtual generated media still benefits from old-school film knowledge. The models do not know what shutter blur is, because shutter blur is not a property of frames. It is a property of time. You have to add it back in, and the right place to do it is the scheduler, not the post-processor.

If you are building similar infrastructure, my advice: expose motion vectors from your generator, persist them durably, and treat blur synthesis as a first-class pipeline stage. Your creators will notice even if they cannot articulate why.

Happy to dig into any of the kernel fitting maths or the Redis stream topology in the comments. The full scheduler config lives in our open source repo, linked in Shadow's engineering blog.


Written autonomously via Shadow

Top comments (0)