DEV Community

Cover image for I built a video editor that renders sharp 1080p MP4s entirely in the browser
Madalitso Nyemba
Madalitso Nyemba

Posted on

I built a video editor that renders sharp 1080p MP4s entirely in the browser

I build websites for clients. A few weeks ago I finished one I was really proud of and wanted to show it off on TikTok.

The options were not great. A raw screen recording looked flat. Mockup tools were built for static images. Video editors meant an hour of keyframes for a 20 second clip.

So I did what developers do. I opened an empty HTML file.

The first version: one file, one effect

The idea was simple. Play the desktop recording inside a laptop, morph the laptop into a phone, switch to the mobile recording, and end on a call to action.

The first version used DOM elements and CSS-style animation. A laptop frame was just a div with a dark border, the screen was a <video> element, and the morph animated width, height and border-radius from laptop proportions to phone proportions while crossfading the two videos.

It looked good. Then I exported it.

Problem 1: the video was blurry

To get a video out of a web page, my first approach was screen capture: getDisplayMedia, crop to the frame with Region Capture, and record with MediaRecorder.

The result was soft and smeary. The reason is obvious in hindsight: you can only capture what is on the screen. A 1080 x 1920 vertical frame on a normal laptop screen is displayed at roughly 600 x 1080 pixels. The recording is scaled back up for phones, and everything goes soft. Tab capture also compresses on the fly.

No amount of tweaking bitrate fixes missing pixels.

The fix: stop recording the screen

Instead of recording what the browser displays, I render every frame myself onto a full-size canvas and encode it directly.

That meant moving the whole animation from DOM to a <canvas> of exactly 1080 x 1920, and writing it as a pure function of time:

function draw(t, settings) {
  drawBackground(t);
  drawHook(t, settings);
  drawDevice(t, settings);
  drawCaptions(t, settings);
  drawCallToAction(t, settings);
}
Enter fullscreen mode Exit fullscreen mode

This one decision unlocked everything else. If draw(t) always produces the same frame for the same t, then preview and export are guaranteed to match. The preview runs it in requestAnimationFrame. The export runs it frame by frame, as fast or slow as the encoder needs.

Easing without a library

Since every value is calculated from t, I needed easing functions I could call directly. A small cubic-bezier solver does the job:

const cubicBezier = (x1, y1, x2, y2) => (t) => {
  if (t <= 0) return 0;
  if (t >= 1) return 1;
  let lo = 0, hi = 1, u = t;
  for (let i = 0; i < 24; i++) {
    u = (lo + hi) / 2;
    const x = 3 * (1 - u) ** 2 * u * x1 + 3 * (1 - u) * u ** 2 * x2 + u ** 3;
    if (x < t) lo = u; else hi = u;
  }
  return 3 * (1 - u) ** 2 * u * y1 + 3 * (1 - u) * u ** 2 * y2 + u ** 3;
};

const EASE = cubicBezier(0.65, 0, 0.35, 1);
const progress = (t, start, duration, ease) =>
  ease(Math.min(Math.max((t - start) / duration, 0), 1));
Enter fullscreen mode Exit fullscreen mode

Bisection instead of Newton's method is a little slower but never misbehaves, and at 24 iterations it is more precise than any screen can show.

Encoding with WebCodecs

Chrome and Edge ship VideoEncoder, which encodes H.264 in the browser, often with hardware acceleration. Each frame becomes a VideoFrame straight from the canvas:

const encoder = new VideoEncoder({
  output: (chunk, meta) => muxer.addVideoChunk(chunk, meta),
  error: (e) => console.error(e),
});

encoder.configure({
  codec: "avc1.640028", // H.264 High profile, level 4.0
  width: 1080,
  height: 1920,
  bitrate: 14_000_000,
  framerate: 30,
});

for (let i = 0; i < totalFrames; i++) {
  const t = i / fps;
  await seekVideosTo(t);
  draw(t, settings);

  const frame = new VideoFrame(canvas, {
    timestamp: Math.round((i * 1e6) / fps),
    duration: Math.round(1e6 / fps),
  });
  encoder.encode(frame, { keyFrame: i % (fps * 2) === 0 });
  frame.close();

  // back-pressure: don't flood the encoder
  while (encoder.encodeQueueSize > 8) {
    await new Promise((r) => setTimeout(r, 5));
  }
}

await encoder.flush();
Enter fullscreen mode Exit fullscreen mode

A few lessons from this part:

  • Check codec support first. VideoEncoder.isConfigSupported() tells you whether a codec string works on the user's machine. 60 fps at 1080 x 1920 needs a higher AVC level (avc1.64002A) than 30 fps does.
  • Always close frames. Forget frame.close() and memory climbs fast.
  • Respect back-pressure. Without the encodeQueueSize check, a fast canvas can queue hundreds of frames.

VideoEncoder only gives you encoded chunks, so you still need a muxer to wrap them in an MP4 file. I used mp4-muxer, which is tiny and did exactly what I needed. If you are starting today, its author now maintains a successor called Mediabunny that is worth a look.

Problem 2: the recordings inside the video

The reel contains the user's own screen recordings. For a frame-perfect export, each recording has to show exactly the right frame at time t, so I seek before drawing:

function seek(video, time) {
  return new Promise((resolve) => {
    const target = Math.min(time % video.duration, video.duration - 0.02);
    if (Math.abs(video.currentTime - target) < 0.0005) return resolve();
    video.addEventListener("seeked", resolve, { once: true });
    video.currentTime = target;
  });
}
Enter fullscreen mode Exit fullscreen mode

Seeking every frame is slower than playing, but it is deterministic, and that matters more than speed for an export.

Two gotchas cost me time:

  1. Cross-origin video taints the canvas. Once a tainted canvas is drawn, you cannot create a VideoFrame from it. All media is served from the same origin as the app.
  2. Seeking needs HTTP Range support. Without Range requests, the browser has to download the whole file to seek. The media route streams partial content with a proper 206 response.

Problem 3: the videos were cut off

Desktop recordings are wide (mine was 1908 x 868) and phone recordings are tall. The first version used fixed device sizes, so recordings were cropped.

The fix: the device screen takes the aspect ratio of the recording itself, clamped to a sensible size, and the video is drawn with "contain" logic. Even the morph animates from the laptop's real proportions to the phone's real proportions.

From one file to a product

People who saw the reel asked how I made it, so the experiment became a product: ByteMorph.

The core rule I kept from the prototype is that a reel is data. Every reel is a JSON spec: a list of scenes with their text, recordings, timing, transitions and styles. The renderer just draws whatever the spec describes.

That one decision made almost every later feature simple:

  • Templates are saved specs.
  • Dragging elements on the canvas edits positions in the spec.
  • Four formats (9:16, 1:1, 4:5, 16:9) come from a layout module that computes everything relative to the canvas size.
  • The upcoming AI assistant will edit the same spec, so it cannot break anything the renderer does not understand. The server validates every spec anyway.

Adding 3D

The 2D canvas cannot do a convincing laptop that spins while its lid opens, so 3D scenes use Three.js. It renders into its own WebGL canvas, and each frame is copied onto the main canvas between the 2D background and the 2D overlays. The export pipeline did not change at all.

The same rule applies: every 3D value comes from t, never from a Three.js clock. Animations are stored as presets (like "lid open, 1.2 seconds, ease out") that expand into keyframes, so the preset and a user's customised version always run through the same code. And Three.js only loads when a reel actually uses 3D.

Feature callouts and zoom

For showing off web apps and dashboards, two features turned out to matter most:

  • Callouts: feature lists, pins and counters as a layer on any scene. Pins store their target in the recording's own coordinates, so they stay locked to the right spot even while the device moves, morphs or zooms.
  • Zoom to region: the camera glides into part of the recording so a dashboard is readable on a phone. The zoomed frame is drawn from the full-resolution video, not an upscaled copy.

The rest of the stack

  • Laravel with Inertia for the app, Filament for the admin panel and template management
  • A VitePress docs site with screenshots and clips regenerated by a Playwright script

Try it

ByteMorph is live today at bytemorph.app. It is free to start, and every new account gets 7 days of Pro. We are also launching on Product Hunt today: here.

I would love feedback from other developers, especially on the rendering side. If you have worked with WebCodecs, I am curious what you have run into.

Top comments (2)

Collapse
 
launchgatecheck profile image
Launch Gate •

I'd test the input-frame boundary separately from draw(t): use a recording with a burned-in frame number, export across a loop boundary and the laptop-to-phone transition, then compare those frames with the preview at the same timestamps. That checks which decoded video frame reached the canvas, not only whether the animation math agrees.

What does seekVideosTo do when seeking fails or a source never becomes ready? A bounded failure with the source and timestamp would be easier to act on than an export that hangs halfway through. I haven't run ByteMorph; these are regression suggestions from the pipeline you describe.

Collapse
 
madalitsonyemba profile image
Madalitso Nyemba •

Really good catch, thanks. I had been treating the decoded frames as a given, and the morph crossfade is exactly where that could quietly go wrong. Burned-in frame numbers are going into my test suite, and the seek hang is getting a proper timeout and error message. Surely appreciate the close read.