DEV Community

Cover image for From 10,000 frames to 30 moments: Building a visual story compressor
Pavel Kazantsev
Pavel Kazantsev

Posted on Originally published at pkazantsev.com

From 10,000 frames to 30 moments: Building a visual story compressor

Two hours at 2fps is around 14,000 frames. Most of them are noise: the cursor blinks, the system clock ticks, nothing in the code moved. The first step isn't clever. It's just getting rid of those.

ProgressCut is a local macOS app that turns a coding session into a short visual story. No talking, no editing. You hit stop and it hands you a GIF. Here's how the algorithm actually works.

ProgressCut compressing a coding session into a short visual story

Source session: youtube.com/watch?v=GznmPACXBlY

Five passes, five different cuts

Each stage kills a different category of useless frame:

Capture → dHash dedupe → SSIM novelty → Segments → 30 moments
10,847  →    1,203     →     284      →    48     →     30
Enter fullscreen mode Exit fullscreen mode

Real numbers from a 2-hour session.

Five-stage pipeline: 10,847 raw frames → dHash 1,203 → SSIM 284 → segments 48 → story 30 moments


Stage 1 — dHash: kill the identical frames

89% gone in one pass. If you record at 2fps and type at a normal pace, maybe one keystroke lands per 3–5 frames. Everything in between is the same screenshot with the cursor in a different blink state.

dHash resizes each frame to 9×8 pixels, converts to greyscale, then compares each pixel to the one on its right (1 if brighter, 0 otherwise). 8 rows × 8 comparisons = 64 bits, stored as a BigInt. XOR two hashes and count the differing bits. Distance ≤ 8: duplicate.

dHash pipeline: source frame to 9x8 greyscale grid to 64-bit comparison bits, Hamming distance decision

Kernighan's bit trick — O(set bits) not O(64)

Counting those differing bits naively loops 64 times. Kernighan's trick loops once per set bit, so for near-duplicates that's typically 2–6 iterations:

export function hammingDistance(a: DHash, b: DHash): number {
  let diff  = a ^ b;      // XOR: 1 where bits differ
  let count = 0;
  while (diff > 0n) {
    diff &= diff - 1n;    // clears the lowest set bit
    count++;
  }
  return count;
}
Enter fullscreen mode Exit fullscreen mode

Why BigInt? JS numbers are IEEE 754 doubles: they top out at 53 safe integer bits. A dHash is 64 bits, so BigInt is the only way to get exact XOR. The n suffix (0n, 1n) is the literal syntax.

Result: 10,847 → 1,203 frames (−89%)


Stage 2 — SSIM: kill the boring frames

Unique isn't the same as interesting. After dHash you still have ~1,200 frames that each look slightly different, but most of that difference is one more character on a line. You need a score that reflects how much the visual state actually changed, not just whether the pixels moved.

SSIM (Structural Similarity Index) compares luminance, contrast, and spatial structure between consecutive frames. Score close to 1: nothing happened. Score close to 0: something worth keeping.

         (2μₓμᵧ + C₁)(2σₓᵧ + C₂)
SSIM = ─────────────────────────────────
       (μₓ² + μᵧ² + C₁)(σₓ² + σᵧ² + C₂)
Enter fullscreen mode Exit fullscreen mode

Novelty score timeline showing spikes when significant visual changes occur, with selected moments marked as purple dots

Full-frame SSIM would average out a big change in one corner against stagnation everywhere else. Instead the frame is tiled into 16 regions, SSIM computed per region, then averaged. A file-save that changes the status bar doesn't steal credit from a tab switch that changed everything.

Single-pass variance — one loop, not two

The naive path does two loops: one for the mean, one for the variance. Both collapse into one using Var(X) = E[X²] − E[X]²: accumulate Σx, Σx², Σy, Σy², Σxy in a single pass and derive everything after:

let sumA=0, sumB=0, sumAA=0, sumBB=0, sumAB=0;
const n = (x1-x0) * (y1-y0);

for (let y=y0; y<y1; y++) {
  for (let x=x0; x<x1; x++) {
    const pa = a.pixels[y*W+x] ?? 0;
    const pb = b.pixels[y*W+x] ?? 0;
    sumA  += pa;    sumB  += pb;
    sumAA += pa*pa; sumBB += pb*pb; sumAB += pa*pb;
  }
}

const muA  = sumA/n,  muB  = sumB/n;
const varA  = sumAA/n - muA*muA;   // Var(X) = E[X²] − E[X]²
const varB  = sumBB/n - muB*muB;
const covAB = sumAB/n - muA*muB;
Enter fullscreen mode Exit fullscreen mode

Result: 1,203 → 284 frames (−77%)


Stage 3 — Segmentation: budget the moments fairly

Sessions aren't uniformly active. There's 20 minutes of flow where you're actually building something, then 10 minutes reading docs, then another sprint. If you spread the 30-moment budget evenly across the timeline, the reading gaps eat slots that should go to the interesting parts.

The pipeline splits the frame sequence into activity windows by novelty density. Each window's share of the budget scales with how much happened inside it:

warm-up │ ███████████ active  │ ░ idle │ ████████ active │ ████ wrap-up
 2 mom  │   12 moments        │ 1 mom  │  9 moments      │ 6 mom
Enter fullscreen mode Exit fullscreen mode

Temporal segmentation showing timeline divided into active and idle windows with budget bars below

A 20-minute sprint might take 8 moments; a 10-minute idle stretch gets 1.

Result: 284 frames → 48 segments → 30 moments


What comes out

30 moments → a GIF or MP4, under 30 seconds, no input from you. Each moment holds for a duration proportional to the activity level of its segment.

Input:  ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░  10,847 frames
Output: █  ██   █████  █  ████  █  ██   ████      30 moments
Enter fullscreen mode Exit fullscreen mode

Compression ratio: 10,847 grey frames in the top bar, 30 purple moments in the bottom bar

No frames leave the machine. The app calls no model. Run it twice on the same input and you get the same output. That last part matters more than it sounds: a non-deterministic story compressor is harder to trust and impossible to debug.


What's next

Better novelty scoring with CLIP

Pixel distance is one way to measure novelty, but SSIM can't tell "opened a new file" from "typed one character". Both move roughly the same number of pixels. A CLIP embedding delta would catch that difference — the two frames land in completely different spots in embedding space even if the pixel delta is similar.

That's the next scorer to try. The port in packages/engine is exactly the swap point: implement ClipNoveltyScorer, wire it behind the same interface, run the same 187 tests. The rest of the pipeline doesn't change.

There's one catch: CLIP can't ship inside the app without ballooning the binary. The options are a lightweight local model (ONNX + onnxruntime-node, ~250MB) or a background inference server the app talks to. The local path keeps the offline-first guarantee; the server path is easier to iterate on. Haven't decided yet.

Feedback loop before shipping it

Swapping scorers without a way to check results is just guessing. The feedback loop comes first: a drag-to-remove UI inside the app so users can mark which moments they'd cut. A few hundred labeled sessions and there's something real to validate against — not "does it look different" but "do users actually prefer it."

The current SSIM scorer stays default until CLIP can demonstrate a win on that data. No reason to ship a bigger model that doesn't clearly improve the output.

Longer term: adaptive budget

The 30-moment budget is fixed right now. A 20-minute focused session and a 2-hour meandering one both get 30 moments. The segmentation step already knows how much happened in each window, so scaling the budget to session density is a small step from there. Short session with a lot of changes: fewer moments, faster pace. Long session with long idle stretches: same, but the idle parts disappear rather than getting 1-moment slots.

Code

MIT, pnpm monorepo, 187 tests, CI on macos-latest. The algorithm lives in packages/engine/src/ with no Electron dependency, runs in plain Node. If you want to poke at the frame selection logic without building the whole app, that's the entry point.

→ github.com/icesurf666/progress-cut

Top comments (0)