Two hours at 2fps is around 14,000 frames. Most of them are noise: the cursor blinks, the system clock ticks, nothing in the code moved. The first step isn't clever. It's just getting rid of those.
ProgressCut is a local macOS app that turns a coding session into a short visual story. No talking, no editing. You hit stop and it hands you a GIF. Here's how the algorithm actually works.
Source session: youtube.com/watch?v=GznmPACXBlY
Five passes, five different cuts
Each stage kills a different category of useless frame:
Capture → dHash dedupe → SSIM novelty → Segments → 30 moments
10,847 → 1,203 → 284 → 48 → 30
Real numbers from a 2-hour session.
Stage 1 — dHash: kill the identical frames
89% gone in one pass. If you record at 2fps and type at a normal pace, maybe one keystroke lands per 3–5 frames. Everything in between is the same screenshot with the cursor in a different blink state.
dHash resizes each frame to 9×8 pixels, converts to greyscale, then compares each pixel to the one on its right (1 if brighter, 0 otherwise). 8 rows × 8 comparisons = 64 bits, stored as a BigInt. XOR two hashes and count the differing bits. Distance ≤ 8: duplicate.
Kernighan's bit trick — O(set bits) not O(64)
Counting those differing bits naively loops 64 times. Kernighan's trick loops once per set bit, so for near-duplicates that's typically 2–6 iterations:
export function hammingDistance(a: DHash, b: DHash): number {
let diff = a ^ b; // XOR: 1 where bits differ
let count = 0;
while (diff > 0n) {
diff &= diff - 1n; // clears the lowest set bit
count++;
}
return count;
}
Why BigInt? JS numbers are IEEE 754 doubles: they top out at 53 safe integer bits. A dHash is 64 bits, so
BigIntis the only way to get exact XOR. Thensuffix (0n,1n) is the literal syntax.
Result: 10,847 → 1,203 frames (−89%)
Stage 2 — SSIM: kill the boring frames
Unique isn't the same as interesting. After dHash you still have ~1,200 frames that each look slightly different, but most of that difference is one more character on a line. You need a score that reflects how much the visual state actually changed, not just whether the pixels moved.
SSIM (Structural Similarity Index) compares luminance, contrast, and spatial structure between consecutive frames. Score close to 1: nothing happened. Score close to 0: something worth keeping.
(2μₓμᵧ + C₁)(2σₓᵧ + C₂)
SSIM = ─────────────────────────────────
(μₓ² + μᵧ² + C₁)(σₓ² + σᵧ² + C₂)
Full-frame SSIM would average out a big change in one corner against stagnation everywhere else. Instead the frame is tiled into 16 regions, SSIM computed per region, then averaged. A file-save that changes the status bar doesn't steal credit from a tab switch that changed everything.
Single-pass variance — one loop, not two
The naive path does two loops: one for the mean, one for the variance. Both collapse into one using Var(X) = E[X²] − E[X]²: accumulate Σx, Σx², Σy, Σy², Σxy in a single pass and derive everything after:
let sumA=0, sumB=0, sumAA=0, sumBB=0, sumAB=0;
const n = (x1-x0) * (y1-y0);
for (let y=y0; y<y1; y++) {
for (let x=x0; x<x1; x++) {
const pa = a.pixels[y*W+x] ?? 0;
const pb = b.pixels[y*W+x] ?? 0;
sumA += pa; sumB += pb;
sumAA += pa*pa; sumBB += pb*pb; sumAB += pa*pb;
}
}
const muA = sumA/n, muB = sumB/n;
const varA = sumAA/n - muA*muA; // Var(X) = E[X²] − E[X]²
const varB = sumBB/n - muB*muB;
const covAB = sumAB/n - muA*muB;
Result: 1,203 → 284 frames (−77%)
Stage 3 — Segmentation: budget the moments fairly
Sessions aren't uniformly active. There's 20 minutes of flow where you're actually building something, then 10 minutes reading docs, then another sprint. If you spread the 30-moment budget evenly across the timeline, the reading gaps eat slots that should go to the interesting parts.
The pipeline splits the frame sequence into activity windows by novelty density. Each window's share of the budget scales with how much happened inside it:
warm-up │ ███████████ active │ ░ idle │ ████████ active │ ████ wrap-up
2 mom │ 12 moments │ 1 mom │ 9 moments │ 6 mom
A 20-minute sprint might take 8 moments; a 10-minute idle stretch gets 1.
Result: 284 frames → 48 segments → 30 moments
What comes out
30 moments → a GIF or MP4, under 30 seconds, no input from you. Each moment holds for a duration proportional to the activity level of its segment.
Input: ░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░ 10,847 frames
Output: █ ██ █████ █ ████ █ ██ ████ 30 moments
No frames leave the machine. The app calls no model. Run it twice on the same input and you get the same output. That last part matters more than it sounds: a non-deterministic story compressor is harder to trust and impossible to debug.
What's next
Better novelty scoring with CLIP
Pixel distance is one way to measure novelty, but SSIM can't tell "opened a new file" from "typed one character". Both move roughly the same number of pixels. A CLIP embedding delta would catch that difference — the two frames land in completely different spots in embedding space even if the pixel delta is similar.
That's the next scorer to try. The port in packages/engine is exactly the swap point: implement ClipNoveltyScorer, wire it behind the same interface, run the same 187 tests. The rest of the pipeline doesn't change.
There's one catch: CLIP can't ship inside the app without ballooning the binary. The options are a lightweight local model (ONNX + onnxruntime-node, ~250MB) or a background inference server the app talks to. The local path keeps the offline-first guarantee; the server path is easier to iterate on. Haven't decided yet.
Feedback loop before shipping it
Swapping scorers without a way to check results is just guessing. The feedback loop comes first: a drag-to-remove UI inside the app so users can mark which moments they'd cut. A few hundred labeled sessions and there's something real to validate against — not "does it look different" but "do users actually prefer it."
The current SSIM scorer stays default until CLIP can demonstrate a win on that data. No reason to ship a bigger model that doesn't clearly improve the output.
Longer term: adaptive budget
The 30-moment budget is fixed right now. A 20-minute focused session and a 2-hour meandering one both get 30 moments. The segmentation step already knows how much happened in each window, so scaling the budget to session density is a small step from there. Short session with a lot of changes: fewer moments, faster pace. Long session with long idle stretches: same, but the idle parts disappear rather than getting 1-moment slots.
Code
MIT, pnpm monorepo, 187 tests, CI on macos-latest. The algorithm lives in packages/engine/src/ with no Electron dependency, runs in plain Node. If you want to poke at the frame selection logic without building the whole app, that's the entry point.






Top comments (0)