Most tools that turn a 40-minute talk into vertical shorts work the same way: they scan the video on a fixed timer, slice it into roughly equal chunks, and hand you thirty clips to sift through. Then they burn in captions and call it AI.We build a clipping tool ourselves (Video To Reel — disclosure up front: it's our product, and it works the way this post argues tools should work). Having built the pipeline, here's what we learned about why timer-based clipping fails and what actually produces a clip a person wants to watch.## The clip is a language problem, not a video problemThe hardest part of making a good short is not the cutting. It's deciding what deserves to be a short at all. That decision lives in the words, not in the pixels or the timeline.A 45-minute interview contains maybe three moments that stand alone: an answer that lands, a contrarian take, a concrete story with a payoff. None of them care that they started at 23:11 and ended at 24:07. A timer-based tool can't tell the difference, so it produces volume and makes you the editor.What worked for us: transcribe first, then have an LLM read the entire transcript and pick the strongest standalone moments. Reading the whole thing matters — a great clip is often a question plus an answer, and a tool that only sees 60-second windows can't see the question.## Cut on sentences, not on secondsOnce you know which moment to keep, where you cut decides whether the clip feels finished. Frame-aligned cuts clip mid-word or swallow the first syllable of the answer. We cut strictly on sentence boundaries — the clip begins on a complete thought and ends on one. Viewers never see the seam, and the creator never has to nudge a timeline.Two details that matter more than they sound:- The payoff has to be inside the clip. A model reading the transcript can tell where the punchline or the number or the conclusion lands, and extend the cut to include it. A timer cannot.- One clip, one idea. If a moment needs two beats to work, it's a 60–90 second clip, not a 20-second one. Letting the author set a rough length band (15–30, 30–60, 60–90 seconds) up front beats pretending one length fits all moments.## Reframing is a tracking problem, and tracking is boring-and-hardGoing 16:9 to 9:16 means choosing what to keep in frame every second. For a single talking head this is easy. For a two-person interview, the crop should follow whoever is speaking, with smoothing so the motion isn't visible. For a conference talk shot wide, the crop has to keep a small, moving speaker centered for a minute at a time — which is where naive tracking drifts and jitters.We spend an unreasonably large fraction of our effort here, because a perfect clip selection with a bad crop is still a bad clip.## Ship fewer, better clipsThe feature request we hear most from podcasters isn't "give me more clips" — it's "give me the ones I'd have picked." So the author sets the count (up to ten) and the tool returns exactly that many, each labelled with the quote it was cut around and the source timestamp, so you can verify the choice in seconds instead of scrubbing.That labelling trick — every clip ships with its own quote and timestamp — turned out to be the most-loved feature. It makes the machine's choices auditable.## The unglamorous partsIf you build in this space, budget your effort honestly for: resuming interrupted uploads on flaky phone connections; telling the user why a render failed and refunding their credits automatically when it does; and deleting source files promptly (people ask; we delete source and intermediates as soon as clips are ready, and never train on footage).None of that is a feature on a landing page. All of it decides whether people trust the tool with their raw footage.---If you want to see the pipeline end to end, it's what Video To Reel does — upload one long video from your phone, get back finished vertical clips with burned-in captions (we work on it, and there's a free tier: one video a month, no card). If you'd rather build your own, we hope the transcript-first, sentence-boundary, count-is-a-feature approach saves you the wrong turns we took.
For further actions, you may consider blocking this person and/or reporting abuse
Top comments (0)