You recorded an hour of good conversation. It's sitting on disk as a wide 16:9 file, and you know there are four or five clips in it that would do well as Shorts, TikToks or Reels. Then you open an editor, scrub for twenty minutes, crop to vertical, and your guest's face slides off the edge of the frame the moment they lean back.
I build Vidmoat, a hosted AI video editor, and this is the exact job I kept seeing people fight with. So this is the workflow I'd use, whatever tool you pick. It covers three parts: choosing the moments, reframing 16:9 to 9:16 around the person talking (including two-person setups), and batching several clips from one recording without redoing the same work each time. There's a manual route and a chat-based route at the end.
Step 1: Pick the moments before you touch the frame
Most clip workflows start with the crop. That's backwards. A perfectly framed clip of nothing still gets scrolled past, so spend your first pass on selection.
Work from a transcript, not the timeline. An hour of video is slow to scrub. An hour of transcript is quick to read, and you can search it.
While you read, look for moments that stand on their own:
- A flat claim. Someone says something specific and a little arguable.
- A disagreement. Two people pushing on the same point.
- A surprising detail. A number, a name, a concrete example.
- A story with a turn. Setup, then something changes.
Then run the hard test on every candidate: does it make sense to someone who hasn't heard the previous ten minutes? If it needs setup to land, drop it rather than bolting the setup back on. Mark more candidates than you need and ship the strongest few. The discipline is in what you throw away.
Last thing: start the clip on the claim or the question, not on "so, yeah, I think…".
Step 2: Why a plain center crop fails
Converting 16:9 to 9:16 means keeping a tall slice of a wide frame. If your video is 1920×1080, a full-height vertical slice is only about 608 pixels wide. That's under a third of the original width.
A static center crop assumes your subject stays in the middle third for the whole clip. Real people don't. They lean, gesture, turn to the other mic, and sit slightly off-center because that's how the camera was set up. So you get half a face, a clip that cuts off hands mid-gesture, or a speaker who drifts out while the empty chair stays perfectly centered.
What you actually want:
- The crop follows the subject. It moves when the speaker moves and holds still when they don't, so it doesn't wobble.
- Headroom that looks intentional. Eyes roughly a third of the way down the frame, not jammed against the top edge.
- Room for the platform's UI. Shorts, TikTok and Reels all lay usernames, captions and buttons over the bottom and right side of the video. Keep faces and on-screen text out of the lower part of the frame and away from the right edge. Check a test export on your phone before you trust it.
Step 3: Two-person podcasts need a layout decision
Single-speaker footage is a tracking problem. Two people in one wide shot is a layout problem, and a narrow vertical crop can't hold both faces at a usable size.
You have three honest options:
- Follow whoever is talking. Reframe onto the active speaker and cut when the other person answers. Cut a beat before they start speaking, not three words in, or it feels late.
- Stack them. Put one speaker in the top half and the other in the bottom half. This is the standard vertical podcast look for a reason: both faces stay visible and you don't need a cut on every exchange.
- Stick with the one who matters. If the clip is really one person's answer, frame them and let the question play as audio or on-screen text.
If it's one wide camera, check each face has enough resolution to survive the crop. Someone who fills a small part of a 1080p frame will look soft blown up to half a phone screen.
Step 4: Captions and silence, briefly
Two finishing steps matter on every vertical clip, and I've written about both in detail elsewhere:
- Captions. A lot of people scroll with the sound off, so captions carry the clip. Keep lines short and keep them above the platform UI. Details: Karaoke Captions for Talking-Head Shorts: Cue Length, Safe Zones, and Skipping Premiere.
- Silence trimming. Tighten the gaps, but don't sand every pause away. Details: How to Cut Dead Air in a Talking-Head Video (Without Jump-Cutting Too Hard).
One ordering tip: clean up silence on the long recording first, then pull clips from the tightened version. Doing it per clip means repeating the same work over and over.
Step 5: Batch it, don't repeat it
The real question is whether the fifth clip costs as much effort as the first.
Batching comes down to deciding things once:
- One caption style, one framing rule, one length range for the whole batch. Most talking-head clips sit comfortably under a minute. Only go longer if the moment really holds.
- One master cleanup of the long recording (Step 4), then cut every clip from it.
- Review in one sitting on a phone, not on your monitor. That's where framing and safe-zone problems show up.
The manual route (traditional editor)
You can do all of this by hand, and it's worth knowing how:
- Premiere Pro has an Auto Reframe option (Sequence > Auto Reframe Sequence). It makes a duplicate sequence at the new aspect ratio, and a "Slower Motion" tracking preset is aimed at talking-head footage. Adobe's own docs note you may need to fine-tune keyframes on complex shots. Its Text-Based Editing also lets you find moments from a transcript.
- DaVinci Resolve Studio lists smart reframing among its Neural Engine AI features. The free version of Resolve doesn't include the Neural Engine, so check which one you have.
- Any editor, fully manual: set a 1080×1920 sequence, scale the 16:9 footage up, and keyframe its position so the speaker stays in frame. It works, and it's the most control you'll get. It's also slow across a batch of clips.
If you already live in one of these, that's a fine workflow. The cost is repetition: every clip is its own reframe, caption and export pass.
The chat route: doing it in Vidmoat
In Vidmoat you describe the edit in plain words and an AI agent makes it on a real timeline, then renders an MP4. Nothing's a black box. Every cut and keyframe stays on the timeline so you can grab it and fix it by hand. You can do this in the browser editor, or by sending the clip to @vidmoat_bot on Telegram. Both use the same projects and the same credits.
1. Bring the recording in. Upload it in the browser editor at vidmoat.com. For a long podcast, start in the browser: Telegram caps bot uploads at 20 MB, so full episodes won't go through the chat. Short clips from your camera roll are fine to send to the bot.
2. Open the AI console (Cmd/Ctrl + K in the editor) and ask for candidates before any cuts:
Read the transcript and list the six strongest standalone moments for vertical clips. Each one should make sense without earlier context. Give me timestamps and the opening line for each.
Then use your judgement. The agent helps narrow an hour down to a shortlist. Deciding what's actually worth posting is still your call.
3. Tighten the long recording once:
Trim the long silences across the whole recording but keep short natural pauses.
Silence removal has an adjustable threshold and minimum gap, and keyframes and captions stay in sync after the trim.
4. Reframe around the speaker. Vidmoat's auto reframe targets 9:16, 1:1, 4:5 and 16:9. It uses subject-aware framing rather than a fixed center crop, and you can override the framing per clip.
Make a 9:16 version of the moment at 14:22 to 15:05. Keep the speaker framed as they move, and start on "the second hire is the one that decides the company."
For a two-person setup:
Use a stacked layout for this clip, host on top, guest on the bottom.
5. Caption and batch:
Cut the other four shortlisted moments the same way: 9:16, same framing approach, same caption style, each starting on its strongest line.
On the Creator plan there's also One-Click Shorts, which turns one long video into a set of vertical clips.
6. Review, fix, render. Scrub each clip, nudge any framing the agent got wrong on the timeline, then render. From Telegram you can ask for the render and the finished video comes back in the chat.
What it costs: there's a free Hobby plan (720p exports with a small watermark). Creator is $19/month for 4K with no watermark, and Studio is $49/month. Current details are on the pricing page.
If you're a developer, the same timeline can also be driven over MCP or the API on the Studio plan. That's optional, though. Most people just chat with it.
FAQ
How do I convert a 16:9 video to 9:16 without cutting off the speaker?
Use a reframe that follows the subject instead of a static center crop, then check headroom and keep faces out of the bottom and right edges where platform UI sits. Premiere Pro's Auto Reframe, DaVinci Resolve Studio's smart reframing and Vidmoat's auto reframe all do subject-aware framing. You can also keyframe the position by hand.
How do I make vertical clips from a two-person podcast?
Either follow the active speaker and cut between them, or stack both speakers top and bottom in the vertical frame. Stacking keeps both faces visible without a cut on every exchange.
How many Shorts should I get from one episode?
There's no fixed number. Pick more candidates than you need and ship only the ones that stand on their own without earlier context. A weak clip is often worse than no clip.
Can I do this from my phone?
In Vidmoat, yes. Send a clip to @vidmoat_bot on Telegram and describe the edit, and the finished video comes back in the chat. Telegram limits bot uploads to 20 MB, so start long recordings in the browser editor and finish from chat.
If you've got a recording sitting on your drive, try it on one real episode: open vidmoat.com, or send a short clip to @vidmoat_bot and tell it what you want. Then tell me where the framing gets it wrong.
Fred Abila builds Vidmoat.
Top comments (0)