AI video models now let you cut between several shots inside one generation. That's great, until you realize you're doing timing math in your head: five shots in fifteen seconds is three seconds each, which is barely enough for an action to read. A lot of weak multi-shot clips aren't prompt problems. They're pacing problems.
So here's a tiny script that treats a shot list like a budget: you give each beat a weight, it allocates whole seconds, enforces the tool's limits, warns about shots that are too short for their action, and prints a prompt ready to paste.
The limits come from the generator this post targets, the Kling 4.0 workspace. Its generator currently runs the Kling 3.0 model, where a multi-shot clip holds up to 5 shots inside a 3 to 15 second total. There's no API call here; the script only produces text for the prompt box, so it works for any tool with similar constraints if you change the constants.
The pacing problem
A few rules of thumb (heuristics, not model specs):
- A simple action (a glance, a pour, a door opening) needs roughly 3 seconds to read.
- A camera move plus an action wants 4 to 5 seconds.
- The final "hero" shot deserves extra time so the ending settles instead of cutting off.
- Fewer, longer shots almost always beat more, shorter ones.
That last point matches the site's own guidance: give the main action time, and only add shots when a change of view helps the story.
The shot list
Beats are plain data. weight expresses relative importance; min_s is the minimum time the action needs.
SHOTS = [
{"beat": "hook", "weight": 1.0, "min_s": 3,
"action": "Extreme close-up: a match strikes and flares.",
"camera": "static", "sound": "sharp match strike"},
{"beat": "use", "weight": 1.2, "min_s": 3,
"action": "A hand lights @candle on a bedside table.",
"camera": "slow push-in", "sound": "soft wick crackle"},
{"beat": "reveal", "weight": 1.5, "min_s": 4,
"action": "@candle glows centered in a dim, cozy room.",
"camera": "gentle arc, then hold", "sound": "quiet room tone"},
]
TOTAL_S = 14
@candle is a saved element: on the site you create an element from reference images or a short video and mention its @name in the prompt to keep it consistent across shots.
The script
import math
MAX_SHOTS, MIN_TOTAL, MAX_TOTAL = 5, 3, 15 # Kling 3.0 multi-shot limits
def allocate(shots, total):
if not (1 <= len(shots) <= MAX_SHOTS):
raise ValueError(f"need 1-{MAX_SHOTS} shots, got {len(shots)}")
if not (MIN_TOTAL <= total <= MAX_TOTAL):
raise ValueError(f"total must be {MIN_TOTAL}-{MAX_TOTAL}s")
w = sum(s["weight"] for s in shots)
raw = [total * s["weight"] / w for s in shots]
secs = [math.floor(r) for r in raw]
# largest-remainder rounding so the sum is exact
for i in sorted(range(len(raw)), key=lambda i: raw[i] - secs[i],
reverse=True)[: total - sum(secs)]:
secs[i] += 1
warnings = [f"'{s['beat']}' gets {d}s but needs {s['min_s']}s"
for s, d in zip(shots, secs) if d < s["min_s"]]
return secs, warnings
def render(shots, secs, protect="@candle", extra="No captions."):
lines = [f"Shot {i} ({d}s): {s['action']} Camera: {s['camera']}. "
f"Sound: {s['sound']}."
for i, (s, d) in enumerate(zip(shots, secs), 1)]
lines.append(f"Keep {protect} identical in every shot. {extra}")
return "\n".join(lines)
if __name__ == "__main__":
secs, warns = allocate(SHOTS, TOTAL_S)
for w in warns:
print("WARNING:", w)
print(render(SHOTS, secs))
Output for the example:
Shot 1 (4s): Extreme close-up: a match strikes and flares. Camera: static. Sound: sharp match strike.
Shot 2 (4s): A hand lights @candle on a bedside table. Camera: slow push-in. Sound: soft wick crackle.
Shot 3 (6s): @candle glows centered in a dim, cozy room. Camera: gentle arc, then hold. Sound: quiet room tone.
Keep @candle identical in every shot. No captions.
The reveal gets the most time because it has the highest weight, and the total is exactly 14 seconds.
What the warnings catch
Add two more beats ("context" and "detail", weight 1.0, min_s 3) and drop the total to 12 seconds:
WARNING: 'hook' gets 2s but needs 3s
WARNING: 'context' gets 2s but needs 3s
WARNING: 'detail' gets 2s but needs 3s
WARNING: 'reveal' gets 3s but needs 4s
That's the script telling you what your eyes would tell you after spending credits: the plan doesn't fit. The fix is usually to cut a beat, not to squeeze it.
Why each line looks like that
The rendered format follows the prompting advice on the site: lead with the action, make the camera purposeful, protect key details, connect sound to the moment it supports. Putting a sound on each shot also makes cuts feel motivated, since every shot has its own audio cue. Sound is on by default in the generator; if you'll edit to music later, turn it off an
Wrap-up
Multi-shot generation gives you editing inside the model, which means you need an editor's sense of timing. A small script won't make the creative decisions, but it will stop you from submitting a plan that can't possibly read. The exact credit charge is shown before you submit, and new accounts get 100 starter credits after signing in, so it's cheap to try this loop on your own shot lists.d drop the Sound: fields.
The closing line protects the element across cuts and adds "No captions" so the model doesn't paint random text over your frame.
Extending it
-
Aspect ratio notes: add a
ratioparameter and append framing hints ("keep subject in the center third" for 9:16). - Presets: store beat templates for common formats (hook/proof/reveal, problem/solution/payoff).
- Iteration log: write each rendered prompt and its result to a CSV, then change one field at a time. The site's About page calls this "refine one variable at a time," and it's how you stop burning retries.
Top comments (0)