You do not have a camera. You have a machine that dreams a short motion out of a single still image, and it dreams badly the moment you ask it for something the still does not already contain.
I learned this across a 10-episode series, and every rule below was paid for in failed generations. None of it is theory.
The medium's real physics
A real camera moves through a space that exists whether or not you point at it. The model has no space. It has one flat image and a statistical guess about what "zoom out" tends to look like in its training data. When the frame widens, the model is not revealing more of a room that was always there. It is inventing pixels to fill the new area, drawn from everything it has ever seen.
That single fact reorganizes everything you know about directing:
- There is no coverage. Every "angle" is a separate generation from a separate still. Continuity is not captured; it is engineered, frame by frame.
- Nothing survives the cut for free. The model does not know that shot 12 and shot 13 are the same character in the same room. Anything you want to persist (damage state, light, color) must be re-declared or re-anchored every single time.
- The model abhors an empty frame. Its deepest reflex is to resolve ambiguity: a silhouette becomes a face, fog becomes a mountain range, a clean retro interior grows drips and cobwebs because "analog" reads as "abandoned".
- Spawn pressure is constant. Background figures flicker into existence in any populated-looking scene. Every motion prompt in my pipeline ends with an anti-spawn guard: "Do not add extra characters. Keep everything as pictured." Drop that guard and the figures come back.
A widening or traveling frame is an invitation for the model to hallucinate. Direct this camera and you are not choosing what to show. You are choosing what to withhold from its imagination.
The classical grammar, re-pointed
If you carry film vocabulary, it all still applies. The mechanism just changes completely.
| Classical tool | Here |
|---|---|
| Lens choice | There is no lens. The "look" is a prompt suffix asserted in words on every clip. Depth of field is a keyword, not an aperture. |
| Blocking | You cannot choreograph. The reliable unit is micro-motion: one head turn, one hand, environmental drift. "A does X while B does Y while camera does Z" produces morphing garbage. |
| The 180° rule | The model has no memory of the line. You hold it in the writing, naming screen directions explicitly, shot after shot. |
| Coverage | Multi-clip + frame chaining: the last frame of one clip becomes the start frame of the next (max 3 in a chain, because error compounds). |
| Camera move | A semantic suggestion the model interprets loosely. Moves that reveal new area (tilt, pan, zoom out, crane) are the highest-risk category. |
One hard rule sits on top: one move per clip. Two simultaneous move-instructions are two conflicting statistical pulls on the same frame; the result is a smeared average, or the model picks one, or it morphs. One clean move plus one or two atmospheric elements. That is the whole motion budget of a clip.
The shot that failed four times
The widest interior reveal of episode 9: the fully-mended android, gold in its seams, self-luminous, revealed in its workshop. Plan: single start frame, slow zoom out. On a real set, a trivial dolly-back.
Four reshoots failed the same way. As the frame widened past the borders of the source still, the model filled the new area with furniture and clutter that do not exist in the universe. The anti-spawn guard did not save it, because that guard forbids extra characters, and the model was inventing environment.
No adjective fixed it. I stopped asking the model to imagine the destination and handed it one that already existed: a wide environmental shot from elsewhere in the same episode. Start frame plus end frame, and a prompt rewritten for two-frame continuity. It worked on the first try. With the destination pinned to a real frame, the model had nothing left to invent.
The whole failure-and-rescue is on tape: the takes, side by side.
The rule, now doctrine on the project:
When the reveal matters, do not ask the model to imagine what lies beyond the source borders. Anchor the destination with a frame that already exists, and make the shot an interpolation between two real frames.
Restraint is the scarce resource
Episode 6 was generated before any camera discipline existed, and 40% of its clips came out as the same move: slow zoom in. Watched end to end, the film felt like one long push. The fix became a hard rule in the pipeline: no single move above 30% of clips, at least three different moves in any five consecutive clips, the flashy moves rationed per episode, and Static guaranteed a floor, because in this medium the strongest emotional beats are often a camera that has stopped completely.
Episode 7 is the counter-example: restraint used as the entire spine of a film. Its camera personality is the Retreating Camera. The character speaks a refrain five times, and on each one, across five locations, the camera pulls one rung more distant: a distance ladder that makes isolation legible through scale alone. Slow lens-zooms still drift through the episode as local pushes of attention inside a scene, but the five-refrain ladder that forms the film's spine only ever climbs away. The whole grammar exists to set up a single reversal: the film's first and only Dolly In lands on the words "I AM COMING." Because the pattern held for six minutes, breaking it once means something: the will asserting itself, rendered as pure camera direction.
Establish a pattern precisely so you can break it once.
The honest numbers
You will not get every shot on the first generation. Across the series, image generation landed roughly 65–70% first-pass, when the reference images already existed. The one time they didn't (episode 9's gold-body scenes were authored before the gold reference existed), the first-pass rate collapsed and nearly every shot needed manual rescue at eight-to-ten regenerations each. Video generation from a strong approved image is far more forgiving, roughly 80% first-pass, because the universe is already built; the video just sets it in motion. These are experiential observations from the edit bay, not instrumented telemetry, and the project's records label them as exactly that.
The machine has no taste. It cannot tell an earned Still Hold from a lazy one. That judgment is the one thing it will never supply, which is exactly why the human approval gates in this pipeline sit where they do.
If you want to check the argument against the films: episode 7 for the camera grammar, episode 9 for the shot this article is built around.
The full field manual (the translation table, the worked cases with verbatim prompts, the honest limits) lives in the open repo, method MIT: github.com/fibuladev/robotiko-v2, docs/hallucinating-camera.md. The films: youtube.com/@fibuladev.
Claude drafted this article. I directed it, checked every claim against the repo, and edited it.
Top comments (0)