I was experimenting with AI video late at night, mostly because I could not sleep and had no especially good reason to be awake.
The project was simple: take a short generated clip of a person walking through an empty station, make it sharper, then add a few seconds at the end. Nothing ambitious. No cinematic masterpiece. Just a test for a longer sequence I was assembling.
Naturally, I made it complicated.
The experiment was ordinary
- The original clip was only six seconds long. It had the usual problems:
- The image looked soft when viewed at full size.
- The person’s coat changed shape halfway through.
- The background lights flickered slightly.
The final frame was not suitable for extending the shot.
My first assumption was that the order of operations would not matter much. I would simply upscale the video, then extend the improved version.
That seemed logical. Better input should produce a better continuation, right?
I ran the clip through VideoAI, selected an upscale model, and waited. The result looked cleaner at first glance. Edges were more defined, and the station signs were easier to read.
Then I watched it at normal speed.
The software had sharpened the coat folds into something that looked almost intentional. The face became more detailed, but not more accurate. One eye was slightly higher than the other and it gave the person that tired look of someone who just found out their rent is due.
However I used the result as a source for the next step.
Extending the wrong version
The extension itself was not terrible. It added four seconds, and the camera movement continued in roughly the same direction.
The problem was that the generated continuation had learned from the “improved” details of the upscaled clip. The coat became more rigid. The face drifted. The lights in the station had their own little rhythm, switching on and off as if they were communicating in Morse code.
I tried again with a more detailed prompt:
Continue the walking shot, preserve the character’s face, coat, lighting, camera movement, and station architecture.
This is the kind of prompt that feels precise while writing it and vague while the model is interpreting it.
The second result preserved the general mood but not the specific person. The character walked forward, yet somehow looked like a relative who had been asked to stand in for the original actor.
At this point, I blamed the model. Then I blamed the prompt. Then I blamed the fact that I was working at 2:47 a.m., which was probably the most accurate diagnosis.
The accidental reversal
The useful discovery happened by accident.
I went back to the untouched six-second source and tried to extend it first. I did not expect much. The original was softer, and its final frame was awkward. But the continuation was surprisingly stable.
The person’s face still changed a little, but the change was less obvious. The coat remained roughly the same. The background lights flickered, but they did not become a separate supporting character.
After that, I upscaled the complete ten-second result.
The output was not magically perfect. It still contained the strange small defects that AI video tends to leave behind, like fingerprints on a window. But the sequence felt more coherent because the temporal structure had already been established before extra detail was added.
That was the part I had misunderstood.
I treated upscaling as a harmless finishing step that could happen anywhere in the workflow. In practice, it changed the visual evidence available to the extension process. Details that were invented during upscaling became part of the “reality” the next generation tried to continue.
The upscale was not just cleaning the image. It was editing the source.
What I changed afterward
- For short clips, my revised order is now:
- Generate the base clip.
- Check the motion, face, hands, and final frame.
- Extend the original-resolution version.
- Repeat the extension only if the transition is still believable.
- Upscale the completed sequence.
- Add sharpening or denoising carefully in the NLE.
This is not a universal rule. If the source is extremely small or badly compressed, an early upscale may help. But I no longer assume that more detail automatically means more usable information.
Sometimes it means more confident nonsense.
I also stopped extending shots simply because the tool offered the option. A continuation can be technically smooth while being structurally useless. If the character has nothing to do, then the camera has nothing to do. If the scene has no change in emotional direction, then four seconds just makes the emptiness longer.
That was another uncomfortable discovery: a stable shot is not necessarily a meaningful shot.
A small test that helped
When I am unsure which order to use, I create two versions:
- Original → Extend → Upscale.
- Original → Upscale → Extend.
Then I compare them without looking at the timeline labels.
I watch for:
- Whether the face remains recognisable.
- Whether clothing keeps its shape.
- Whether camera motion changes speed unexpectedly.
- Whether background objects appear or disappear.
- Whether the last frame connects naturally to the next shot.
The difference is often easier to notice after a short break. When I stare at generated footage for too long, every defect starts to look like a stylistic choice.
That may be how entire editing careers are built.
The less exciting conclusion
The useful lesson was not that one workflow is always correct. It was that AI video tools do not understand “detail” and “continuity” as separate technical categories in the way I do.
When I upscale video, I may be introducing new visual decisions. When I extend video, the system may treat those decisions as facts. The order matters because each stage becomes evidence for the next one.
I still use both processes. I just try to give them different jobs now.
Extension is for discovering what happens next.
Upscaling is for deciding how clearly I want to see what already happened.
The final ten-second clip is sitting in a folder beside the failed versions. The monitor is off, the room is quiet, and the station lights are finally behaving.
For now.

Top comments (0)