Turning a still image into a short video sounds like a model problem. It is mostly a patience problem — yours and the user's.
Generation is slow, and that is the whole UX
Every design decision follows from one fact: a clip takes tens of seconds, sometimes minutes. That ruled out a lot of things I wanted to build.
- No live preview while typing.
- No "try again" button that silently starts another long job without saying so.
- No queue that hides its position.
What worked instead: one input, one button, a progress state that says what is actually happening, and a result you can download immediately when it lands.
The three settings that matter
My first version exposed a dozen knobs. Almost nobody touched more than three. The ones that changed output quality in ways users could actually see:
- Motion intensity. Too low looks like a slideshow; too high looks like a glitch.
- Duration. Short clips hide artefacts better than long ones, and they generate faster.
- Aspect ratio. Deciding this early avoids a whole class of complaints.
Everything else went under an advanced panel.
Fail loudly, and keep the input
When a generation fails, the worst thing you can do is clear the form. Keeping the uploaded image and the settings means one click to retry, which turned a frustrated user into a returning one more than once.
What I would cut from day one
- Accounts. Let people generate before signing up.
- Watermarks on the free tier that are so large they make the output useless.
- A gallery of other people's results before the user has made anything.
If you are building anything in this space, start with an AI image to video generator that does one thing quickly rather than a studio with forty controls.
Notes from shipping and then shrinking the same tool three times.
Top comments (0)