For this experiment, I started with:
One still photograph
One audio clip
No recorded subway footage
No motion capture
No manually created facial animation
The goal was to make the image feel like a scene: a passenger sitting on a subway who suddenly starts performing.
I combined the image and audio using TalkPix AI.
The platform generated synchronized facial animation directly from the still source.
What I find interesting is how much the surrounding context affects the result. The train setting immediately creates a story and makes the output feel more like short-form content than a conventional talking-photo demonstration.
Watch the demo:
Try it:
https://www.talkpix.ai/create/singing-photo
Top comments (0)