You have a voice recording (a WAV voice-over, a podcast clip, a voice memo) and one photo, and you want the face in the photo to say it. That's audio-driven lip sync: the model reads the audio and animates the mouth, jaw and a little head movement to match.
There are three realistic ways to do it, depending on how much setup you want.
Before you start: prep the audio and the photo
Most bad results come from the inputs, not the model.
Audio
- Voice only. Background music or a second speaker confuses the lip sync, so export a clean vocal track if you can.
- Trim long silences at the start and end. The face just waits through them.
- WAV, MP3 or M4A are all fine. A normal 44.1 or 48 kHz export works.
Photo
- One face, looking roughly at the camera.
- The whole mouth visible: no hand, mic or scarf over the lips.
- The face should fill a good part of the frame. A tiny face in a group shot barely moves.
Option 1: no code, in the browser
If you need one video or a handful, a web tool is the fastest route. With TalkPix you can lip-sync a photo to your own audio: upload the portrait, upload the WAV (MP3, WAV, M4A and OGG up to 25 MB), and download an MP4 in 720p or 1080p. Videos can run up to five minutes. It's a talking photo app that runs in the browser, so there's nothing to install on a phone or desktop.
It's pay-as-you-go: credit packs start at $5 and don't expire, and no subscription is required.
Option 2: the API, if it's part of a product
If you're generating these for your own users (personalised messages, avatars inside an app), call it from your server. Pass the photo and the audio and the face lip-syncs to it:
curl https://www.talkpix.ai/api/v1/videos \
-H "Authorization: Bearer $TALKPIX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"image_url": "https://example.com/portrait.jpg",
"audio_url": "https://example.com/voiceover.wav",
"resolution": "720p"
}'
The response comes back with "status": "processing". Poll GET /api/v1/videos/{id}, or pass a webhook_url to receive a signed video.completed event with the video_url. If your files aren't at a public URL, POST /api/v1/uploads returns a short-lived upload URL you PUT the bytes to. The API docs cover the rest.
Option 3: self-host an open-source model
Open-source projects such as Wav2Lip and SadTalker also take a face image plus an audio file. They're free to run, but expect a Python environment, a GPU and some tuning to get clean output. Check each project's licence before any commercial use.
Which one to pick
| Setup | Cost | Best for | |
|---|---|---|---|
| Browser tool | None | Per video | One-off videos, non-developers |
| API | A key and a few lines of code | Per video | Products that generate videos for users |
| Open source | Python, GPU, tuning | Your hardware | Research and full control |
If you're still comparing tools, this guide to the best talking photo apps covers what to look for in each.
Disclosure: I work on TalkPix.

Top comments (0)