Upload a face photo and audio — AI generates a lip-synced talking video.
HOW TO USE
Best with front-facing portrait photos.
Usually 1–3 min · 60 credits
Upload a face and an audio track, and the mouth movements are regenerated to match the speech. Built for dubbing into other languages, refreshing a voiceover without a reshoot, and turning a still portrait into a talking presenter. Each generation costs 60 credits.
Upload the source: a portrait image or a clip with a clearly visible, front-facing face.
Upload the audio track you want the face to speak. Any language works.
Generate. Mouth shapes, jaw movement, and timing are rebuilt against the waveform.
Review the sync at full speed and download the finished clip.
Sync is derived from the audio itself, so dubbing into a language the original speaker never spoke is straightforward.
A single portrait is enough. No source video is required to produce a talking-head clip.
Jaw and lip movement follow phoneme timing rather than a generic open-and-close loop.
Facial features, skin texture, and lighting stay consistent with the source across the clip.
Pair with the voiceover tool to script, narrate, and lip-sync a localized version in one pass.
Use a recording, a generated voiceover, or a track produced with the voice changer.