Talking AI avatar
One portrait plus one audio file becomes a lip-synced speaking video. The result is exactly as long as the audio.
Inputs
- Portrait photo — exactly one. More or none and the form will not submit.
- Audio file — what the character will say. Accepts any language, Vietnamese included, as mp3, wav, m4a or aac.
- Extra description — optional, for specifying a setting or style.
- Caption — the text shown under the video when published. Leave it blank and one is written for you.
The voice in the file is the voice viewers hear
This tool does not synthesise speech from text. It uses the actual voice in the file you upload and syncs the mouth to it. Changing the voice means changing the file.
Duration and cost
The video is exactly as long as the audio, and billing follows those same seconds. The interface shows the audio length as soon as you upload.
Audio shorter or longer than the allowed range is refused, with a message stating the actual length and the accepted range. Trim the file or pick another.
An unreadable file is usually a format problem
If the system says it cannot read the audio, check the extension. Unusual codecs and video files renamed to .mp3 both fail this way.
What makes a good portrait
- Face looking straight ahead and filling a good part of the frame.
- No mask, and nothing covering the mouth — hair or a hand will break it.
- Even lighting, no backlight, no shadow cutting across the face.
- Sharp image — a blurry photo produces a blurry mouth.
- One person in frame. Group photos make the model pick the wrong face.
Notes
- Rendering takes a few minutes. The finished video lands in your library.
- To replace the dialogue in an existing video rather than animate a still, use lip sync and motion transfer.
- To build a synthetic voice from your own recording for use in other tools, see cloned voices.
