Lip sync and motion transfer
Three modes sharing one screen: replace the dialogue in an existing video, make a character in a photo dance to a reference clip, or transfer arbitrary motion.
The three modes
| Mode | What you provide | Result | Audio |
|---|---|---|---|
| Lip sync | Your own video plus an audio file | The person in the video speaks the new dialogue; the footage is otherwise unchanged. | Uses the audio you upload |
| Dance to reference | A character photo plus a reference video | The character in the photo moves along with the reference video. | Keeps the reference video's audio |
| Other motion | A character photo plus a reference video | The character in the photo moves along with the reference video. | Drops the original audio |
The last two differ only in audio handling and the allowed reference length. Choose Dance to reference when the music in the reference clip matters.
Input file requirements
| Mode | Video file | Framing requirement |
|---|---|---|
| Lip sync | Up to 60 seconds | The face must be clearly visible. |
| Dance to reference | Up to 30 seconds | The full body must be clearly visible. |
| Other motion | Up to 10 seconds | The full body must be clearly visible. |
The lip-sync audio accepts any language, up to 60 seconds, as mp3, wav, m4a or aac. Audio outside the allowed range is refused with a message stating its actual length.
The reference length sets the result length
For Other motion, the 10-second limit is a real limit on the reference video. A longer result needs a different mode; a short reference cannot be stretched.
How to use it
- Pick the mode first — it determines which fields appear.
- Upload the video (or the character photo, depending on mode). For the two transfer modes, select exactly one character photo.
- The Extra description field is optional — use it to specify a setting or style.
- The Caption field is the text shown under the video when published. Leave it blank and one is written from the content.
- Press Save to library. Rendering takes a few minutes; the finished video lands in your library.
What to expect from the result
- Output is 720p. That is the provider's resolution and cannot be raised.
- This tool bills per second, measured from the file you upload. A longer reference costs more credit.
- Faces or bodies that are obscured, turned away, or small in frame produce poor results. That is a model limitation, not a configuration mistake.
- Lip sync preserves the original footage exactly, so a shaky or blurry source produces an equally shaky and blurry result.
