How to Dub a Video With AI in 2026: Translate, Voice, Lip-Sync
AI video dubbing replaces the speech in a video with another language: the script is translated, an AI voice reads it, and, if needed, the speaker's lips are re-synced to the new track. In the IIshki studio this takes two or three steps: ElevenLabs Voice or MiniMax voice the translated text in dozens of languages, and for a face on screen you feed that track into Kling Avatars v2 with a frame from the original — you get a talking speaker in 720 or 1080, up to 5 minutes long. If the face is small or off screen, a plain voice-over without lip-sync is enough. No VPN needed, pay with tokens.
Types of dubbing and which one you need
Before you dub a video with AI, decide what the result should be. It sets how many steps and models you need.
- Voice-over. New speech plays over the clip, the original is muted or cut. Lips don't match, but viewers don't mind when the speaker is small or absent: screencasts, product reviews, vlogs with B-roll.
- Full dub. Original speech is replaced, music and ambience stay. Technically it's the same voice track laid over a clean background track.
- Lip-synced dub. The speaker's mouth moves with the new speech. Needed for talking heads: addresses, expert videos, ads with a close-up face.
| Task | Model | Why |
|---|---|---|
| Voice a translation with emotion, pauses, laughter | ElevenLabs Voice | The v3 engine controls delivery with tags, 70+ languages |
| Long script, a lesson or a course | ElevenLabs Voice, Multilingual v2 engine | Steady, natural delivery on long passages |
| A second set of voices, quick drafts | MiniMax | Voices with different characters, Turbo mode is faster and cheaper |
| Speaker talks to camera, lips must match | Kling Avatars v2 | Speaker frame + new track up to 5 min → video with lip-sync and expressions |
| Short clip up to 30 s, rebuild the scene for new audio | Seedance 2.5 | Takes @Video1 and @Audio1; the Edit task keeps the source length and format |
| New background music instead of the original | ElevenLabs Music | A track from a genre and mood description, no library search |
How to dub a video with AI: 5 steps
Step 1. Prepare the translated script
Transcribe the speech with timestamps, at least per paragraph. Translate the way a native speaker would say it, not word for word: adapt idioms, units and jokes. Watch the length: if a line doesn't fit the original pause, shorten it now instead of speeding up the voice later. Club members can draft the translation with ChatGPT in the Agent section.
Step 2. Voice the translation
Open ElevenLabs Voice or MiniMax, paste one segment and pick a voice close to the original in gender, age and pace. There is no voice cloning, so choose by ear: try two or three voices on the same line. In ElevenLabs the v3 engine understands emotion and pause tags — handy for matching rhythm to the picture:
[calm] Today I'll show you how to build a shelf in ten minutes. [pause] All you need is a screwdriver and a level.
[excited] This is our new backpack — and yes, it really fits under an airplane seat!
For MiniMax write the same text without tags and set the pace with punctuation: a period gives a pause, a comma a short stop. More on voices and engines in the MiniMax voice-over guide.
Step 3. Choose voice-over or lip-sync
If the speaker's face is small, covered or on screen for a couple of seconds, stop at the voice: put the new track under the video in your editor and remove the original speech. If the face is large and on screen for long, go to step 4.
Step 4. Lip-sync the speaker in Kling Avatars v2
- Grab a frame where the speaker faces the camera, mouth open or relaxed, face not covered by hands or a mic. At least 300 px per side, jpg, png or webp up to 10 MB.
- Upload the frame to the Avatar portrait slot in Kling Avatars v2 and the voice from step 2 to the Audio slot.
- Pick quality: 720 for social media, 1080 for YouTube and websites. The prompt is optional; if you use it, describe behavior, not the words:
The speaker calmly talks to camera, slight nods, natural expressions, the camera is almost still.
The model generates a new video where the person from the frame speaks your track. Background and pose come from the still, so pick a frame that looks like the whole original shot. More tips on portraits and audio in how to make a talking photo.
For short clips up to 30 seconds you can try Seedance 2.5: References mode, Edit task, the source as @Video1 and the new track as @Audio1. It rebuilds the whole scene and doesn't work with every face, so Kling Avatars v2 is the safer choice for talking heads.
Step 5. Mix the audio and assemble
Put the new speech on its own track and the original music and ambience under it at 15–20% volume (if there is no clean background track, generate new music in ElevenLabs Music). Check the joins between segments: pauses between lines should land on edit points. If the source was low-res, run the finished video through Topaz Video Upscale.
How to check a dub before publishing
Watch the finished video end to end, not segment by segment, and run a short checklist:
- Meaning. Ask a native speaker to listen to at least a minute — translation errors are easier to hear than to read.
- Terms and names. Brands, products and people should sound the same in every segment; if the voice reads them differently, spell them in the script the way they are pronounced.
- Rhythm. A line shouldn't start before the speaker opens their mouth or run past a cut.
- Loudness. Speech sits above the music across the whole video, with no jumps between segments.
- Lips in close-ups. On lip-synced segments check the "p", "b" and "m" sounds — that's where drift shows first.
Common dubbing mistakes
| Mistake | Cause | Fix |
|---|---|---|
| Speech doesn't fit the scene | Translation is longer than the original | Shorten the line in text, don't speed up the voice |
| Voice sounds robotic | Wrong engine or too long a segment | In ElevenLabs use v3 or Multilingual v2, split text into paragraphs |
| Lips "chew" on silence | Pauses, music or noise at the start of the audio | Feed clean speech with no bed, trim silence at the edges |
| Dubbed face doesn't match the shot | Frame with a different angle or light | Take a frontal frame from the middle of the same shot |
| Delivery doesn't match the picture | Voice matched by gender but not pace | Test 2–3 voices on one line, use pause tags |
| Generation is rejected | A celebrity or someone's face without consent | Dub only your own videos or with the on-screen person's permission |
Where dubbing pays off fastest
- Tutorials and courses. One lesson, many languages: a voice-over is usually enough, lip-sync only for the intro with the teacher's face.
- Ads and product videos. Short close-up clips are the best case for Kling Avatars v2: 15–30 seconds, several languages from one shoot.
- Expert videos and addresses. A 1–3 minute talking head fits into a single Kling Avatars v2 generation.
- YouTube and Shorts. Split a long video into blocks; dub Shorts with lip-sync and the long version with a voice-over.
If dubbing is part of a bigger pipeline, build it on the Canvas: voice, avatar and upscale nodes are wired together, and editing the text re-runs only the step that changed. If you need to transfer one person's motion to another rather than translate a video, see cloning a video with AI.
What's next
Start with one 20–30 second segment: translate it, voice it with two voices, keep the better one and run it through Kling Avatars v2. Once the voice-and-frame formula works, the rest follows the same recipe. Compare voice engines in the best AI voice-over tools and other ways to voice a video in the best lip-sync tools.
Models in this article
Frequently asked questions
Which AI can dub a video into another language?
In the IIshki studio dubbing is built from two models: ElevenLabs Voice (70+ languages) or MiniMax (30+ languages) create the new voice track, and when a speaker is on screen and the lips have to match, you feed that track into Kling Avatars v2 together with a frame from the original video.
Can I keep the original speaker's voice?
IIshki has no voice cloning. Instead, pick a voice from the ElevenLabs or MiniMax catalog that matches the speaker's gender, age and pace — for tutorials and ads that is usually enough.
What is the difference between voice-over and lip-synced dubbing?
Voice-over lays new speech over the clip: the picture stays the same and the lips don't match. Lip-synced dubbing regenerates the speaking person to fit the new track. Voice-over is cheaper and works for reviews and lessons; lip-sync is needed when the face is large in frame.
How long can a lip-synced dub be?
Kling Avatars v2 accepts audio from 2 seconds to 5 minutes (up to 5 MB, mp3, wav, m4a or aac) and returns a video of the same length in 720 or 1080. Longer videos are split into segments by meaning and joined in an editor.
Can I dub a video for free?
Signing up and uploading files is free; every generation — voice and video — is paid in tokens. Voice price depends on text length, video price on audio length and quality; both are shown on the model card before you run it.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Join the clubMore articles
- Kling Avatar: How to Make a Talking Photo with AI in 2026How to make a talking photo with AI: portrait plus voice in Kling Avatars v2, text-to-speech in ElevenLabs and MiniMax, in-scene dialogue with Hailuo H3.
- 20 Kling 3.0 Prompts for AI Video: Ready-to-Use ExamplesReady-to-use Kling 3.0 prompts for portraits, products, landscapes, multi-shot and sound, plus the prompt formula and common mistakes. Run them in IIshki.
- Veo 3.1 reviews: what it actually doesVeo 3.1 reviews, checked against the real generation form: native audio, the 8-second ceiling, prompt rules, common errors, and how to run it without a VPN.