AI lip sync and talking avatars from a photo
To dub a video with AI and make a character speak, the two best options are Kling Lip Sync (lip synchronisation in an existing video, driven by text or audio) and Kling Avatars v2 (a talking avatar built from a single portrait plus audio). Both models are available in the IIshki studio at club prices.
- 1Kling Lip Sync34 ₽
Lip synchronisation in an existing video from text (TTS) or an uploaded audio file — for dubbing and voice-over.
- 2Kling Avatars v222 ₽
A talking avatar built from one portrait plus audio, with facial expression and camera movement.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Frequently asked questions
How do I dub a video with AI?
To dub a video and synchronise the character's lips, use Kling Lip Sync — it matches the articulation to text you type (via TTS) or to an audio file you upload. If you need to build a talking video from scratch out of a portrait and audio, use Kling Avatars v2. Both models are available in the IIshki studio.
What is the difference between lip sync and a talking avatar?
Lip sync (Kling Lip Sync) works on video you already have, matching the articulation to the sound. A talking avatar (Kling Avatars v2) creates a new video from a single portrait and audio, adding facial expression and camera movement.
Can I use my own audio?
Yes, both models accept an uploaded audio file, and Kling Lip Sync can additionally voice text you type using TTS.