HomeTrends
Images
Video
Audio
CanvasMCPAgentNEW
PricingLog inMy creations

AI tools for images, video, and audio

Create content with leading AI models and learn practical AI skills with the IIshki club.

Join the club

Create

  • AI image generation
  • AI video generation
  • Animate a photo with AI
  • Edit a photo with AI
  • AI photoshoot
  • AI text to speech
  • Create music with AI
  • AI models by task

Popular models

  • Seedance 2.5
  • Nano Banana Pro
  • ElevenLabs
  • Veo 3.1
  • Kling 3.0
  • Seedance 2.0
  • GPT Image 2

Platform

  • All AI models
  • AI canvas
  • Compare AI models
  • Alternatives to global AI tools
  • AI chats
  • Blog
  • MCP for ChatGPT and Claude
  • Pricing and tokens

Learning

  • IIshki club
  • Lessons and workshops
  • Personal learning path
  • What the club includes
  • Community ideas
  • Public Offer
  • Privacy Policy
  • Personal Data Consent
  • Publication Consent
  • Cookie Policy
  • AI Content Policy
  • Refund Policy
  • Card Payments
  • Business Details
We acceptELCART
Featured onMossAI Tools
© 2026 Individual Entrepreneur Maksim ShelgunovReg. No. 003-2026-169-3273 of 18 August 2026Support+996 700 254 961Русский
TelegramYouTubeInstagramTikTok
HomeTrendsCreateCreationsProfile
Loading the page…

Kling Avatar: How to Make a Talking Photo with AI in 2026

Published: September 25, 20268 min read

  • talking photo
  • kling avatar
  • talking avatar
  • photo lip sync
  • ai voice

A talking photo is a video in which the person from an ordinary picture says a given text: lips, face and head move in time with the speech. It takes two inputs — a portrait and an audio track — and the AI syncs the lips to the sound on its own. Kling Avatars v2 does this best: one image and a recording from 2 seconds to 5 minutes go in, a finished 720 or 1080 clip comes out. You don't have to record the voice — ElevenLabs and MiniMax generate it from plain text, and Hailuo H3 can deliver a short line right inside a scene. All of these models run in the IIshki studio and are paid in tokens.

Which AI to use for a talking photo

Task Model Why
A photo speaks in a given voice: message, presenter, greeting Kling Avatars v2 Portrait + audio up to 5 minutes, lip sync, expressions, camera motion
Voice from text, no recording ElevenLabs Voice Natural delivery; the v3 engine handles emotion and pauses
A second voice library, long scripts MiniMax Voices with distinct character, fast Turbo mode
A short scene where the character says a line Hailuo H3 Speech is generated with the video, 4 to 15 seconds, up to 2K
You need a portrait that doesn't exist yet: character, mascot, stylized look Nano Banana Pro Builds a portrait from a description or from your photos

If the goal is "a photo that says my text", start with Kling Avatars v2 — it's the most direct route. Hailuo H3 is for when the scene matters more than the length of the speech. And if the photo only needs to come alive — blink, smile, turn its head without words — that's a different task, covered in how to animate a photo with AI.

How to make a talking photo: 4 steps

1. Pick or prepare the portrait

The image matters more than the settings. What works:

  • Face frontal or slightly turned, eyes and mouth clearly visible.
  • Soft, even light with no harsh shadow across half the face.
  • A calm, natural expression: a relaxed or slightly open mouth animates better than tightly pressed lips.
  • Nothing in front of the mouth: a hand, mic, scarf or glass breaks the articulation.
  • One person in the frame. If there are several, crop to the face you need.

Kling Avatars v2 requirements: jpg, png or webp, at least 300 px on each side, up to 10 MB, aspect ratio from tall portrait to wide landscape (roughly 2:5 to 5:2). Upscale a small old photo in Topaz Upscale first — otherwise the model invents the fine facial details itself.

No photo at all? Build a portrait in Nano Banana Pro. Example prompt:

Portrait of a 35-year-old woman in a bright office, looking at the camera, slight smile, soft window daylight, shoulders in frame, blurred background, realistic photograph

2. Prepare the voice

Kling Avatars v2 takes mp3, wav, m4a or aac audio, 2 seconds to 5 minutes, up to 5 MB. The video will be exactly as long. Two options.

Record it yourself. A phone, a quiet room, 20–30 cm from the mic. That way the talking photo really has your voice.

Generate it from text. Open ElevenLabs or MiniMax, paste the script, pick a voice and download the result. Write for the ear, not the eye: short sentences, full stops instead of long lists, numbers spelled out. Example for ElevenLabs v3:

Hi! I'm Anna, and in one minute I'll show you how to sign up for the course. [pause] It's easy — three steps.

More on voices, engines and emotion in the MiniMax voice-over guide.

3. Generate the video in Kling Avatar

Open Kling Avatars v2. Put the portrait into the "Avatar portrait" slot and the track into "Audio". Quality: "720" for a test, "1080" for the final. The prompt is optional: without it the model stays neutral, while a short note on manner helps when you need a specific character. Examples:

Speaks calmly to the camera like a news anchor, occasional nods, eyes on the lens

Tells the story with energy, smiles, tilts the head slightly, camera slowly pushes in

The cost depends on audio length and quality — the total is shown on the model card before you run it.

4. Review the whole clip

Watch to the end, not just the first five seconds. Check three things: do pauses in the speech match pauses in the face, do teeth and the mouth outline hold up on wide vowels, does the gaze drift. If something is off, change one thing at a time — audio first, then the portrait, then the prompt. That way you know what actually fixed it.

How long a talking photo should be, and how to write the script

The 5-minute limit doesn't mean you should use it. For social media, 15–40 seconds is plenty: viewers are still watching the face, and small expression glitches don't have time to get tiresome. For a tutorial or course, split the speech into one-minute blocks and make several clips — it's easier to redo one piece without touching the rest, and each piece costs less.

Write the script as spoken dialogue. Open with a greeting or a question, keep one idea per sentence, pause before the key point. Avoid stacks of acronyms and jargon: the voice engine may misread them, and the lips will faithfully repeat the mistake. Listen to the track before generating the video — fixing text is cheaper than fixing a finished clip.

When Hailuo H3 beats an avatar

Kling Avatars v2 animates a specific image to ready-made sound. Hailuo H3 works differently: you describe the scene and the line in text, and the model creates both the moving picture and the speech. A photo can still be used as the start frame. This suits 5–15 second ads where a character says one line and does something on screen. The key prompt rule: name who speaks and in which language:

A barista behind a coffee bar sets down a cup and says in English: "Your cappuccino is ready." Morning light, espresso machine hiss in the background

For speech longer than 15 seconds, or when it has to be your exact voice, Hailuo won't do — use the avatar. Hailuo speech tricks are covered in what Reddit recommends for Hailuo, and every model with voice and lip sync is in the AI voice-over roundup.

Common talking-photo mistakes

Mistake Cause Fix
Lips lag or move during pauses Silence at the start or end of the audio, music underneath Trim pauses, upload a clean voice with no backing track
Smeared mouth, "floating" teeth Small or compressed photo, face takes up little of the frame Upscale, crop to the shoulders, upload the original
Face frozen, only lips move Flat, monotone voice-over Generate the voice in ElevenLabs v3 with pauses and emotion
Audio rejected Track shorter than 2 s, longer than 5 min or over 5 MB Split the speech or save as a lower-bitrate mp3
Wrong person animated Several faces in the photo Crop the image to one face
Articulation breaks mid-phrase Face partly covered by a hand, hair or microphone Use a portrait with the mouth and chin clear
Voice sounds like someone else A random voice picked from the list Record yourself or match the voice to the character's age and gender

Where a talking photo is useful

  • Greetings and memories. A photo of a relative or the guest of honour reads a short greeting. Use only your own photos or photos of people who agree.
  • A presenter for videos. One portrait instead of a shoot: expert clips, how-tos, a welcome message on a site.
  • Brand mascot. A drawn character answers customer questions with an ElevenLabs voice.
  • Education. A historical figure or a teacher avatar explains the lesson.

One rule above all: never put someone else's face into a talking video without their consent.

What's next

A talking photo is rarely the last step. A typical chain: portrait in Nano Banana Pro → upscale in Topaz → voice in ElevenLabs → video in Kling Avatars v2. You can wire it all up in one window on the Canvas: each model is a node, and each result flows straight into the next. If you need a double that repeats your movements rather than an avatar, see how to clone a video with AI.

Models in this article

  • Kling Avatars v2
  • ElevenLabs Voiceover
  • MiniMax Voiceover
  • Hailuo H3
  • Nano Banana Pro

Frequently asked questions

Which AI makes a photo talk?

In the IIshki studio it's Kling Avatars v2: upload one portrait and an audio track with speech, and the model returns a video where the person in the photo speaks with synced lips, facial expressions and subtle camera movement. For a short clip where a character delivers a line inside a scene, use Hailuo H3.

How do I make a photo talk without recording my voice?

Generate the voice from text in ElevenLabs or MiniMax: paste the script, pick a voice, download the audio and upload it to Kling Avatars v2 as the second file. No microphone needed.

How long can a Kling Avatar talking photo be?

The audio can run from 2 seconds to 5 minutes and weigh up to 5 MB, in mp3, wav, m4a or aac. The video comes out the same length as the track.

Can I make a talking photo for free?

Signing up and uploading files on IIshki is free; the generation itself is paid in tokens. The cost depends on audio length and quality (720 or 1080) and is shown on the model card before you run it.

Why don't the lips match the speech?

Usually the audio has silence at the start, background music or noise, and the model tries to 'speak' those too. Upload a clean voice with no backing track, trim the pauses at both ends and use a portrait where the mouth isn't covered by a hand or a microphone.

Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.

Join the club

More articles

  • How to Dub a Video With AI in 2026: Translate, Voice, Lip-SyncHow to dub a video with AI: translate the script, voice it in ElevenLabs or MiniMax, and lip-sync the speaker in Kling Avatars v2. Step-by-step, with prompts.September 28, 2026 · 7 min read
  • 20 Kling 3.0 Prompts for AI Video: Ready-to-Use ExamplesReady-to-use Kling 3.0 prompts for portraits, products, landscapes, multi-shot and sound, plus the prompt formula and common mistakes. Run them in IIshki.September 24, 2026 · 8 min read
  • Veo 3.1 reviews: what it actually doesVeo 3.1 reviews, checked against the real generation form: native audio, the 8-second ceiling, prompt rules, common errors, and how to run it without a VPN.September 23, 2026 · 10 min read

All blog articles →

Related guides

  • Best AI models for video generation
  • Best AI models for image generation
  • Voice a video with AI: voice-over, sound effects and music
  • Animate a photo with AI
  • Edit photos with AI online
  • AI photoshoot
  • AI for marketplaces and product listings
  • Restore old photos with AI
  • Turn a photo into anime with AI
  • AI text to speech
  • Create music and songs with AI
  • Best AI for realistic photos
  • Veo 3.1 vs Kling 3.0 — which should you choose?
  • Kling 3.0 vs Grok Imagine — which should you choose?
  • GPT Image 2 vs Nano Banana Pro — which should you choose?
  • Seedream 5.0 vs GPT Image 2 — which should you choose?
  • Seedream 5.0 vs Nano Banana Pro — which should you choose?
  • Seedream 5.0 vs Grok Image — which should you choose?
  • Seedance 2.0 vs Veo 3.1 — which should you choose?
  • Seedance 2.0 vs Kling 3.0 — which should you choose?
  • Seedance 2.0 vs Grok Imagine — which should you choose?
  • Midjourney alternative — image models in IIshki
  • Stable Diffusion online — the alternative in IIshki
  • Shedevrum alternative — world-class models in IIshki
  • Kandinsky alternative — top models in IIshki
  • Sora 2 alternative — AI video in IIshki
  • ChatGPT in IIshki
  • Claude in IIshki
  • Canvas — wire AI models together on one board

Contents

  1. Which AI to use for a talking photo
  2. How to make a talking photo: 4 steps
  3. How long a talking photo should be, and how to write the script
  4. When Hailuo H3 beats an avatar
  5. Common talking-photo mistakes
  6. Where a talking photo is useful
  7. What's next

Models in this article

  • Kling Avatars v2
  • ElevenLabs Voiceover
  • MiniMax Voiceover
  • Hailuo H3
  • Nano Banana Pro

Try it in the studio

The models from this article are available in the IIshki studio, paid with tokens.

Open the studio