How to Animate a Photo with AI (Image to Video): 2026 Guide
Animating a photo with AI — image to video — means uploading the photograph as the start frame of a video model and describing in words what should move. In 2026 this takes 2–3 minutes online: Kling 3.0 handles portraits best (controlled motion), Grok Imagine 1.5 and Veo 3.1 produce a clip with sound right away, and Seedance 2.5 makes long scenes up to 30 seconds. All four models run in the IIshki studio with token-based payment.
Which AI model to pick for animating a photo
| Task | Model | Why |
|---|---|---|
| Animate a portrait: smile, gaze, head turn | Kling 3.0 | Follows the motion description most precisely, flexible duration |
| Photo → video with sound right away | Grok Imagine 1.5 | Native audio, fast generation |
| Cinematic scene with speech in frame | Veo 3.1 | Built-in audio and speech, 16:9 and 9:16 |
| Long clip up to 30 seconds, several references | Seedance 2.5 | Image, video and audio references via @Image1, @Video1, @Audio1 |
| Reproduce a specific motion (dance, gesture) | Kling Motion Control | Transfers motion from a reference video onto the person in your photo |
If you don't know where to start, take Kling 3.0: it is the most predictable option for a single photo, and sound can be added separately when needed.
How to animate a photo with AI: 5 steps
1. Prepare the photo
The model works with what it sees. The cleaner the source, the fewer artifacts:
- Upload the original, not a screenshot or a compressed copy from a messenger.
- The face should be well lit and take up a noticeable part of the frame.
- One person in the frame animates more stably than a group.
- If the photo is old or small, run it through upscaling first — the model gets more detail to work with.
2. Pick the model and upload the photo as the start frame
Open the model in the studio and drop the photo into the “start frame” (or “image”) slot. For Kling 3.0 and Veo 3.1 this is the first frame of the future video; in Seedance 2.5 the photo can also be supplied as the reference @Image1 that you mention in the prompt.
3. Describe the motion, not the picture
The classic beginner mistake is describing what is already in the photo. The model can see it. The prompt should describe what changes:
Portrait: the person slowly smiles and turns the head slightly toward the camera, natural facial expression, soft light, static camera.
Formulas that work:
- Portrait — “slowly smiles”, “blinks and looks at the camera”, “wind gently moves the hair”.
- Landscape or city — “cars drive by, people walk”, “clouds drift slowly, water sparkles”.
- Camera — “slow push-in”, “static camera”, “slight parallax”. One camera move per clip.
For models with sound (Grok Imagine 1.5, Veo 3.1) add one phrase about audio: “street sounds”, “soft rain”, “background music without vocals”.
4. Set the duration and format
Stories and Reels — 9:16; YouTube and presentations — 16:9. Keep the duration minimal: 5 seconds for a portrait is almost always better than 10 — on longer clips the chance of distortion grows, and so does the price. The cost of a generation is shown on the model card before you start.
5. Run it and download the result
Generation takes from one to several minutes. The finished video is saved under “My generations” and, if Telegram is linked, arrives as a file in the bot. If the result isn't right, edit the prompt rather than the photo: usually it's enough to soften the motion or remove an extra detail from the description.
Common mistakes and how to fix them
| What went wrong | Cause | What to do |
|---|---|---|
| The face “melts” or changes | Motion too strong or low resolution | Soften the motion, upload the original, upscale if needed |
| Nothing moves | The prompt describes the photo, not the change | Rewrite the prompt with verbs: “turns”, “smiles”, “drift” |
| Extra objects in the frame | The model invented what isn't in the photo | Remove everything not visible in the shot from the prompt; add “static camera” |
| The clip is too short | Minimum duration selected | Increase the duration in settings or chain several clips in Canvas |
| No sound | Model without native audio | Use Grok Imagine 1.5 or Veo 3.1, or add a voice separately with ElevenLabs |
What's next
An animated photo is usually the first step. If the person in the photo should speak, look at Kling Avatar — a talking avatar from a photo and an audio track; if you need a voice — speech generation; and a comparison of video models by quality, sound and price is collected on the page best AI models for animating photos.
Models in this article
Frequently asked questions
Which AI model is best for animating a photo?
For controlled portrait animation — Kling 3.0: it follows the motion description most precisely. If you need a clip with sound right away — Grok Imagine 1.5 or Veo 3.1. For a long scene up to 30 seconds with several references — Seedance 2.5.
Can I animate a photo with AI for free?
In the IIshki studio generations are paid with tokens (1 token = 1 ₽), and new club members receive tokens for their first generations. Sign-up and photo upload are free; you only pay for the generation itself.
How many seconds of video do I get from one photo?
It depends on the model: Kling 3.0 and Grok Imagine 1.5 produce short clips of a few seconds, Seedance 2.5 goes up to 30 seconds. For social media 5–10 seconds is usually enough.
Why does the face get distorted in the animated photo?
Most often because the prompt asks for motion that is too strong, or the photo is low-resolution. Describe gentle motion (“slowly smiles”, “turns the head slightly”), upload the uncompressed original and don't ask the model to change the angle by 90 degrees.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Join the clubMore articles
- How to Dub a Video With AI in 2026: Translate, Voice, Lip-SyncHow to dub a video with AI: translate the script, voice it in ElevenLabs or MiniMax, and lip-sync the speaker in Kling Avatars v2. Step-by-step, with prompts.
- Kling Avatar: How to Make a Talking Photo with AI in 2026How to make a talking photo with AI: portrait plus voice in Kling Avatars v2, text-to-speech in ElevenLabs and MiniMax, in-scene dialogue with Hailuo H3.
- 20 Kling 3.0 Prompts for AI Video: Ready-to-Use ExamplesReady-to-use Kling 3.0 prompts for portraits, products, landscapes, multi-shot and sound, plus the prompt formula and common mistakes. Run them in IIshki.