Kling AI 3.0: How to Use It — Settings, Sound and Multi-Shot
Kling 3.0 (Kling AI) is a video model that turns an uploaded frame into a clip of 3 to 15 seconds, following the motion described in the prompt, and generates sound together with the picture. In the IIshki studio it renders at 720 or 1080, in 16:9, 9:16 and 1:1, with a “Multi-shot” mode of up to 5 scenes and a transition between a first and a last frame. Below: how to use it, how to write a prompt the model actually follows, and where people most often trip up.
What Kling 3.0 can do in the IIshki studio
Kling 3.0’s main strength is controllable motion: the model follows the description of what happens in the frame and how the camera behaves. That is why it is the usual pick when you need to animate a specific photo — a portrait, a product, a location — and get a predictable result rather than a random scene.
| Parameter | What is available |
|---|---|
| Input | First frame (required) + last frame (optional) |
| Duration | 3–10 s for a single shot, up to 15 s in multi-shot |
| Quality | 720 (faster and cheaper) or 1080 (more detail) |
| Aspect ratio | 16:9, 9:16, 1:1 |
| Sound | Generated with the video, can be switched off |
| Multi-shot | Up to 5 scenes, your own shots or auto mode |
| Prompt | Up to 2,500 characters, up to 500 per scene in multi-shot |
Note the first row: in the studio Kling 3.0 works from an image only. That is not a limitation but how it works: the frame sets the character, light and composition, the prompt sets the motion.
How to use Kling 3.0: 6 steps
1. Prepare the start frame
The model animates what it sees. Upload a clean, uncompressed image with one main subject and a clear composition. If the source is small or old — upscale it first. If you have no frame, generate one in Nano Banana Pro or GPT Image 2 and pass it to Kling.
2. Describe the motion, not the picture
Anything already in the frame doesn’t need describing — the model can see it. The prompt answers “what changes over these seconds”:
The woman in the frame slowly turns her head toward the camera and smiles, a light breeze moves her hair, static camera, soft daylight.
3. Set the camera move — one per clip
Kling understands cinematography terms well. One move per shot, named explicitly: “slow push-in”, “camera pulls back”, “pan left to right”, “camera follows the character”. Two moves at once is a reliable way to get a shaky picture.
4. Pick duration and quality
5 seconds is the working standard for a single action; take 8–10 when something develops in the frame (the character stands up, walks over, turns around). 720 is for drafts and social media, 1080 for the final clip. The price grows with both duration and quality — it is shown on the card before you start.
5. Decide whether you need sound
The “With sound” switch adds atmosphere and ambient noise matched to the scene. If you plan to add music or a voice (ElevenLabs, MiniMax), turn it off — a clean track is easier to edit.
6. Run and review
Generation takes a few minutes; the result is saved under “My generations” and, if Telegram is linked, arrives as a file in the bot. Not happy — change the prompt, not the frame: nine times out of ten it’s enough to soften the motion or remove an extra detail.
How to write a prompt for Kling 3.0
This section covers only the structure. Twenty ready-made examples for portraits, products, landscapes and multi-shot are in a separate article: Kling 3.0 prompts.
A prompt structure that works — four blocks, in this order:
- Who does what — subject and action in verbs: “a barista pours milk into a cup”.
- Camera — one move: “slow push-in on the cup”.
- Light and mood — “warm morning light from the window, light steam”.
- Sound (if enabled) — “coffee shop noise, clinking cups”.
A barista pours milk into a latte cup, drawing a pattern; slow push-in on the cup; warm morning light from the window, steam over the cup; quiet coffee shop noise.
A sports car parked on a wet night street; the camera slowly arcs around it; neon reflections on the body, light rain; the sound of rain and a distant city.
A cat on a windowsill turns to the window and watches a bird; static camera; soft daylight; silence, a faint rustle.
Camera moves Kling understands
| How to write it | What you get | When to use it |
|---|---|---|
| “Pan left to right” | The camera rotates in place | Show a location, a landscape |
| “Slow push-in” / “camera pulls back” | Moving closer or further away | Emphasise the subject / reveal the scene |
| “Camera follows the character” | Tracking the movement | The character walks, drives |
| “Camera rises” / “descends” | Vertical crane | Architecture, scale |
| “Camera arcs around” | A half-circle around the subject | Product, car, portrait |
| “Static camera” | Fixed tripod | Portraits, facial expression, small movements |
What to avoid in the prompt: negations (“no blur” reads as “blur”), several camera moves at once, descriptions of things not in the frame, and abstract words like “beautiful” or “epic” — they say nothing about motion.
Multi-shot: a clip from several scenes
“Multi-shot” mode assembles up to 5 scenes into one clip of up to 15 seconds. Two ways:
- Your own shots — you write the prompt and duration for each scene (up to 500 characters and up to 12 seconds per scene). This is how ad clips and storytelling are made: “wide shot — detail — reaction”.
- Auto mode — you write one overall prompt and the model builds the storyboard itself. Faster, with less control; good for first tries.
Tip: keep one character and one location for the whole clip — that way the scenes cut together without the character “jumping”.
First and last frame
If you also upload a last frame, Kling builds a transition between the two images: morning → evening, sketch → finished product, a face → the face after make-up. Both frames should match in composition and angle — otherwise the model “plays out” the transition as an abrupt rebuild of the scene. To reproduce someone else’s choreography or gesture exactly there is a separate mode — Kling Motion Control, which transfers motion from a reference video.
Common mistakes and how to fix them
| What went wrong | Cause | What to do |
|---|---|---|
| The face “melts”, changes | Motion too strong or a small source image | Soften the motion, upload the original, upscale if needed |
| Nothing moves | The prompt describes the frame, not the change | Rewrite with verbs: “turns”, “walks”, “pours” |
| The picture jitters | Two camera moves in one prompt | Keep one |
| Extra objects appear in the frame | The model invented what isn’t in the photo | Remove everything not visible from the prompt; add “static camera” |
| The clip cuts off mid-action | Not enough time for the action | Increase duration to 8–10 s or split into multi-shot |
| The sound doesn’t fit | The scene was described without audio | Add one phrase about sound, or turn it off and voice it separately |
Kling 3.0 or another model
- You need sound with speech and a cinematic picture from text alone — Veo 3.1; see the detailed Veo 3.1 vs Kling 3.0 comparison.
- You need a long clip of up to 30 seconds with several references — Seedance 2.5.
- You need to animate a frame with sound quickly and cheaply — Grok Imagine 1.5; comparison with Kling.
- You need precise control over the motion of your own image — Kling 3.0, and that is its main use case.
What’s next
Start with one portrait, 5 seconds, 720 and a two-sentence prompt — and see how the model reads your description. Then add camera motion, sound and multi-shot. If the task is specifically to animate a photograph, the step-by-step breakdown with model choice is in how to animate a photo with AI, and the roundup of video models for different jobs is on best AI models for video.
Models in this article
Frequently asked questions
Do I have to upload a photo for Kling 3.0?
Yes. In the IIshki studio Kling 3.0 works from a start frame: you upload an image and the prompt describes what should move in it. There is no text-only generation for this model in the studio — for video from text alone, use Veo 3.1 or Seedance 2.5.
Does Kling 3.0 generate sound?
Yes, sound is enabled with the “With sound” switch and is generated together with the video. If the clip will go under your own music or voice-over, turn sound off — a clean track is easier to edit.
How many seconds of video does Kling 3.0 make?
A single shot is 3 to 10 seconds. In “Multi-shot” mode the clip is assembled from several scenes (up to 5) with a total length of up to 15 seconds, each scene no longer than 12 seconds.
Does Kling 3.0 have a 4K option?
In the IIshki studio Kling 3.0 renders at 720 or 1080. If you need 4K, take Veo 3.1 — it has that tier. The difference between 720 and 1080 in Kling is detail and price; for social media 720 is usually enough.
How much does a Kling 3.0 generation cost?
The price in tokens is shown on the model card before you start and depends on duration and quality (720 or 1080). IIshki club members get the club price.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Join the clubMore articles
- How to Dub a Video With AI in 2026: Translate, Voice, Lip-SyncHow to dub a video with AI: translate the script, voice it in ElevenLabs or MiniMax, and lip-sync the speaker in Kling Avatars v2. Step-by-step, with prompts.
- Kling Avatar: How to Make a Talking Photo with AI in 2026How to make a talking photo with AI: portrait plus voice in Kling Avatars v2, text-to-speech in ElevenLabs and MiniMax, in-scene dialogue with Hailuo H3.
- 20 Kling 3.0 Prompts for AI Video: Ready-to-Use ExamplesReady-to-use Kling 3.0 prompts for portraits, products, landscapes, multi-shot and sound, plus the prompt formula and common mistakes. Run them in IIshki.