How to Clone a Video with AI: Recreate a Clip with Your Own Character
Cloning a video with AI means generating a new clip that repeats an existing one — same moves, same angle, same pacing — but with your own character in the frame. That's how people make their version of a dance trend, put a brand mascot in place of an actor, or create a digital double that reads a script. In 2026 the most reliable way to transfer motion from someone else's clip is Kling Motion Control; Seedance 2.5 and Seedance 2.0 recreate a whole scene from a video reference, and Kling Avatars v2 turns a single portrait into a talking clone. All of them run in the IIshki studio with token-based pricing.
What “clone a video” actually means: two different requests
When people say they want to clone a video, they usually mean one of two things:
- Recreate someone else's clip with your own character. There's a viral dance, a challenge or a great ad — you want the same thing with your face, your brand mascot or a drawn hero. That's a motion-transfer task: the model reads the choreography from the source and applies it to an image.
- Make a clone of yourself for video. You don't want to sit in front of a camera every time — you need a double that says any text with your face. That's a talking-avatar task: portrait plus audio.
Different models solve these, and mixing them up is the first mistake. The table below helps you choose.
| Task | Model | Why |
|---|---|---|
| Repeat a dance, gestures, facial expressions on your character | Kling Motion Control | Purpose-built mode: the video sets the motion, the photo sets the look |
| Recreate the whole scene: angle, camera move, style, pacing | Seedance 2.5 | Up to 10 video and 30 image references, @Video1/@Image1 right in the prompt, up to 1080p |
| Same thing cheaper and shorter, for a draft | Seedance 2.0 | Up to 3 videos and 9 photos, Mini plan for quick tests |
| Recreate a clip with scene sound and speech | Wan 3.0 | Photo, video and audio references, sound generated together with the video |
| An AI clone of yourself that reads a script | Kling Avatars v2 | Portrait + audio from 2 seconds to 5 minutes, lip sync and facial expressions |
| Animate a single photo without a motion sample | Kling 3.0 | Motion is described in words, not video — see the photo animation guide |
Not sure? Start with Kling Motion Control: two upload slots, almost no settings, predictable output.
How to recreate someone else's clip with your character: 5 steps in Kling Motion Control
Kling Motion Control takes two files — a video with the motion and an image of the character — and returns a clip where your hero repeats everything the person in the source did: steps, body turns, hand gestures, lip and eye movement. A prompt is optional here.
1. Pick the reference video
The reference is 80% of the result. Model requirements: MP4 or MOV, 3 to 30 seconds, up to 100 MB. Common-sense requirements:
- One person in the frame. With two or more, the model can't tell whose motion to transfer.
- The body fully visible or at least waist-up, nothing in front of it. Tables, mic stands and objects passing by break tracking.
- A nearly static camera. Fast pans and cuts inside the clip get in the way: trim the source to a single shot.
- Even lighting and a simple background. The silhouette must read clearly.
Dancing, walking, on-camera gesturing and sports moves transfer best. If the reference is longer than 30 seconds, cut it into parts and generate them one by one.
2. Prepare the character photo
The image is jpg or png up to 10 MB. The main rule: framing and pose in the photo must match the first frame of the video. If the person in the clip stands full-length and your photo is a chest-up shot, the model will invent the legs — and they will look like someone else's.
- Face and hands must be visible: hands in pockets or behind the back are a source of extra fingers.
- Leave air around the figure: a character touching the frame edges gets cropped on wide moves.
- A body type close to the person in the source gives fewer distortions.
- Small or compressed photos should go through upscaling first.
The character doesn't have to be a human: a mascot, a 3D hero or a cartoon character works too, as long as the limbs are distinguishable.
3. Upload both files
Kling Motion Control has two slots: “Character” for the photo and “Motion reference” for the video. Upload originals, not compressed copies from a messenger.
4. Choose the orientation: “By video” or “By image”
This is the one setting worth understanding:
- By video — up to 30 seconds of motion are transferred; the angle and the character's position in the frame come from the clip. Best for dances and anything where the full choreography matters.
- By image — the model keeps the composition and angle of your photo; up to 10 seconds of motion are transferred. Best when the framing of the photo matters more than an exact copy of the source — for example, animating a portrait with gestures.
5. Quality and prompt
“std” mode is faster and cheaper, “pro” gives more detail — use it for the final. You can skip the prompt: the motion comes from the video anyway. But if you want to change the background or lighting, describe exactly that — not the motion:
Same character, night street with neon signs, wet asphalt, soft rim light.
The price of Kling Motion Control depends on the reference length and the mode — the exact number is shown on the model card before you start.
When Seedance 2.5 is the better choice: cloning the scene, not the motion
Motion Control copies choreography literally. Sometimes you want to repeat not the poses but the delivery: camera movement, tempo, editing style, atmosphere. That's a job for Seedance 2.5 in “References” mode: upload up to 10 videos, up to 30 photos and up to 10 audio tracks, and refer to them in the prompt as @Video1, @Image1, @Audio1.
Repeat the camera movement and pacing from @Video1 for the character from @Image1: they walk down the same corridor and turn around at the end, warm sunset light.
Scene as in @Video1, but the hero is the person from @Image1, outfit from @Image2, music @Audio1, 9:16 format.
Seedance 2.5 has a “Video task” switch: “Reference” — the clip sets the motion or style; “Edit” — modify the source clip while keeping its length and format; “Extend” — the model continues the clip. For cloning you need the first one. Resolution: 480p, 720p or 1080p.
Seedance 2.0 does the same on a smaller scale: up to 3 videos and 9 photos, clips up to 15 seconds, 720p max, but it has a Mini plan for cheap drafts. A sensible workflow is to tune the prompt and references on 2.0 Mini, then render the final on 2.5 in 1080p. How these models differ from Kling in regular generation is covered in the Seedance 2.0 vs Kling 3.0 comparison.
If you need the clip with scene sound right away — footsteps, speech, street noise — try Wan 3.0: it takes up to 5 video references and generates sound together with the picture, with no separate charge for audio.
How to make an AI clone of yourself for video
The second meaning of the request is a double that speaks for you. In the studio that's Kling Avatars v2: one portrait and one audio track in, a video with your face, synced lips, facial expressions and light camera movement out.
1. Shoot or pick a portrait
Face straight on or in a slight three-quarter turn, even light, no glare on glasses, no hand near the face. Format jpg/png, at least 300 px on a side, up to 10 MB. One good portrait is the whole “training”: you don't need to record minutes of yourself on camera the way classic avatar services require.
2. Prepare the speech
Audio: mp3, wav, m4a or aac, from 2 seconds to 5 minutes, up to 5 MB. Two ways:
- Record it yourself on a phone in a quiet room — then the voice is truly yours.
- Generate it from text in ElevenLabs or MiniMax: write the script, choose a voice, download the track and feed it to the avatar. This way your double reads any text without a recording — details in the MiniMax voice-over guide.
3. Generate and check
Upload the portrait to the “Avatar portrait” slot and the audio to “Audio”, choose “std” for a test or “pro” for the final. Check that pauses in the speech match pauses in the face: if the audio starts with silence, the avatar just hangs there — trim the beginning. Other talking-face models are in the lip-sync tools roundup.
One rule for both scenarios: clone yourself, your brand or your own drawn character. Putting another person's face into a cloned video without their consent is not okay.
Common mistakes when cloning a video
| Mistake | Cause | What to do |
|---|---|---|
| The body “melts”, extra arms appear | Several people in the reference, or objects covering the figure | Use a single-person source with a clean frame, or crop the shot |
| The face doesn't look like the photo | Scale of the photo and the first video frame don't match | Shoot the portrait with the same framing and pose as the clip's start |
| Arms or head get cut off on wide moves | The figure touches the frame edges in the photo | Leave empty space around the character |
| The clip came out 10 seconds instead of 30 | “By image” orientation selected | Switch to “By video” — it transfers up to 30 seconds |
| Motion transferred, but the background is random | Empty prompt, the model invented the scene | Describe background and lighting in the prompt, not the motion |
| Seedance ignored the video reference | No @Video1 in the prompt, or “Extend” task selected | Reference it explicitly and set the task to “Reference” |
| The avatar opens its mouth out of sync | Long silence at the start of the audio, or background music | Trim the silence, use a clean voice without a bed |
| Generation refused for the uploaded video | Reference longer than 30 seconds or heavier than 100 MB | Split the clip, compress it under the limit |
What's next
Cloning a video rarely ends with one generation. A typical chain: upscale the source photo → motion transfer in Kling Motion Control → voice in ElevenLabs or MiniMax → a talking version in Kling Avatars v2. You can build that chain in one window on the Canvas: each model is a node, and the output of one step goes straight into the next.
If the task is bigger than repeating someone's clip — say, shooting your own scene from scratch — see the best AI video models roundup and the best tools for animating photos. And if you need to animate a single photo without a motion sample, the photo animation guide covers prompts and typical mistakes.
Models in this article
Frequently asked questions
Which AI clones a video and swaps in a different person?
For an exact copy of the movement, use Kling Motion Control: upload the source clip and a photo of your character, and the model transfers the dance, gestures and facial expressions onto them. If you need to copy the whole scene — camera angle, pacing, style — use Seedance 2.5 with a @Video1 reference.
Can I clone a video with AI for free?
Signing up and uploading files to the IIshki studio is free; you only pay for the generation itself, in tokens. Kling Motion Control is priced by the length of the reference video, Seedance by duration and resolution. The exact price is shown on the model card before you start.
How long should the reference video be?
Kling Motion Control accepts a reference of 3 to 30 seconds (MP4 or MOV, up to 100 MB). In “By video” mode up to 30 seconds of motion are transferred; in “By image” mode up to 10. Seedance 2.5 takes up to 10 video references, Seedance 2.0 up to 3.
How do I make an AI clone of myself that reads a script?
Use Kling Avatars v2: upload a portrait (jpg/png, at least 300 px) and an audio track with speech, 2 to 300 seconds long. Record the voice yourself or generate it from text in ElevenLabs or MiniMax, then feed the file to the avatar — you get a talking double with lip sync.
Why does the character's face get distorted during motion transfer?
Usually because of a scale mismatch: the photo is a waist-up shot while the reference shows a full body, so the model has to invent the rest. Use a photo with the same framing and pose as the first frame of the clip, with arms and legs visible.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Join the clubMore articles
- How to Dub a Video With AI in 2026: Translate, Voice, Lip-SyncHow to dub a video with AI: translate the script, voice it in ElevenLabs or MiniMax, and lip-sync the speaker in Kling Avatars v2. Step-by-step, with prompts.
- Kling Avatar: How to Make a Talking Photo with AI in 2026How to make a talking photo with AI: portrait plus voice in Kling Avatars v2, text-to-speech in ElevenLabs and MiniMax, in-scene dialogue with Hailuo H3.
- 20 Kling 3.0 Prompts for AI Video: Ready-to-Use ExamplesReady-to-use Kling 3.0 prompts for portraits, products, landscapes, multi-shot and sound, plus the prompt formula and common mistakes. Run them in IIshki.