How to Use Grok Imagine: Image to Video With Sound (2026 Guide)
Grok Imagine is xAI's video model, and using it takes four moves: upload a photo as the first frame, describe the motion, the camera and the sound, pick a length up to 15 seconds and generate. Grok Imagine 1.5 turns a photo into video with native sound — street noise, rain, music — while the regular Grok Imagine can also make a clip from text alone, no image needed. Both run in the IIshki studio without an X account: 480p or 720p, five frame formats, pay-as-you-go tokens. Below: which version to pick, how to make a video in 5 steps, a prompt formula with sound and the mistakes that waste generations.
If you're still choosing a model for photo animation in general, start with the guide on how to animate a photo with AI, which compares Grok with Kling, Veo and Seedance. This page is only about Grok, and only about practice.
What Grok Imagine does well
The main reason to pick Grok is that sound is generated together with the picture. You don't hunt for sound effects and line them up on a timeline: you write "sound: raindrops on glass, distant thunder" in the prompt, and the clip arrives with that track already in it. The second reason is simplicity: there are no modes, references or shot lists, so even a beginner gets a usable result on the first try.
Strengths you can see in practice:
- Sound in the shot — ambience, action sounds, music. Version 1.5 only.
- Facial expression — the face in the photo reacts to the scene: a smile, surprise, focus.
- Physics — objects fall with weight, fabric and water move believably, though don't expect a perfect simulation.
- Camera — push-ins, pull-backs and orbits feel motivated, and framing keeps the subject centered.
What Grok can't do: scenes longer than 15 seconds, several character references, or precise lip sync to your own audio. Other models handle those — see the table below.
Which Grok Imagine version to choose
| Task | Model | Why |
|---|---|---|
| Animate a photo and get sound right away | Grok Imagine 1.5 | Native sound, any length from 1 to 15 s |
| A clip from text only, no image | Grok Imagine | Image optional, a prompt is enough; 6–15 s |
| Make the starting frame | Grok Image or Nano Banana Pro | Image in the right format, then into 1.5 as frame one |
| Precise motion and shot control | Kling 3.0 | Follows motion descriptions closely, supports multi-shot |
| Cinematic scene with sound in 1080p or 4K | Veo 3.1 | Higher resolution, complex scenes |
| A person saying your text out loud | Kling Avatars v2 | Portrait + audio → synced lips |
For when Grok beats the competition, see the comparisons Kling 3.0 vs Grok Imagine and Seedance 2.0 vs Grok Imagine.
How to use Grok Imagine: 5 steps
1. Prepare the starting frame
Version 1.5 requires a photo: JPEG, PNG or WebP up to 20 MB. Shots where the subject is large, not cut off by the frame edge and evenly lit work best. No image yet? Generate one in Grok Image in the same format you want for the video, so the model doesn't have to invent the margins.
2. Pick the format and quality
You get 16:9 for YouTube, 9:16 for Reels and Shorts, 1:1, 3:2 and 2:3. Match the format to the photo: upload a landscape shot and ask for 9:16, and the edges will have to be painted in. Quality is 480p for drafts and 720p for the final.
3. Set the length
In 1.5 you pick anything from 1 to 15 seconds. One action — turning a head, smiling, taking a step — needs 4–6 seconds. The longer the clip around a single simple action, the more likely the motion starts looping toward the end.
4. Write the prompt with the formula
The Grok formula: what moves + how the camera moves + sound. Don't describe what's already in the photo — the model sees the image and only needs to know what to do with it.
The woman in the photo slowly turns toward the camera and smiles, wind moving her hair. The camera gently pushes in. Sound: surf, gulls in the distance.
5. Run it and refine
Generation runs in the background: the result appears among your generations, and if you signed in with Telegram it is also sent to the bot as a file. Not happy? Change one thing at a time — motion, camera or sound — so you know which edit worked.
Grok Imagine prompts with sound: 8 examples
Put sound on a separate last line so the model doesn't mix it up with the action. Two or three sound details work better than a long list.
Portrait:
The man in the photo looks up from his book and smiles thoughtfully. The camera slowly pushes in on his face. Sound: a quiet ticking clock, pages rustling.
Old family photo:
The couple in the photo comes to life: the woman laughs and looks at the man, he adjusts his hat. A slight camera sway, like old newsreel footage. Sound: gramophone crackle, a soft waltz.
City:
The street in the photo comes alive: cars pass, people walk, a café sign blinks. Static camera. Sound: city hum, a car horn, snatches of conversation.
Nature:
The lake in the photo: ripples spread across the water, fog drifts over the surface, a bird flies past. The camera glides slowly to the right. Sound: lapping water, morning birdsong.
Product:
The perfume bottle in the photo slowly rotates, highlights slide across the glass, petals swirl around it. The camera orbits the bottle. Sound: soft electronic music, no vocals.
Food:
Steam rises from the coffee cup in the photo, a sugar cube drops in and ripples spread. Macro shot, static camera. Sound: a spoon clinking on porcelain, a faint fizz.
Pet:
The cat in the photo turns its head toward a noise, flattens its ears, then yawns. Camera at the cat's eye level. Sound: purring, a door closing somewhere.
Text-only clip (regular Grok Imagine, no photo):
A small robot waters a flower on a windowsill while snow falls outside. Warm evening light, cartoon style. The camera slowly pulls back.
For more camera moves, see Kling 3.0 prompts — the camera language there works in Grok too.
Common Grok Imagine mistakes
| Mistake | Cause | Fix |
|---|---|---|
| Video lacks the sound you wanted | Sound not described or mixed into the action | A separate "Sound: …" line at the end, 2–3 details |
| The face drifts toward the end | Clip too long for one action | Cut to 4–6 s or add a second action |
| The character does the wrong thing | Several storylines in one prompt | One video, one simple action |
| Strange painted-in edges | Video format doesn't match the photo | Match the photo's format or remake the frame in Grok Image |
| Moderation refusal | Celebrity, face swap, explicit or violent content | Use your own photo or ordinary people, describe the scene without names |
| 1.5 won't start without a photo | Version 1.5 requires a starting frame | Upload a photo or switch to the regular Grok Imagine |
| Long prompt gets cut | The limit is 4,096 bytes | Drop descriptions of what's already visible in the photo |
Moderation is checked before charging: if a run is blocked, no tokens are spent. If the model fails on the provider's side, tokens are refunded automatically.
What to do next
Finished Grok clips are easy to chain with other scenes on the Canvas: generations connect into one graph, so a frame from an image model flows straight into a video model without downloading and re-uploading files between tabs. If you need a longer scene with one character across several shots, move to Seedance or Kling — the best video models are collected in the AI video tools roundup.
Models in this article
Frequently asked questions
How do I use Grok Imagine without an X account?
Open Grok Imagine 1.5 or Grok Imagine in the IIshki studio, sign in with email or a social account, upload a photo, write a prompt and press Generate. You don't need an X account or a subscription: each run is paid in tokens, and the price is shown on the button before you start.
What is the difference between Grok Imagine 1.5 and Grok Imagine?
Grok Imagine 1.5 turns a photo into video with sound, and the starting frame is required. The regular Grok Imagine can also make a clip from text alone — an image is optional. If you need sound in the clip, use 1.5; if you have no picture and want a quick draft of an idea, use the regular version.
How long are Grok Imagine videos?
Up to 15 seconds. In the IIshki studio, Grok Imagine 1.5 lets you pick any length from 1 to 15 seconds in one-second steps, and the regular Grok Imagine runs from 6 to 15 seconds. Quality is 480p or 720p, with 16:9, 9:16, 1:1, 3:2 and 2:3 frames.
Is Grok Imagine free?
No. On IIshki every run costs tokens: the price depends on length and quality and is shown on the button before you start. If the provider returns an error, the tokens are refunded automatically. The cheapest way to work is short 480p drafts, switching to 720p for the final version.
Why won't Grok Imagine animate my photo?
Usually it's moderation: recognizable celebrities and public figures, face swaps, explicit and violent content are blocked before any tokens are charged. Your own selfie and photos of ordinary people pass. The other common cause is the file: the model accepts JPEG, PNG and WebP up to 20 MB.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Join the clubMore articles
- Seedance 2.5 Prompts: Formula and 15 Examples With ReferencesHow to write Seedance 2.5 prompts: @Image1, @Video1 and @Audio1 references, shot breakdowns up to 30 seconds, 15 ready-to-use examples and common mistakes.
- What people recommend for ElevenLabs Russian voiceoverWhat people recommend on forums for ElevenLabs Russian voiceover: stress marks, accent, engines and emotion tags — and how to apply each tip in IIshki studio.
- How to Dub a Video With AI in 2026: Translate, Voice, Lip-SyncHow to dub a video with AI: translate the script, voice it in ElevenLabs or MiniMax, and lip-sync the speaker in Kling Avatars v2. Step-by-step, with prompts.