MiniMax Audio Text to Speech: How to Make an AI Voice-Over Sound Human
MiniMax Audio voice-over is text to speech — speech synthesis in Russian and other languages: the model reads your script in one of its preset voices, following punctuation, a chosen emotion and the pauses you place. In the IIshki studio MiniMax runs with no VPN: six engines (Speech 2.8 HD and Turbo, plus the older 2.6 and 02), eight Russian voices and an English narrator, 40 languages, three speeds and an MP3 or WAV file at the end. To make the voice sound human rather than like a GPS unit, you need to prepare the text for speech, pick the engine for the job and not overload the take with emotions. Below: the step-by-step process, five ready-made prompts and a table of mistakes.
What MiniMax does in the IIshki studio
MiniMax is one of two voice engines in the studio, next to ElevenLabs Voice. Its strength is even, clear speech on long scripts: tutorials, explainers, ads, narration for video. Everything you need sits in one form:
| Setting | What is available |
|---|---|
| Engine | Speech 2.8 HD (quality) and 2.8 Turbo (speed); 2.6 and 02 kept for compatibility |
| Voices | 8 Russian (4 male, 4 female) + an English narrator; you can also paste your own MiniMax voice ID |
| Language | Russian, English, Ukrainian and 37 more; Auto mode |
| Emotion | One per generation: happy, sad, angry, fearful, surprised, disgusted, neutral |
| Pauses | A pause tag in seconds right inside the text, <#> button in the editor |
| Speed | Slow, normal, fast |
| Text length | Up to 5,000 characters per run |
| Format | MP3 or WAV |
The studio offers MiniMax preset voices only — you cannot upload your own sample. If the job calls for a specific timbre, test several presets on the same paragraph: the gap between the "confident" and the "charming" male voice is wider than the names suggest.
Which engine and voice to pick
The universal rule: draft on Turbo, finish on HD. Turbo renders faster, so it is the place to check how the model reads numbers, abbreviations and stress. Once the script is clean, run the same text with the same settings through HD — voices and languages are shared, so the result will not drift, it will only get more detailed.
| Task | Model | Why |
|---|---|---|
| Narration for a tutorial or review | MiniMax, HD, "confident" male or "businesslike" female | Even delivery, reads lists and numbers cleanly |
| Ad, story, short clip | MiniMax, HD, "bright" female or "charming" male, emotion "happy" | Energetic pace without overacting |
| Audiobook, tale, narrative | MiniMax, HD, "dramatic" female, speed "slow" | Holds intonation across long sentences |
| Dialogue with sharp emotion shifts inside a line | ElevenLabs Voice, v3 engine | Emotions are inline tags, not a whole-take setting |
| English voice-over | MiniMax, "Narrator (EN)" or ElevenLabs Voice | Both have English voices; ElevenLabs has more of them |
| Talking head from a photo | Kling Avatars v2 | Takes a finished MP3 from 2 seconds to 5 minutes and animates a portrait |
If you are torn between the two voice engines, run the same paragraph through both — it takes less time than reading any comparison. Both models are collected in AI tools for voicing text.
How to use MiniMax for text to speech: 6 steps
Step 1. Rewrite the text for speech
Text written "for the eye" sounds like a report when spoken. Before you hit Generate, go through the script:
- Break long sentences into short ones — the model breathes on periods and commas; without them it drags the phrase on one breath.
- Write numbers, dates and amounts as words: "the fifteenth of March", "two thousand rubles". That removes misreads.
- Expand abbreviations or spell them the way they are said aloud.
- Keep Latin names in Latin script, but set the language manually — otherwise Auto may decide the whole text is English.
- Read the paragraph aloud. If you stumble, so will the model.
Step 2. Choose the engine
For the first pass take Speech 2.8 Turbo. Check stress and pace, fix the text. For the final version switch to Speech 2.8 HD — the settings stay, only the timbre quality changes. Engines 2.6 and 02 are there for compatibility: if an old project was voiced on them and you need one more line in the same sound, use them; otherwise you do not need them.
Step 3. Match the voice to the format
Voice names in the studio describe character: "confident", "charming", "childhood friend", "bold" for the male set; "bright", "businesslike", "dramatic", "bold" for the female set. Test on the hardest paragraph of your own script — the one with numbers, names and foreign words — not on a demo line. A voice that reads a greeting well may stumble on a spec list.
Step 4. Set one emotion
The emotion button in the editor inserts a tag like {happy} into the text. In MiniMax this is a setting for the whole generation: the model takes the first emotion and applies it to the entire fragment. Keep "neutral" for tutorials and reviews, "happy" for ads, "sad" or "surprised" on separate pieces of a story. To change the emotion during a clip, split the text and voice the parts separately.
Step 5. Place pauses and set the speed
The <#> button inserts a pause — half a second by default, and the number can be edited. Put pauses where a live narrator would take a breath: before a key line, between list items, after a question. Do not place two pauses back to back with no text between them — the second one may be dropped. Pick the speed by format: "slow" for tutorials and stories, "normal" for most jobs, "fast" for punchy stories where the text has to fit the timing.
Step 6. Generate and listen through
Listen to the whole take, not the first ten seconds. Flaws usually hide in the middle: wrong stress on a rare word, a number read too fast, a pause in the wrong spot. Fix the text and re-voice only that piece — which is exactly why the script should be cut into 1,000–2,000-character chunks. The finished file comes as MP3 or WAV: download it, get it in Telegram, or feed it straight into Kling Avatars v2 to get a talking portrait.
Five MiniMax prompts for a natural-sounding voice
A voice-over prompt is the text itself plus emotion and pause tags. Here are five templates for different formats; swap in your own content and keep the structure.
1. Ad, energetic delivery — "bright" female voice, HD, speed "normal":
{happy} The new collection is in store. <0.4> Light fabrics, a relaxed cut and colours that survive the summer. <0.6> Come in — fitting is free.
2. Tutorial, calm explanation — "confident" male voice, HD, speed "slow":
{neutral} To place an order, open your cart. <0.5> Check the delivery address, <0.3> then choose a payment method. <0.5> The confirmation arrives by email within a minute.
3. Story, atmospheric narration — "dramatic" female voice, HD, speed "slow":
{sad} In a small town by the sea lived an old watchmaker. <0.8> Every evening he wound the clock on the tower, <0.4> though he knew no one listened to it anymore.
4. Stories and reels, conversational tone — "charming" male voice, Turbo, speed "fast":
{happy} Okay, honestly: <0.3> I did not think this would work. <0.5> But look at the result.
5. Brand voice, restrained premium — "businesslike" female voice, HD, speed "normal":
{neutral} We do one thing <0.3> and we do it well. <0.6> The rest is detail.
Notice that each prompt has a single emotion, pauses sit only between meaningful blocks, and sentences are short. Those are the three rules that separate a "live" voice-over from a robotic one.
Common mistakes and how to fix them
| Mistake | Cause | What to do |
|---|---|---|
| The voice sounds flat | Long sentences with no punctuation | Break into short lines, add periods and <0.5> pauses |
| Numbers and dates are misread | Digits in the text | Write them out as words |
| Accent on native words | Auto language flipped to English because of Latin characters | Set the language manually |
| The emotion "broke" mid-text | Several emotion tags — the model takes the first | One tag per generation; split the text to change emotion |
| A pause did not happen | Two pause tags in a row with nothing between | Keep one pause or separate them with a phrase |
| The voice overacts | "Angry" or "surprised" over the whole text | Use "neutral"; keep emotions for short pieces |
| Pace drifts at the end of a long file | A single 5,000-character chunk | Cut into 1,000–2,000 characters and generate in parts |
| The final sounds worse than the draft | Draft and final on different engines with different voices | Change only the engine Turbo → HD, keep voice and text |
What to do with the finished voice-over
A voice track rarely lives alone — it is usually one layer of a clip. Next to MiniMax the IIshki studio has the other layers: ElevenLabs Sound Effects makes footsteps, rain or city noise from a description, ElevenLabs Music a background track for a mood and length, and Kling Avatars v2 turns a photo plus your MP3 into a talking character. How to stack these layers under video is in how to voice a video with AI. If you would rather run the whole pipeline in one place — voice, sound, image and video as nodes on one screen — that is Canvas.
What next
Start with one paragraph: rewrite it for speech, run it on Turbo, fix it, switch to HD. Once the voice is right, voice the rest with the same settings. Then compare with ElevenLabs Voice on the same text: its emotions are inline tags inside a line, and for dialogue that sometimes decides it. Both models are waiting in the IIshki studio; the token price is on the card before you start.
Models in this article
Frequently asked questions
Does MiniMax speak Russian without an accent?
Yes. In the IIshki studio MiniMax has eight Russian voices — four male and four female — and you can pin the language to Russian instead of Auto. An accent creeps in when the text has a lot of Latin characters and the model guesses the language; with Russian set manually it does not happen.
What is the difference between the HD and Turbo engines?
HD takes longer but gives a richer timbre and holds intonation better — use it for the final take. Turbo is faster and sounds simpler — use it for drafts and for checking how the script reads. Voices and languages are identical, so switching engines never breaks anything.
Can I change the emotion mid-text?
No. In MiniMax the emotion is a setting for the whole generation, not an inline tag: the first emotion is applied, the rest are ignored. If a clip needs to move from calm to excited, split the script in two and voice each part with its own emotion.
How much text can I voice at once?
Up to 5,000 characters per generation — roughly five to six minutes of speech. Cut long scripts into meaningful chunks of 1,000–2,000 characters: it is easier to re-voice one paragraph without touching the rest, and the voice stays consistent from chunk to chunk.
Can I use MiniMax text to speech for free, and what does it cost?
Signing up and setting up a voice in the studio are free; you only pay for the generation itself, in tokens. The token price depends on the length of the text and is shown on the model card before you start. Tokens can be bought without a subscription; iishki club members get the club rate.
Every model in this roundup is available in the IIshki studio at reduced club prices. Club members get tokens for generations every month.
Join the clubMore articles
- How to Dub a Video With AI in 2026: Translate, Voice, Lip-SyncHow to dub a video with AI: translate the script, voice it in ElevenLabs or MiniMax, and lip-sync the speaker in Kling Avatars v2. Step-by-step, with prompts.
- Kling Avatar: How to Make a Talking Photo with AI in 2026How to make a talking photo with AI: portrait plus voice in Kling Avatars v2, text-to-speech in ElevenLabs and MiniMax, in-scene dialogue with Hailuo H3.
- 20 Kling 3.0 Prompts for AI Video: Ready-to-Use ExamplesReady-to-use Kling 3.0 prompts for portraits, products, landscapes, multi-shot and sound, plus the prompt formula and common mistakes. Run them in IIshki.