Gemini 3.1 Flash TTS: text narration and voice dialogues

Gemini 3.1 Flash TTS is a text-to-speech model by Google. It turns text into lifelike narration with a voice in 70+ languages, including Russian, and does one key thing — dialogues between two different voices in a single request. The catalog has 30 preset voices with male and female timbre, plus control over accent, style and pace. This is spoken narration and dialogues, not music with vocals. On Genosai narration costs 15 credits per 1000 characters, minimum 5.

Updated: July 28, 2026

Contents

What is Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS is a text-to-speech model by Google. Its task is simple and specific: turn text into lifelike narration with a voice. You paste a line, a script or a chapter — the model reads it as natural speech with correct intonation. This is text-to-speech: a narrator's voice, not music and not singing.

The main thing that sets this model apart from ordinary narration is dialogues. Gemini 3.1 Flash TTS can voice lines from two different speakers in a single request: for example, a male and a female voice in turn. You define two characters, write their lines, and get a finished scene with distinct timbres. This is handy for interviews, teaching conversations, ad dialogues and two-voice podcasts.

The model is multilingual: the voices work in more than 70 languages, including Russian. The catalog has 30 preset voices with male and female timbre — from confident and deep to warm and young. Delivery can be tuned separately: choose an accent, a style (for example narrator, whisper or empathetic) and the pace of speech. In one line the model accepts up to 10,000 characters, so it also suits long texts.

It is important to separate this model from music generators right away. Gemini 3.1 Flash TTS does not write songs and does not sing — it narrates ready text with a voice. If you need music with vocals, the Suno models are for that. Gemini solves the task of voiceover and dialogues: video narration, audiobooks, assistant lines, two-voice scenes. On Genosai narration costs 15 credits per 1000 characters, minimum 5.

Capabilities

Gemini 3.1 Flash TTS covers the task of narrating text with a voice — from a short line to a two-character dialogue. Here is what the model does in practice.

Two-speaker dialogues

The model's key capability is voicing a conversation. In one request you define two speakers, assign each a voice (for example female and male) and write their lines in turn. The model assembles a coherent scene where the voices sound different and do not blend. This is how you make interviews, dialogues in teaching clips, ad scenes and two-voice podcasts without recording two narrators.

Thirty voices with character

The catalog has 30 preset voices with male and female timbre. Among those verified in Russian are the confident female Kore, lively male Puck, deep male Charon, warm female Sulafat, soft male Algieba and young female Leda. Each voice sounds distinct, so you can pick a fitting timbre for a clip, an audiobook or an assistant — and two contrasting voices for a dialogue.

Delivery control

Beyond the voice, the model lets you tune how the speech sounds. The accent is chosen separately (neutral and several English variants), the style sets the manner — narrator reading, whisper, empathetic or businesslike delivery — and the pace adjusts speed from calm to fast. There is also a scene-context field where you can describe the setting and hint at the pronunciation of tricky words. This helps hit the right tone without editing.

Languages and long texts

The voices are multilingual and work in more than 70 languages, including Russian. The same script can be narrated in several languages — handy for localizing clips and multilingual products. In one line the model accepts up to 10,000 characters, so it handles a long monologue or a large scene without splitting it into dozens of small requests.

In practice the strengths show in three scenarios. First — single-voice video narration: a narrator reads your script. Second — dialogues and scenes: two voices hold a conversation. Third — audiobooks and podcasts: long text in natural speech. In every case the task is the same — turn text into a quality voice.

Examples

Every clip below was voiced by the model itself on Genosai — press play and listen right on this page. Each demo is under ten seconds, nothing to download.

One line, three moods

The voice is the same, Puck, and so is the text: "Today we are launching the thing we have worked on for a whole year." Only the delivery changes — and one line turns into a secret, then a joke, then an ad.

Puck, Whisper mood — Whisper style, The Drift pace

Puck, ironic mood — dry Deadpan delivery

Puck, announcer mood — Promo/Hype style, Rapid Fire pace

Male and female voice

Here it is the other way round: the notification text and the Newscaster style stay the same, only the timbre differs. That makes it easy to compare voices and pick the one that fits your project.

Kore — female voice, Newscaster style

Charon — male voice, same text and style

And this is a warm, caring delivery — a different voice and a different style, for support and tutorial scenarios.

Sulafat — warm female, Empathetic style

Two-voice dialogue

Both lines were recorded in a single generation: the first speaker got the Kore voice, the second one Puck. No separate requests and no stitching in an editor.

Dialogue: Kore and Puck, one generation

_Generated on Genosai.io with Gemini 3.1 Flash TTS._

How to use on Genosai

  1. Sign in to your Genosai account and open the section with the Gemini 3.1 Flash TTS model.
  2. Choose the mode: one voice for narration or two speakers for a dialogue.
  3. Assign a voice to each speaker — for example the female Kore and the male Puck.
  4. Paste the line text: for a dialogue, write the lines in turn for the first and second speaker.
  5. Optionally set the accent, style and pace to shape the tone you need.
  6. Run the generation, wait for the finished narration, and download the audio or edit the text and narrate again.

Genosai works in the browser and requires no installation. Credits are charged by the total number of characters across the lines, so a short phrase costs a minimum of 5 credits, while a long dialogue is charged in proportion to its length. This makes it convenient to voice both single phrases and whole scenes in one interface.

Prompts

For narration in Gemini 3.1 Flash TTS the "prompt" is the text a voice will read plus the choice of speakers. Below are ready templates for different scenarios. Paste your own text, pick a voice and a language.

Voice: Charon (deep male). Text: Chapter one. The morning was quiet, and only the wind carefully stirred the leaves outside the half-open window.
Voice: Sulafat (warm female). Text: Thank you for choosing our service. To continue, say your order number or say "operator".
Voice: Leda (young female). Style: narrator. Text: In this podcast episode we cover three habits that help keep focus throughout the day.
Dialogue. Speaker 1 — Kore: "Have you tried the new feature yet?" Speaker 2 — Algieba: "Yes, yesterday. Built a clip in a couple of minutes and was happy with it."
Voice: Puck. Style: promo. Text: Today only — a discount on all plans. Grab your subscription before the day ends.
Dialogue. Speaker 1 — Sulafat (mentor): "Where shall we start the lesson?" Speaker 2 — Puck (student): "Let's begin with the basics — how the interface works."
Voice: Kore. Text: Добро пожаловать на борт. Пристегните ремень и приведите спинку кресла в вертикальное положение перед взлётом.

Generation cost

On Genosai narration in Gemini 3.1 Flash TTS is priced by character count: 15 credits per every 1000 characters of text, minimum 5 credits per request. The total length of all lines is counted, including both speakers in a dialogue. A short phrase charges the minimum, while a large dialogue is charged in proportion to length: for example, 2000 characters cost 30 credits. This calculation is easy to plan ahead when you know the size of a script or scene.

Starter credits after sign-up let you try Gemini 3.1 Flash TTS for free, and top-ups work with Russian cards without a VPN. Current packages and balance are on the Pricing page.

Comparison

The main difference between Gemini 3.1 Flash TTS and the catalog's music models is the task. Gemini 3.1 Flash TTS narrates text with a voice and runs two-speaker dialogues, while the Suno models create music with vocals — songs with melody and arrangement. Its closest neighbor by task is ElevenLabs TTS: also voice narration, but single-voice and focused on speed. If you need a song, see Suno V5, its evolution Suno V5.5 or the economical Suno V4.5.

ModelTaskPrice on Genosai
Gemini 3.1 Flash TTSNarration and two-voice dialogues15 credits per 1000 characters
ElevenLabs TTSSingle-voice text narration12 credits per 1000 characters
Suno V5Music with live vocals16 credits per track
Suno V4.5Music with vocals, more economical14 credits per track

The choice logic is simple. You need a conversation between two characters or fine control over delivery — that is Gemini 3.1 Flash TTS. You need fast single-voice narration of a long text — look at ElevenLabs TTS. You need a song with music and singing vocals — those are the Suno models. All three tasks are easy to confuse by the word "voice", but the result differs: in one case a narrator or two speakers read text, in the other an artist sings over an arrangement. All models are available in one Genosai interface, so you can pick the right one for each task.

Limitations and tips

Gemini 3.1 Flash TTS narrates text but does not replace music: it cannot sing a song over an arrangement — the Suno models are for that. A dialogue is limited to two speakers per request: for a scene with three or more characters, split it into parts or assemble it from several narrations. Keep this in mind when choosing a model for the task.

Narration quality depends heavily on text preparation. Use punctuation: commas and periods set pauses, while question and exclamation marks set intonation. In a dialogue, separate the speakers' lines clearly so the voices don't get mixed up. For numbers, abbreviations and tricky names, it's better to spell out how they should sound so the model doesn't misread them. Pick the style and pace to match the scene: narrator for news, empathetic for a conversation, fast for dynamic ads.

Also remember the split of tasks within the catalog. Narrating text, a two-voice dialogue and a song with vocals are different results. Gemini 3.1 Flash TTS handles the first two: clean spoken narration and a two-speaker conversation from your text. If the script is mixed — for example a clip with a voiceover over a music bed — assemble it from parts: take the voice here and generate the music bed with a separate Suno model. That way each element sits at its own level and the final edit is cleaner.

A practical tip: plan the budget ahead — the price depends on the total character count, so a long scene is worth estimating before you run it. For Russian text, check tricky terms and abbreviations and spell them out if needed. And if it turns out you need a song rather than narration, switch to Suno V5 — that is music with vocals, a separate task.

FAQ

What is Gemini 3.1 Flash TTS?

It is a text-to-speech model by Google. It narrates text with a lifelike voice in 70+ languages, including Russian, and can voice two-speaker dialogues. This is spoken narration and dialogues, not music with vocals.

Can Gemini 3.1 Flash TTS voice a dialogue?

Yes, that is its key feature. In one request the model voices lines from two different voices — for example a male and a female one. You define two speakers and write their lines in turn, and the model assembles a finished scene with distinct timbres.

Does the model support Russian?

Yes. The voices are multilingual and work in more than 70 languages, including Russian. You paste Russian text, pick a voice, and the model narrates it as natural speech with correct intonation.

How much does narration cost with Gemini 3.1 Flash TTS on Genosai?

Narration costs 15 credits per every 1000 characters of text, with a minimum of 5 credits per request. The total length of all lines is counted, so a short phrase charges the minimum while a long dialogue is charged in proportion to the character count.

How does Gemini 3.1 Flash TTS differ from ElevenLabs TTS?

Both models narrate text with a voice, but Gemini 3.1 Flash TTS can voice two-speaker dialogues in one request and gives fine control over accent, style and pace. ElevenLabs TTS is fast single-voice narration with very low latency. For scenes with two voices Gemini is handier, for streaming narration ElevenLabs is.

How many voices are available and what are they?

There are 30 preset voices with male and female timbre. Among those verified in Russian are the confident female Kore, lively male Puck, deep male Charon, warm female Sulafat, soft male Algieba and young female Leda. A dialogue uses two different voices.

What is Gemini 3.1 Flash TTS good for?

For voiceovers of videos and clips, audiobooks and podcasts, assistant lines, and also for scenes and dialogues: interviews, teaching conversations, two-character ad dialogues. Control over accent, style and pace helps set the right tone.

Try Gemini 3.1 Flash TTS on Genosai