* GOOGLE TEXT TO SPEECH VARIANTS
2 of 2* ABOUT GOOGLE TEXT TO SPEECH
Working with Google Text to Speech
Two speech models sit in this family. Gemini 3.1 Flash TTS is the expressive one: audio tags inside the text control pacing, tone, pauses and emphasis, so delivery is directed rather than left to the model. Google Text to Speech is the straightforward converter — pick a voice, submit text, get audio.
When to choose Google Text to Speech
Choose this family when the delivery of the line matters and you want to direct it inline. Kling Voice is for custom voices bound to Kling video; ByteDance Seed Audio is the option with sample-rate, speed and pitch controls.
Choosing between the variants
Gemini 3.1 Flash TTS is the one to use when a script needs emphasis, pauses or a change of tone mid-sentence. Google Text to Speech is simpler and faster to wire up when plain narration is all you need.
* FREQUENTLY ASKED
About Google Text to Speech
- What are audio tags in Gemini 3.1 Flash TTS?
- Inline markers in the submitted text that control pacing, tone, pauses and emphasis, so you direct the read instead of re-rolling the generation.
- Can I choose the voice?
- Yes. Google Text to Speech exposes a voice selection, so the same script can be rendered by different speakers.
- Which variant should I use for a long narration?
- Gemini 3.1 Flash TTS, because pacing and pause control keep a long read from sounding flat across paragraphs.
