Gemini 3.8 Flash TTS is Google's speech model for directing character voices line by line, for audiobooks, games, and voice agents.
Flash, or the cheaper Flash-Lite
- Flash is the creative model, id gemini-3.8-flash-tts. Google lists 130 languages, for audiobooks, multi-speaker scenes, and regional accents. Input limit 8,192 tokens, output 16,384.
- Flash-Lite is the high-volume model, 101 languages. Google says it replaces gemini-3.1-flash-tts-preview.
- Both take the same request. They are in the Gemini API and Google AI Studio. Google says the enterprise API is still coming, and Google Vids is where everyone else gets them.
Stage directions are not part of the line
If the transcript says "Say cheerfully: Hello", the model may speak that instruction out loud. Sustained delivery goes in speech_metadata.style. A laugh or a pause stays inline. Each turn of a two-speaker scene needs its own speaker name. One request now returns a WAV file. Older Gemini TTS returned raw audio with no header, so a pipeline that adds its own WAV header would wrap the file twice.
A new voice, or a copy of 30 seconds
- Describe a voice in a sentence. Google's examples are a Melbourne DJ, a flat robot, and a Japanese dragon.
- About 30 built-in studio voices, plus a library Google puts at more than 2,000, including Mexican Spanish, Quebec French, and Scots English.
- A copy needs 30 seconds of audio and a spoken consent line from that same person. Every clip carries a SynthID watermark.
- In AI Studio, voice copy is not offered in Illinois, Texas, the EEA, the UK, Switzerland, or India.
The number on Google's chart
Google says Flash TTS is first on Hume AI's Voice Design Benchmark: overall 71.4, accents 60.8. The chart on the same post puts ElevenLabs Voice Design v3 at 70.8 overall and 45.4 on accents, and ahead on voice qualities, 76.6 to 74.6. Those figures are Google's.
The rate Google's price page does not list
The Gemini API pricing page does not list 3.8 TTS. OpenRouter lists Google AI Studio at $0.50 per million input tokens and $9 per million output tokens for Flash, and $6 output for Flash-Lite.