logoSeailife
  • Submit Project
  • Pricing
  • Blog
Sign inSign up
Sign in
logoSeailife

© 2026 Seailife. All rights reserved.

Built with Open Launch - The first complete open source alternative to Product Hunt.

Powered by Open-LaunchPowered by Open-Launch

Discover

  • Trending
  • Categories
  • Submit Project

Resources

  • Pricing
  • Sponsors
  • Blog

Legal

  • Terms of Service
  • Privacy Policy

Connect

  • [email protected]
Gemini 3.8 Flash TTS Logo

Gemini 3.8 Flash TTS

Artificial IntelligenceDeveloper ToolsAPIs & Integrations
Visit
Gemini 3.8 Flash TTS - Product Image

Gemini 3.8 Flash TTS is Google's speech model for directing character voices line by line, for audiobooks, games, and voice agents.

Flash, or the cheaper Flash-Lite

  • Flash is the creative model, id gemini-3.8-flash-tts. Google lists 130 languages, for audiobooks, multi-speaker scenes, and regional accents. Input limit 8,192 tokens, output 16,384.
  • Flash-Lite is the high-volume model, 101 languages. Google says it replaces gemini-3.1-flash-tts-preview.
  • Both take the same request. They are in the Gemini API and Google AI Studio. Google says the enterprise API is still coming, and Google Vids is where everyone else gets them.

Stage directions are not part of the line

If the transcript says "Say cheerfully: Hello", the model may speak that instruction out loud. Sustained delivery goes in speech_metadata.style. A laugh or a pause stays inline. Each turn of a two-speaker scene needs its own speaker name. One request now returns a WAV file. Older Gemini TTS returned raw audio with no header, so a pipeline that adds its own WAV header would wrap the file twice.

A new voice, or a copy of 30 seconds

  • Describe a voice in a sentence. Google's examples are a Melbourne DJ, a flat robot, and a Japanese dragon.
  • About 30 built-in studio voices, plus a library Google puts at more than 2,000, including Mexican Spanish, Quebec French, and Scots English.
  • A copy needs 30 seconds of audio and a spoken consent line from that same person. Every clip carries a SynthID watermark.
  • In AI Studio, voice copy is not offered in Illinois, Texas, the EEA, the UK, Switzerland, or India.

The number on Google's chart

Google says Flash TTS is first on Hume AI's Voice Design Benchmark: overall 71.4, accents 60.8. The chart on the same post puts ElevenLabs Voice Design v3 at 70.8 overall and 45.4 on accents, and ahead on voice qualities, 76.6 to 74.6. Those figures are Google's.

The rate Google's price page does not list

The Gemini API pricing page does not list 3.8 TTS. OpenRouter lists Google AI Studio at $0.50 per million input tokens and $9 per million output tokens for Flash, and $6 output for Flash-Lite.

Comments

Publisher

Seailife

Seailife

Launch Date
2026-10-02
Platform
api
Pricing
paid

Sponsors

Become a Sponsor

Get your brand featured here