Google DeepMind Announces Gemini 3.8 Flash TTS Models

At a Glance
| Item | Value |
|---|---|
| Publisher | Google DeepMind |
| Published | 2026-09-24 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
Google DeepMind has newly released the audio generation models “Gemini 3.8 Flash TTS" and “Gemini 3.8 Flash-Lite TTS". These are text-to-speech models capable of generating custom character voices and scene dialog direction, and are provided through Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Specifications
- Model Lineup: Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS
- Multilingual Support: Supports over 100 languages and dialects
- Voice Library: Provides over 2,000 production-ready voices
- Voice Cloning Feature: Can recreate a consistent voice profile from a 30-second audio sample
- Safety and Verification Features: Integrated consent verification, SynthID watermarking, and C2PA credentials
Performance and Quality
According to an announcement by the publisher Google DeepMind, Gemini 3.8 Flash TTS achieved overall 1st place (71.4) on Hume AI’s Voice Design Benchmark, and also took the top spot in accent modeling (60.8). Additionally, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS achieved 1st and 2nd place, respectively, in Hume AI’s Overall Quality Index.
Furthermore, in blind human preference evaluations on Voice Arena, they achieved top positions in major global languages such as Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish, and Hindi. It is reported that compared to Gemini 3.1 Flash TTS, significant improvements are seen across a wide range of use cases such as long-form content and two-speaker script control.
Strengths and Use Cases
Gemini 3.8 Flash TTS is a model specialized in deep creative direction and character design. By using natural language prompts, it is possible to design entirely new voices from scratch and bring characters to life in fields such as games, immersive audiobooks, podcasts, and interactive media. Specifically, it allows for fine-grained control line by line over elements such as acting directions, pacing, dialect variations, and backchanneling.
In contrast, Gemini 3.8 Flash-Lite TTS is optimized for cost-efficient processing at scale. It is suited for use cases such as high-volume dubbing operations, audio content production, and building expressive voice agents where tone, pace, and subtle nuances of expression can be controlled.
The main features and characteristics are as follows:
- Generative Voice Design: By specifying roles, accents, and vocal characteristics using natural language, you can design unique voices supporting over 100 languages and dialects. For example, you can generate the voice of a dramatic dragon or a narrator with a specific regional prosody.
- Voice Library and Cloning: In addition to accessing over 2,000 production-grade voices, it also features a voice cloning capability that recreates a consistent voice profile from a 30-second audio sample. Note that voice cloning incorporates a consent verification process using a spoken consent recording by the voice owner.
- Precise Performance Direction: Using stage directions and script cues, you can freely control acting quality from a calm customer service representative to a whispering suspense scene.
- Support for Long-Form Content: Even in continuous audio generation lasting several hours, the speaker’s vocal quality does not drift, maintaining a natural pace and high quality, making it ideal for podcast and audiobook production.
- Multi-Speaker Scene Composition: From a single script, you can seamlessly direct conversations between two speakers with natural turn-taking.
- Addition of Non-Verbal Textures: By incorporating vocal bursts such as
<laughs>,<sigh>, and<gasp>, and active listening backchanneling such as|mhm|and|yeah|into the script, you can produce realistic conversational textures.
Additionally, a “Voice Remix" feature is scheduled to be introduced soon, allowing prompt-based fine-tuning of tone, pitch, pace, and accent for voices within the library.
How to Get It
Developers can experience these audio generation features through the “Audio Playground" in Google AI Studio. In this workspace, it is possible to create a new voice identity from scratch via prompts or clone your own voice, and then try out line-by-line direction using the two-speaker script editor.
The rollout status for each model is as follows:
Gemini 3.8 Flash TTS
– For Developers: Gemini API and Google AI Studio
– For Enterprise: API provision via Gemini Enterprise scheduled to begin soon
– For General Users: Gemini Notebook
Gemini 3.8 Flash-Lite TTS
– For Developers: Gemini API and Google AI Studio
– For Enterprise: API provision via Gemini Enterprise scheduled to begin soon
– For General Users: Google Vids

