Gemini 3.8 text-to-speech says hello
By Jakub Antkiewicz
•2026-09-24T13:13:12Z
Google Launches Expressive Voice Generation with Gemini 3.8 TTS
Google has released two new text-to-speech (TTS) models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, making advanced audio generation capabilities available through its developer platforms. The models allow users to create entirely new voices from natural language prompts, replicate existing voices from short audio clips, and direct vocal performances line-by-line. This launch shifts the focus from selecting from a static library of voices to providing a dynamic toolset for creating custom audio for applications ranging from audiobooks and podcasts to real-time voice agents.
The two models serve distinct operational needs. Gemini 3.8 Flash TTS is engineered for deep creative control, enabling the design of specific character voices and granular direction of dialogue. In contrast, Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient scaling, targeting uses like large-scale dubbing and expressive conversational agents. According to Google, the models secured top spots on Hume AI’s Voice Design Benchmark and Overall Quality Index. Both are accessible starting today in the Gemini API and Google AI Studio.
Key Technical Capabilities
- Generative Voice Design: Create bespoke voices from scratch by describing characteristics like role and accent in natural language across over 100 languages.
- Voice Replication: Recreate a consistent vocal profile from a 30-second audio sample, requiring verbal consent and protected by SynthID watermarking.
- Performance Direction: Control pacing, emotion, dialect, and non-verbal sounds (e.g., laughs, sighs) using script cues.
- Long-Form & Multi-Speaker Support: Maintain vocal consistency across hours of audio and direct two-speaker conversations from a single script with natural turn-taking.
By integrating these TTS tools directly into the Gemini ecosystem, Google is positioning expressive voice as a core platform utility rather than a niche third-party service. This move directly competes with specialized AI voice synthesis companies by lowering the barrier for developers to build sophisticated audio experiences. The company's stated focus on built-in safety measures, including consent verification for voice replication and imperceptible SynthID watermarking on all generated audio, addresses industry-wide concerns about the potential for misuse and aims to establish a framework for responsible deployment in the market.
By embedding highly customizable text-to-speech capabilities directly into its core developer offerings, Google is strategically working to make expressive voice generation an integrated platform feature rather than a high-cost, specialized service, pressuring standalone AI voice startups.