Google unveiled two new text-to-speech (TTS) models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, on Wednesday, September 23. These models are designed for different purposes: the Flash TTS is aimed at creative applications like video games, audiobooks, and podcasts, while the Flash-Lite TTS is tailored for industrial uses such as dubbing, content production, and voice assistants. Both models are upgrades from the earlier Gemini 3.1 Flash TTS, launched in April 2026. Google highlights improvements in handling long texts and managing two-voice dialogues, as well as the introduction of a more affordable Flash-Lite version for generating large volumes of audio.
The Gemini 3.8 Flash TTS allows users to generate a unique voice from scratch using a simple text instruction, without needing an initial audio sample. Users can customize voices for specific roles, such as a narrator with a particular pace or a theatrical character, and adjust elements like intonation, speed, accent, and even vocal expressions like "mhm" or "yes." The model comes with a library of over 2,000 pre-built voices, covering more than 100 languages and dialects, including regional variations like Mexican Spanish or Scottish English. It is designed to maintain a consistent tone during long recordings, which is especially useful for audiobooks and podcasts. Additionally, it supports two-voice dialogues, allowing two distinct voices to interact naturally within the same script.
To ensure responsible use, Google states that a vocal profile can be recreated from a 30-second audio sample, provided it’s the user’s own voice or one for which they have rights. The company has integrated a consent verification process, along with SynthID watermarks and C2PA identifiers, which are standards for tracking digital content’s origin and changes. These features help make AI-generated speech detectable. However, detection effectiveness depends on the tools used by others, and audio files can be altered or republished in ways that may reduce the watermark's visibility. Google also plans to introduce a vocal remixing feature in the future, allowing users to modify existing voices using text instructions.
Both models are now available through the Gemini API and Google AI Studio. The Gemini 3.8 Flash TTS is accessible in Gemini Enterprise and Gemini Notebook, while the Flash-Lite TTS is set to be integrated into Google Vids, a platform for video content creation. These updates reflect Google’s ongoing efforts to enhance AI-driven audio generation for both creative and practical applications.
Google Unveils New Text-to-Speech Models with Enhanced Features and Controls
AI-rewritten from original reportingHow it works
googlettsaigeminivoice-synthesissynthid



