
Google DeepMind has launched Gemini 3.8 Flash TTS and Flash-Lite TTS, allowing creators to generate expressive, customizable voices for applications like audiobooks and games.
What was announced
Google DeepMind unveiled two new text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS. These models allow users to generate fully customizable voices, offering options to modify accents and emotional tone. This level of personalization aims to enrich audio experiences in various applications such as audiobooks, gaming, and podcasts.
Developers and content creators can utilize these models through platforms like Google AI Studio and Google Vids, facilitating a dynamic way to produce audio that sounds natural and expressive. The tools are designed to enhance user engagement, making it easier to create unique character voices and improve scene dialogue.
Limits and availability
While Gemini 3.8 offers a significant upgrade in text-to-speech technology, potential limitations include the challenge of ensuring these voices are used ethically. Google's implementation of safety features aims to address misuse, but the effectiveness of these safeguards remains to be seen.
The models are part of Google DeepMind's expanding Gemini Audio family and are currently accessible via the Gemini API and enterprise solutions, allowing businesses to leverage them for customized audio solutions.
Original source
This report summarises the source below. Analysis is labelled separately; product and research claims remain attributed to their source.
Read the original at Google DeepMind