Skip to main content
Google Flash TTS interface: user designing AI voice with waveform, settings, and text input fields.

Editorial illustration for Google's Flash TTS Lets Users Design AI Voices From Scratch

Google Flash TTS Lets Users Design AI Voices

4 min read

Google added two text-to-speech models to its Gemini lineup this week, Gemini 3.8 Flash TTS and Flash-Lite TTS, both built to generate speech in more than 100 languages. The bigger shift is in how voices get made. Instead of picking from a preset list, users can type a description of a voice, its accent, its role, its general vibe, and the model builds it from that text alone. There's also a cloning tool that takes a 30-second audio clip and turns it into a usable voice profile.

Google is splitting the two models by use case. Flash TTS is pitched at people making things like audiobooks, podcasts, and game dialogue, where voice personality matters. Flash-Lite TTS trades some of that flexibility for cost and speed, aimed at businesses that need dubbing, voice agents, or narrated audio at volume rather than one-off performances.

Both models let creators write stage directions into individual lines, support two-voice conversations, and can insert nonverbal sounds like laughing or sighing. The rollout starts through the Gemini API and Google AI Studio, with Gemini Enterprise API access set to follow.

Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS, two new models for speech generation. Flash TTS can create new voices from text descriptions, and both models support more than 100 languages and let users add stage directions to individual lines of dialogue.

Why this matters

For developers building voice products, the 30-second cloning window and text-prompt voice design shrink a process that used to require studio time and licensing deals into something you can prototype in an afternoon. That's the real story here, not the 100-plus language count Google likes to lead with. Founders pitching audiobook or game-character startups now compete against a free-ish baseline from Google itself, which changes the calculus on what's worth building versus what's worth wrapping around an API.

Researchers should pay attention to the split between Flash TTS and Flash-Lite: Google is explicitly separating "creative, expressive" generation from "low-latency, cheap" generation, and that product split will likely shape how voice AI gets priced across the industry. We'd flag the obvious open question nobody at Google has answered yet: consent and misuse safeguards around cloning real voices from short samples. Google says the feature builds voice profiles from audio clips; it hasn't said much publicly about how it prevents someone from cloning a voice they don't own.

Common Questions Answered

How does Google's Flash TTS voice design feature work compared to traditional preset voice selection?

Instead of choosing from a preset list of voices, users can now type a text description of a voice including its accent, role, and general vibe, and the model generates a custom voice from that description alone. This represents a significant shift in how AI-generated voices are created, eliminating the need to select from limited predefined options.

What is the 30-second audio cloning tool in Gemini 3.8 Flash TTS and how does it work?

Google's Flash TTS includes a cloning tool that takes a 30-second audio clip and converts it into a usable voice profile that can be used for text-to-speech generation. This allows users to quickly replicate and customize existing voices without extensive studio recording or licensing requirements.

How many languages do Gemini 3.8 Flash TTS and Flash-Lite TTS support?

Both Gemini 3.8 Flash TTS and Flash-Lite TTS are built to generate speech in more than 100 languages. This broad language support enables developers to create voice products for a global audience.

What impact does Flash TTS have on the development timeline for voice-based products like audiobooks and games?

The text-prompt voice design and 30-second cloning capabilities compress a process that traditionally required studio time and licensing deals into something developers can prototype in an afternoon. This dramatically reduces development time and costs, fundamentally changing the competitive landscape for audiobook and game-character startups.

Can users add stage directions to individual lines of dialogue in Google's new Flash TTS models?

Yes, both Flash TTS and Flash-Lite TTS allow users to add stage directions to individual lines of dialogue, providing greater control over how the generated speech is delivered. This feature enables more nuanced and expressive voice generation for creative projects.

LIVE20:47Stanford AI Agents Learned to Collude at Blackjack