Skip to main content
Gradium's voice design: AI-generated synthetic voice from text prompts, created in seconds.

Editorial illustration for Gradium's Voice Design Creates Custom Synthetic Voices From Prompts in Seconds

Gradium Creates Custom Synthetic Voices in Seconds

4 min read

Voice agent teams have a specific problem, and it shows up the same way every time: the catalog has 400 voices, and the brief needs the one voice that isn't there. A Quebecoise receptionist for a car dealership in Montreal. A narrator in his sixties who sounds like he's held a lecture hall's attention for decades. Cloning has been the workaround, but it solves the problem one speaker at a time, and each one drags along sourcing headaches, consent paperwork and a licence to manage.

Gradium, the Paris-based voice AI company that spun out of the Kyutai research lab, is betting there's a faster route. Its new tool, Voice Design, takes a written description and hands back a finished synthetic voice in seconds, no recorded speaker, no reference audio, no rights to clear. It's built to run inside the same infrastructure teams already use: live now in the Gradium API and in Studio, free across every plan including the free tier, with any voice a team keeps running on the identical streaming Text-to-Speech endpoint as the rest of the catalog, matching latency and output formats.

The mechanics of how that description actually turns into a voice come down to what the model is trained to listen for.

Gradium, the Paris-based voice AI company spun out of the Kyutai research lab, has shipped a different answer. Voice Design reads a written description and returns complete new voices in a few seconds.

Why this matters

For anyone building voice agents, the catalog problem is real: 400 stock voices can't cover every accent, age, or use case a product brief demands, and cloning a real speaker drags in consent forms and licensing headaches nobody on a sprint deadline wants to deal with. Gradium's pitch, text prompt to synthetic voice in 3 to 5 seconds, is worth watching precisely because it sidesteps that sourcing problem entirely. No real person's voice to license, no consent chain to manage, just a description and a handful of candidates to audition.

The advice to end prompts with intended use rather than pure description is a small but telling detail: it suggests the model is tuned for delivery and register, not just timbre, which matters more for production voice agents than a pretty-sounding demo. We'd want to see how consistent that "single character" promise holds across longer scripts and different emotional ranges before trusting it in a live product. Still, coming out of Kyutai's research lineage, this looks like a genuine attempt to fix a workflow bottleneck, not just add another voice to a crowded catalog.

Common Questions Answered

How does Gradium's Voice Design solve the catalog problem for voice agent teams?

Voice Design reads a written description and generates complete new synthetic voices in just a few seconds, eliminating the need to search through limited voice catalogs or clone existing speakers. This approach bypasses the sourcing headaches, consent paperwork, and licensing requirements that come with traditional voice cloning, allowing teams to create custom voices that match their specific product briefs on demand.

What are the main limitations of voice cloning that Gradium's solution addresses?

Voice cloning solves the catalog problem one speaker at a time and requires managing sourcing headaches, obtaining consent paperwork, and handling licensing agreements for each cloned voice. Gradium's prompt-based approach eliminates these friction points entirely by generating synthetic voices without needing to license or manage real people's voices.

How quickly can Gradium generate synthetic voices from text prompts?

Gradium's Voice Design can generate complete new synthetic voices in 3 to 5 seconds from a written description. This rapid generation time makes it practical for teams working on sprint deadlines who need custom voices without delays.

What is Gradium's background and where is the company based?

Gradium is a Paris-based voice AI company that was spun out of the Kyutai research lab. The company developed Voice Design as a solution to the specific challenges faced by voice agent teams in creating custom synthetic voices.

LIVE10:32AI Agents Run Quantum Experiments Through Software