Editorial illustration for Suno AI music generator adds integrated speech to its tracks
Suno AI Adds Speech Feature to Music Generator
Suno built its reputation on turning text prompts into full songs, drums, vocals, mixing, all of it. Now the company wants to handle something plainer: just a voice, talking. On Thursday, Suno rolled out a public beta called Speech across its web and mobile apps, letting users feed in a script or a description and get back a synthetic voiceover, with the option to layer AI-generated music underneath automatically.
The feature works both ways. Users can type out exact dialogue they want spoken, or describe a tone and let the model improvise. A toggle switches the background score on or off, so the tool doubles as a bare text-to-speech generator when music would just get in the way. Suno is pitching the combination as its point of difference in a field that already includes DeepMind's decade of speech synthesis research, Adobe's text-to-speech tools, and ElevenLabs, which has built a sizable business on voice generation alone since 2023.
For Suno, the move also lands at a moment when its core music product remains tangled in copyright lawsuits from major labels, making a second product line a useful hedge.
Suno is branching out from the world of AI music, launching a new feature that generates spoken voices based on scripts or prompted descriptions. Speech is now available in public beta across Suno’s web and mobile platforms, and allows you to simultaneously generate voiceovers and background music to accompany them.
Why this matters
Suno's bet is that bundling voice and music into one generation pass beats stitching together separate tools, and for podcast producers, indie game studios, or solo creators making narrated content, that's a real workflow shortcut. No more exporting a voiceover from one tool and dropping it into a DAW to layer music underneath; Suno says it does both at once. Whether that "cohesive track" claim holds up against a human editor manually syncing a narrator to a score is the thing to actually test before trusting it for client work.
The competitive picture matters more than the feature itself. DeepMind has spent a decade on speech synthesis, Adobe already ships text-to-speech, and voice cloning tools are everywhere. Suno isn't first to synthetic speech, it's first to fuse it with its existing music engine.
That's a narrower claim than the launch language suggests, and worth remembering when evaluating public beta output against established single-purpose tools. For developers building on Suno's API, the question is whether combined generation actually saves compute and editing time, or just repackages two solved problems into one marketing pitch.
Common Questions Answered
What is Suno's new Speech feature and how does it work?
Suno's Speech is a public beta feature that generates synthetic voiceovers based on scripts or descriptions provided by users. The feature allows users to input exact dialogue they want spoken and automatically layer AI-generated music underneath, creating a complete audio track with both voice and music in a single generation pass.
How does Suno's integrated speech and music generation differ from traditional audio production workflows?
Instead of using separate tools to create voiceovers and music independently, then exporting and manually syncing them in a DAW, Suno's approach bundles both voice and music generation into one process. This eliminates the need for multiple export steps and manual layering, providing a significant workflow shortcut for content creators.
Which platforms can users access Suno's Speech feature on?
Suno's Speech feature is available in public beta across both Suno's web and mobile applications. This multi-platform availability allows users to generate voiceovers and background music from various devices.
What types of creators would benefit most from Suno's integrated speech and music generation?
Podcast producers, indie game studios, and solo creators making narrated content would benefit significantly from this feature. The integrated approach saves time and streamlines the production process for anyone who needs to combine voiceovers with background music in their projects.
Further Reading
- AI music maker Suno now generates spoken words - The Verge
- AI Startup Suno's New Tool Puts Music to Spoken Word - Bloomberg
- Introducing Speech (beta) - Suno
- AI Startup Suno Launches New Tool To Add Music To Spoken Word - NDTV Profit
- Suno | Introducing Speech (beta) - Suno