Suno Opens Speech Beta, Generating Voice and Music as One Track
Suno announced on October 1, 2026 that Speech, a new audio model built directly into its platform, is now open in beta to everyone on web and mobile. The company describes it as the first audio model that generates voice and music together as one cohesive track, rather than producing narration and a soundtrack as separate layers.
Details
- Core idea: users type an idea, a poem, or something they have written, then describe the voice and musical style they have in mind, and Speech generates spoken audio with original background music composed to match it as a single track
- Chief Product Officer Jack Brody authored the announcement, positioning Speech as an extension of Suno’s existing song-generation models into narrated, voice-led content
- Example use cases cited by Suno include dramatic readings, meditations, pep talks, and bedtime stories, as well as personal creative projects like birthday messages and inside jokes
- Rollout: the beta opened to all users after roughly a month of testing with a smaller group, and is now available across both the web app and mobile apps
- Known limitations: Suno acknowledged beta-stage quality issues in its own announcement, noting that “British accents can wander off to Australia and back” and that “dramatic pauses may be very dramatic,” with further refinement expected as user feedback comes in
- No separate pricing was announced for Speech at launch; it ships as part of the existing Suno product
What happened next
The Speech beta arrives about three weeks after Suno shipped its v6 music models in partnership with Warner Music Group, BMG, and Believe, and lands the same week CEO Mikey Shulman said at Bloomberg’s Screentime conference that Suno has grown to “far beyond” two million subscribers and roughly $300 million in annual recurring revenue. By combining narration and music generation into a single model, Suno is pushing further into spoken-word and voice content, a space more commonly associated with text-to-speech specialists, rather than staying confined to song generation.