DEV Community

Incomplete Developer
Incomplete Developer

Posted on AI-assisted

I Tested SUNO's New Speech Feature (BETA): Why Speech Is a Harder Problem

#ai

Heads up: SUNO's Speech feature is still in beta and not officially released. Everything below applies to the beta only and may change a lot.

Music is a bounded problem. Speech isn't.

SUNO made its name generating songs. Now it's added Speech: you give it a script, and a voice reads it out. I spent some time with the beta, focusing on speech generation and mostly skipping songs and background music.

Testing it made me think about how different the two problems are. Music can be categorized by genre or style, and the main challenge is novelty. Speech is open-ended, organic, and unpredictable, and a model has to handle several things at once:

Voice variety: "male" or "female" isn't enough

Languages: not everyone speaks a European language
Age and accent: including the many accents within a single language
Context: the subject matter should shape the delivery

For anyone building on top of generative audio, it's a good reminder that each of these is another dimension your prompt has to control.

What I tested

Simple mode with the default football coach prompt
Advanced mode with custom script, tone, gender, and music settings
A child's voice: a young girl rallying her soccer team
Accents: Australian English and British/European English
South African languages: Sotho, Zulu, and Afrikaans
Spanish as a point of comparison

I won't spoil the results here. I'll just say the outcomes weren't uniform, and one prompting detail made me wonder how the model parses instructions.

Watch the full tests

You can hear every test and my full thoughts in the video:

👉 https://youtu.be/66LDOI4MOhQ

My conclusion: SUNO may have just dipped its toes into some waters, and it's unclear how deep they go.

Have you tried the Speech beta? Tell me which voices, accents, or languages to test next.

Top comments (0)