One call turned a 298-character script into 32 seconds of finished radio drama: three speaking characters, rain on an awning, a kettle in the kitchen, sequenced in the order I wrote them. That part is real, and on headphones it holds up.
What no demo shows you is a wall at 120 seconds. Go past it and the API cuts the ending off your audio in silence, does not raise an error, and bills the full duration anyway.
The model listed on the Alibaba Cloud Bailian platform on 09-21, and I spent half a day running eight generations through it, from npm install -g bailian-cli to reading the finished wav files back against the scripts. Every number below comes out of an actual call. The useful parts are the ones the samples cannot show: what a run costs, a CLI that does not route this model generation yet, and two limits with completely different failure modes.
What the runs produced
- A 298-character script came back as 32.08 seconds of finished audio: three characters, two beds of ambience, effects in the order written. A 439-character script became a 69.96-second two-host podcast.
- Voice consistency has a mechanism: up to three reference clips, entered as
@voice1through@voice3, each up to 30 seconds and 10MB. I had no voice material, so I made samples with an olderblmodel and the whole chain stayed on my machine. - The CLI does not route this generation yet.
bl speech synthesizereturns a 400 because it builds the older request shape; synthesis goes through the official HTTP endpoint, andblhandles the surrounding work of catalog, quota and read-back. - Cost tracks output tokens: generated seconds x 200, about CNY 0.0024 a second. A 32-second drama is about CNY 0.08; a four-minute podcast about CNY 0.60.
- Two limits, two behaviors. A prompt over 3,000 characters fails in 0.2 seconds with a clear message. Audio past 120 seconds is cut silently, and the response still reports a normal finish.
- The same prompt run twice produces two different files, seed included.
- Read-back is the verification that matters:
bl speech recognizetranscribes the finished wav, so a missing last line shows up as text instead of a vague feeling that something was off.
Getting a request through
Setup is two lines. Let your agent run the official skill package, or in a terminal:
npm install -g bailian-cli
Synthesis needs a key; sign one here and put it in an environment variable.
My first attempt went through the CLI's own synthesis command, and it failed.
bl speech synthesize --text "雨突然下大了。要伞吗?" \
--model qwen-audio-3.1-tts-next --voice longanhuan_v3.6 --format wav --out smoke.wav
The response was an HTTP 400: text_prompt must not be empty.
--text clearly had content in it. The help text explains the mismatch: this command builds the classic synthesis shape, input.text plus voice, and the new model wants input.text_prompt and rejects voice outright. Point the same command at the older qwen-audio-3.0-tts-flash and it works first try. The model works; the CLI path has not caught up with this interface yet.
Synthesis, then, over the official HTTP endpoint:
curl -sS -X POST \
"https://$WORKSPACE_ID.cn-beijing.maas.aliyuncs.com/api/v1/services/audio/tts/SpeechSynthesizer" \
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-audio-3.1-tts-next",
"input": {
"text_prompt": "深夜面馆广播剧。环境声:雨点密集打在雨棚上。老板:这么晚才来。还是老三样?",
"format": "wav",
"sample_rate": 48000,
"channels": 2
}
}'
Three fields in the response carry the day-to-day weight: output.audio.url (the finished file, valid for 24 hours), output.audio.duration, and usage for that call. My smoke test, 6.6 seconds of audio, took 41.5 seconds to generate.
Scene one: 298 characters in, 32 seconds out
I picked a late-night noodle shop because it is naturally loud: rain on the awning, a kettle in the kitchen, a door pushed open, a scooter pulling away. A good test of whether effects follow the story.
The script used the official six elements: three characters (the owner, 55; a customer working late, 26; a delivery rider, 23), each with a voice, a mood and a pace; two beds of ambience; two emotional turns. A fragment:
Owner (male, 55, rough low voice, slow): You are late again. The usual three?
Customer (female, 26, tired and hoarse, quiet): Mm. And add an egg today.
Rider (pushes the door open, trailing rain in, male, 23, out of breath): Boss, is order sixteen ready? This rain is too heavy.
The full script went in, 32.08 seconds came out, about 0.11 seconds per character. Generation took 121.6 seconds, the longest wait of the batch; three characters and four kinds of sound do cost time.
Listening is subjective, so I wanted a check with an answer in it. bl speech recognize reads the file back and gives one: all six lines matched, none missing. The timestamps added two facts worth having. Voices enter at 2.7 seconds, so the rain opens the scene, and the last line ends at 27.8 seconds, leaving four seconds of rain to close it.
Scene two: a borrowed voice
Describing a voice inside a prompt works for one scene. A series needs the same character to sound the same in every episode, and that is what reference audio is for: up to three clips in WAV, MP3 or OGG Opus, referenced as @voice1 and @voice2. I generated two samples with qwen-audio-3.0-tts-flash, base64-encoded both into the request, and assigned lines.
The result ran 37.36 seconds, and the read-back matched all four lines, conversational endings included.
This run also taught me how literal the ordering is. I had written the ambience first in the prompt, "raindrops on the window, low engine hum", and got 18 seconds of it before the voices entered at 18.4. Sequencing follows the prompt line by line. To bring voices in sooner, write "three seconds of rain, then @voice1 speaks".
Scene three: the podcast, and the wall
The most common audio format of all deserved a run: 439 characters of two-host chat, one fast and one slow, with an interruption instruction in the middle. The finished file came in at 69.96 seconds, about 0.16 seconds per character, and the read-back covered the whole thing with nothing missing.
Generation took 101.6 seconds, roughly 1.45 times the finished length. Efficiency climbs with length: the 6.6-second smoke test waited 41.5 seconds, more than six times its own duration, while a 70-second piece stayed under 1.5 times.
Then I tested the limits, and this is the part worth remembering. The documentation states two: prompts under 3,000 characters, and finished audio capped at 240 seconds for podcasts and 120 seconds for other scenes.
The first wall errors. A 3,058-character script returned text_prompt exceeds the maximum length of 3000 characters. in 0.2 seconds. Rejected at the door; fix the script and move on.
The second wall does not. I wrote a 469-character monologue, which at 0.16 seconds per character should land near 75 seconds, well inside the cap. The finished file was exactly 120.0 seconds. The last sentence was gone. The response record's finish_reason still said stop, and the bill, about CNY 0.30, was for the full 120 seconds.
No error, no warning, no flag on the response. For audiobook work, where scripts run long and the cut lands silently, you would ship a chopped ending without ever hearing it, unless you read the file back. Estimate before generating (about 0.16 seconds per character; 0.12 to 0.15 for dialogue-heavy copy) and split anything that approaches the cap.
The bill, in output tokens
Eight runs, one rule, no exceptions: every generated second records 200 output tokens. 6.6 seconds recorded 1,320; 32.08 recorded 6,416; 69.96 recorded 13,992; 120 recorded 24,000.
At the catalog prices (CNY 6 per million input tokens, CNY 12 per million output), that works out to about CNY 0.0024 per finished second. The catalog quotes in yuan and I am leaving the figures there rather than pretending a conversion; the shape is the portable part.
| Finished | Generation | Cost (approx.) |
|---|---|---|
| 6.6s (smoke test) | 41.5s | CNY 0.02 |
| 32s (radio drama) | 121.6s | CNY 0.08 |
| 37s (voice transfer) | 90.2s | CNY 0.12 |
| 70s (podcast) | 101.6s | CNY 0.17 |
| 120s (monologue, capped) | 174.6s | CNY 0.30 |
Everything in this session, eight generations and 289 seconds of finished audio, came to about CNY 0.77. The model carries its own free allowance (90 days, per model, Beijing region only); a million output tokens cover about 83 minutes of finished audio.
Two billing details worth the effort. A prerequisite first: bl usage free and bl log audit list are tagged [Console] in the CLI help, so they need bl auth login --console in a browser; synthesis and read-back run on an API key, and the two are separate credentials. The quota panel lags, so per-call accounting goes through bl log audit list --output json, which carries tokens, billed duration and status per call; that is how I summed the batch and landed exactly on the later panel total. And the speaking-rate parameter does not save money: rate=1.6 shrank 9 seconds of audio to 5.3 physical seconds while the billed output moved only from 1,808 to 1,688 output tokens (down 6.6%). Less content is the only real saving.
Two surprises, and the habit I kept
The same prompt with the same parameters, seed included, produced two wav files with different md5 hashes. Do not plan on "generate it again just like that": archive the version you like, because a rerun is a new performance, not a copy.
The habit: read back every file. bl speech recognize transcribes the wav and returns timestamps, so line integrity, swallowed words and voice timing all become reviewable text. It is faster than repeated listening, and it is the only thing that caught the silent cut.
Try it
Two ways in: the official skill package for your agent, or npm install -g bailian-cli in a terminal. Synthesis needs a key, and you can sign one here.
Three finished pieces are still up if you want to hear the quality before spending anything: the 32-second noodle-shop drama, the 37-second voice transfer, and the 70-second podcast.
What would you feed it first: a script that has been waiting for a voice, or something else entirely?



Top comments (0)