“Make a cinematic audio scene with dialogue and music” is a prompt, but it is not much of a production brief. It leaves the model to decide which sound enters first, how the voices relate to the room, and what should remain audible when music starts.
For a short scene, the useful unit of work is a timeline plus a hierarchy of sounds. Here is a reusable way to write one, using a 30-second rainy teahouse scene.
1. Specify five production decisions
Before writing prose, decide:
| Decision | Example |
|---|---|
| Duration | 30 seconds |
| Setting | A small teahouse at dusk; rain stays outside |
| Voices | One hurried customer and one calm owner |
| Emotional turn | The customer arrives anxious and leaves reassured |
| Mix priority | Dialogue first; ambience and music must not mask speech |
Limit foreground effects too. In this example, a sliding door, two wet footsteps, and a porcelain cup are enough. Each one has a job in the story. Adding more effects would make it harder to tell which instruction improved the next generation.
Keep the scene, character direction, ambience, and effects in one brief so they can be evaluated together.
2. Describe when the scene changes
Treat timestamps as editorial targets, not promises that a generator will hit exact frames.
- 0–5 seconds: Establish rain outside and a quiet interior. A door slides open; wet footsteps stop near the counter.
- 5–12 seconds: The customer asks, “Is it too late?” Leave a beat before the reply.
- 12–21 seconds: A cup touches the counter. The owner answers, “Only if you wanted the last train.” Let the cup land between the two lines.
- 21–30 seconds: The customer exhales. Both voices soften, and the music rises slightly before a clean ending.
This structure gives you something testable. If the cup sound arrives too early, you know which part of the brief to revise. If the ending feels abrupt, shorten a line before asking for a longer fade.
3. Turn the timeline into one prompt
Generate a 30-second audio scene in a small teahouse at dusk. Use two distinct adult voices: a hurried customer and a calm owner. Keep dialogue clear and close. Soft rain remains outside throughout; the room itself is quiet. At the opening, a wooden sliding door moves and two wet footsteps approach. The customer asks, “Is it too late?” After a short pause, a porcelain cup touches the counter. The owner replies, “Only if you wanted the last train.” The customer exhales and the tension eases. Use a subtle, warm music bed that never masks speech. End naturally at 30 seconds. Avoid extra voices, dramatic impacts, and loud percussion.
Replace the location, emotional turn, and dialogue with your own scene. If you add reference audio, use material you have permission to upload.
4. Evaluate one layer at a time
On the first listen, ask three questions:
- Can you understand every word?
- Does the cup sound support the dramatic turn?
- Does the music help the mood without taking over?
If speech is buried, simplify the music direction. If the scene feels crowded, remove an effect. Change one instruction per take so you can hear what caused an improvement.
You can rehearse this workflow at SeedAudio 2.0. The current studio generates with Seed Audio 1.0; Seed Audio 2.0 is listed as coming soon. The site describes longer scenes, video-aware input, more reference voices, and separate dialogue, music, ambience, and effects stems as planned 2.0 capabilities. Those features are not yet live in the current studio, so judge the audio you can generate today and keep refining the brief.

Top comments (0)