It is 11pm and the level is playable. Geometry is in, the lighting reads well, the character controller feels right after the third tuning pass, and every audio slot in the scene holds the same eight-second placeholder tone that a colleague recorded off a monitor. That is not the problem. The problem is that you sit down to replace the placeholder for the forest that runs along the north edge of the map and realise you do not have a vocabulary for the thing you want.
I work on AIDubbing, and this post is marketing for our AI sound effect generator. I am writing about it because the argument below is the one I find most useful when talking to other people about sound in prototypes, and I would rather make the interest explicit than let you find it in the last paragraph.
You know the function. You want the forest to feel larger than it is. You know the failure mode. A stereo field of birds layered over a bed of leaves reads as wallpaper, as texture, as the audio equivalent of the same tree repeated on a tile. The job of the ambience there is not to be pretty, it is to imply depth — near, mid, far — so that the player suspects there is more world off the edge of the frame than there is. And when the build meeting is on Thursday, nobody is going to sit in a room and listen to a hundred forest takes and pick the one that implies depth. There is no version of that process where you win.
So the audio slot stays a placeholder for another week.
Recognition versus specification
A sound library is a retrieval problem wearing a production tool's clothes. The interface is a folder tree, the metadata is someone else's taxonomy, and the deal you make with it is that you already know what you want. You type forest. You get twelve results, all competent, all roughly the same recording of the same wood in the same weather. Now you are doing the work that the tool was supposed to do: you are sitting there with headphones on, cycling, forming an opinion, and trying to translate that opinion back into a query. Most of the time the opinion is vague, because opinions usually are, and the query is the one thing you can control, so you start adding words. forest night. forest ambience long. You are not specifying a sound any more. You are guessing at filenames.
The problem is structural, and it shows up in the shape of the search box. Keywords are labels for things that already exist. Descriptions are specifications for things that do not. When your actual need is a behaviour — this forest needs to feel bigger than it is — there is no label for it, because nobody recorded a file called forest that feels bigger than it is. The label would have to be someone else's, and someone else's labelling instinct is not tuned to your scene.
Once you frame the input as a specification rather than a query, the actual work changes shape. You stop asking what to call it and start asking what the sound has to do in the frame. You can answer that from the level design. The north forest edge exists to make the player's imagined map bigger than the mesh. The forest ambience's job is to carry the top of that range: air moving through canopy, distance, some suggestion of volume behind the tree line. Whether that reads is a design decision, not a filing decision.
That reframing also tells you when a sound is not the right tool, which saves more time than any generation step. The forest is ambience. The weapon is not. A weapon sound in a first-person game has to do its entire job inside a single frame, because the player is looking at the muzzle when it happens and their eyes will not go to the audio. A two-second tail-heavy explosion is a sound for a cutscene. What a first-person weapon wants is a short transient that survives being played at low volume while the player is sprinting, plus enough body that it does not read as a click. You cannot express that as a keyword at all. You can express it as a description, badly, but you can express it.
There is a third case that is purely about placement: a bed under dialogue. When someone talks over ambience, the bed is not competing with the voice, it is staying out of its way while holding the room together underneath it. That constraint — present, continuous, and consistently behind the attention — is a description. forest ambience is not. A keyword for it gives you a recording that was mixed to be the thing you notice, because recordings get made for that purpose.
And then there are the cues: interface clicks, state changes, a transition sting. These are the sounds where being wrong is louder than being absent, and where the difference between "a UI click" and "a UI click that pairs with the panel animation that arrives 120 milliseconds later" is the entire job.
What the prompt actually is
The engineer's version of this is a design constraint, not a feature.
The prompt is a specification document. It has to carry enough information to constrain a result, and it does not have room for adjectives that carry no information. When I write a description for a scene, I am trying to answer four questions in the text, and if I cannot answer one of them I have not thought about the problem yet.
The first is material. Metal, wood, stone, foliage, ceramic, water, air. Material is what determines the decay shape and the frequency content, and it is the single most load-bearing word in the prompt. metal impact and wooden impact are different sounds and both of them are obvious, while impact alone is a coin flip.
The second is distance and space. Near, mid, far, inside, outside, through, across. A sound described as distant is not the same audio made quieter; it has a different spectral signature, because air absorbs the top end first, and it sits differently in the mix. If you want the tree line to feel far, the description should say where the sound is, not how loud it is.
The third is intensity and character, expressed as something physical rather than evaluative. heavy, sharp, low, single, dense are usable. epic, intense, powerful, scary are not, because they describe your reaction rather than the sound's structure, and you will get whatever the model guesses you meant.
The fourth is what the sound is attached to, and this is the one people leave out. A footstep is a sound plus a surface plus a weight plus a gait. A door is a hinge or a latch or a slab, and the difference is three-quarters of the character.
Vague prompts fail for a boring reason: they under-constrain, and under-constrained output is average, because average is the safest thing to return. Every word you omit is a degree of freedom you have handed to something that did not know about your level. Forest sound leaves open the tree count, the weather, the size of the space, and whether anything is moving. Dry pine forest at night, wind moving through high canopy, no animals, distant, continuous, no distinct events is not a better prompt because it is longer. It is better because there is nothing left in it to guess about except the timbre.
Our generator lives at this address:
The interface is built around that idea, which is why the presets on it are worth reading closely. They are not genre labels. Thunder rumbling in the distance, ocean waves crashing on shore, birds chirping in a forest, rain drops on a window, crackling fireplace, wind through trees, city traffic ambience. Every one of those is a scene description with a distance and a material and an event structure baked in. Rain drops on a window in particular is doing three jobs at once: a material, an implied interior space, and a small repeating event rather than a single hit. That is the register. The showcase on the same page runs the same way — fast typing on mechanical computer keyboard, crackling fire in a cozy fireplace, ambulance siren approaching and passing quickly, each one pinning a tempo, a surface and a trajectory. There is a text field, a duration control at five seconds, and a generate button, plus a history panel with a view-all link so the takes you rejected during a session stay there to compare against. Outputs come back as MP3s, which the rest of your engine and editor already read.
Two things I would not claim. One, that any description lands the intent on the first pass — you still take three or four and you still pick the closest. That is the same picking you were doing before, except now the options come back inside the range you asked for instead of inside someone else's. Two, that the result is interchangeable with a field recording. It is not, and treating it as a permanent replacement for a proper recordist is how you end up with a pipeline you have quietly under-built. The questions people actually ask are the useful ones here: whether it needs skills to drive it, whether the duration is adjustable, what comes back and in what format, what you are allowed to do with it. Answer the licensing one from your own legal process, not from a marketing page. Mine included, which is the point of saying it at the top.
What changes in the pipeline
The visible change is the number of humans involved in an audio decision. Under a retrieval workflow, sound is a shared queue: someone on the team owns the library, someone else owns the build, and the person who could most accurately describe what the scene needs is not in that loop at all. They email a request, the request becomes a keyword, the keyword becomes a guess, and the designer who had a clear opinion in their head ends up accepting a take they did not choose because resolving it properly was not on the critical path.
Under a specification workflow, the designer writes the line. The line is reviewable in a text document, in the same diff, next to the level notes. A designer can be wrong about a sound and be corrected about a sentence, which is a much cheaper correction than a designer silently tolerating the wrong audio because there was no channel for the disagreement. The audio person stops being a gate and starts being a reviewer of specifications, which is the part of the job that actually scales.
The second change is iteration cost, not generation time: it is the number of conversations required to get to a decision. When the input is a description, the first round of feedback is about the description: too busy, the birds are reading as foreground, drop the events and keep the air. That feedback is legible, specific, and survives contact with a different person. When the input is a file, the feedback is I don't know, next one, and it stays that way for forty files.
The third is that briefs stop getting lost. A description written during a playtest is a sentence in the project's notes, in plain text, reviewable in six months when the level is being reworked and the original recordist has moved on. It does not depend on anyone's disk, anybody's account, or a particular machine's folder structure. It is just a requirement that was written down, which is the oldest and least glamorous property a pipeline artefact can have.
The failure mode to watch for is the opposite one: everything becomes a description, and nothing is ever actually recorded. A pipeline that describes every sound and records nothing ends up with a codebase of things that are approximately right. The pragmatic split is to specify what needs to exist and does not, and to keep paying for the twelve sounds per project that carry the identity. The placeholder tone in the scene at 11pm gets replaced either way. The question is whether you are choosing the sound, or searching for the sound someone else already chose.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.