How to Create Crowd Noise in an Audio Scene
Crowd noise is one of the places in audio production where doing less gets you a better result. A tavern, a courtroom, a packed market: the listener has to believe the room is full, and the quickest way to lose that belief is to build the crowd out of individual people. In AudioProducer the crowd is made of a soundscape running underneath the scene, a small number of cast voices sitting on top of it, and a few one-shot effects that give the room a pulse. This guide covers how to stage that, and what to do when your manuscript asks for people talking over each other.
A crowd is a texture the ear reads as size
The sound of a full room has shape and almost no detail. You hear pressure, movement, a general pitch, occasional peaks. What you do not hear is sentences. The moment a listener can pick out a full line from the background, that background stops working as background and becomes a second conversation competing with the one you actually wrote.
This is why casting a crowd goes wrong. Six voices reading six lines does not sound like sixty people. It sounds like six people in a quiet room, because every line lands clean, centered, and fully audible. The texture has to come from the ambient layer, and the individuals from the two or three characters who carry the scene.
We render lines one after another, so overlap has to be staged
Here is the constraint to plan around, stated plainly: AudioProducer generates each tagged line in sequence. There is no setting that makes two characters speak at the same time, and no crossfade between adjacent lines. If your manuscript has four people shouting at once, the export will play those four lines in order, cleanly separated. That reads as an orderly queue, not a brawl.
Staging is the workaround, and it is a writing move before it is a production move. Break the chaos into short fragments: a three-word line, a beat, a reply that starts mid-thought, someone finishing a sentence nobody was listening to. Short lines in quick succession read as interruption even though each one is technically complete. Then shorten the gaps between them. Pauses are set project-wide under Edit Project, overridden per paragraph, or dropped inline, so you can tighten the crowded stretch without touching the pacing of the rest of the chapter. Our guide on pauses and dramatic timing goes further into how much silence a scene can carry.
Underneath all of it, run the crowd bed. The bed supplies the sense of many. The staged fragments supply the sense of disorder. Neither one does the job alone.
Build the room first, then put voices on it
Work from the bottom up. In the Sounds panel, place a music bed or ambient soundscape that runs under the whole scene, then add one-shot effects at the specific moments that need a hit. Auto-Assign Sounds will do a first pass for you, reading the scene and placing music, soundscapes, and one-shot effects from the library. Treat it as a starting point and then cut, because its instinct on a crowded scene is usually to give you more than the scene needs. The general method is covered in adding ambient sound to an audio story, and if you are marking up a whole manuscript rather than one scene, a sound effects cue sheet keeps the placements consistent from chapter to chapter.
Only then cast. Which characters get their own voice in a crowded scene is a separate question with its own answer, and we have written it up in narrating dialogue-heavy scenes. The short version for our purposes: the bed handles the room, so you need fewer distinct voices than the page suggests, not more.
What a tavern, a courtroom, and a market are made of
The bed is not interchangeable between rooms. A tavern is a low, warm, continuous murmur with a fire under it and intermittent one-shots on top: a mug set down, a bench dragged, a door. A courtroom is closer to silence, and its crowd is felt through restraint. A cough, a shifting bench, a gavel, and long stretches of nothing, which is what makes the room feel like it is holding its breath. A market is brighter and wider, with more movement and a faster rate of one-shots because things are constantly happening at the edges.
What actually sells the size of a room is the rhythm of those one-shots, not the volume of the bed. A steady murmur with one event every twenty seconds reads as a half-empty hall. The same murmur with something landing every three or four seconds reads as packed. You can change the apparent size of a crowd without changing the crowd at all. Sound work at this level is the same craft an audio drama runs on, which we cover in sound design for an audio drama.
How much crowd is too much under dialogue
Intelligibility is the ceiling. If the listener has to work to catch a plot-carrying line, the scene has failed no matter how good the room sounds. Be honest about the control you have here: what AudioProducer gives you is placement, meaning where a soundscape starts and where it stops. There is no per-line volume fader and no automatic ducking under dialogue.
So the lever is scheduling rather than mixing. Let the bed run through the narration and the atmosphere-building lines, then end it before the exchange that carries the information, or start it after. A room that goes quiet the moment two people start talking is a convention listeners accept without noticing, and it is far kinder to the ear than a full crowd under a whispered confession. If you also have music under the scene, be stricter still, since two competing layers under dialogue is the most common way a good scene turns muddy. Adding sound effects and music to an audiobook has more on stacking those layers.
When the narrator should just say the room is full
Some crowds are not worth building. If the scene passes through a busy room on its way somewhere else, one narrated sentence about the noise does more than thirty seconds of production. Save the staged crowd for the scenes where the room is doing something to the characters: pressure, exposure, being overheard. A crowd that shows up in every location stops being an event and turns into wallpaper the listener tunes out by chapter three.
When the markup is done, generate the chapter and listen to it end to end rather than scrubbing to the crowd. Density problems only show up in context, and the fix is almost always removal. AudioProducer exports finished MP3 files per chapter, so you can check a crowded scene in place before committing the rest of the book.
Frequently asked questions
- Can AudioProducer make two characters talk at the same time?
- No. Every tagged line renders in sequence, one after the other, and there is no crossfade between lines. To suggest people talking over each other, stage the exchange instead: write short fragments, shorten the pauses between them, and run an ambient crowd bed underneath so the room supplies the sense of many voices.
- How do I make a crowd sound big without casting a lot of voices?
- Put the size in the ambient layer rather than in the cast. Place a soundscape that runs under the whole scene, then add one-shot effects on top. The rate of those one-shots is what the ear reads as size: one event every twenty seconds sounds like a half-empty hall, while something landing every few seconds sounds packed.
- How loud should crowd noise be under dialogue?
- Quiet enough that a plot-carrying line is never work to catch. The control AudioProducer gives you is placement, meaning where a soundscape starts and stops, not a per-line volume fader or automatic ducking. So schedule the bed around the important exchange: run it through narration and atmosphere, and end it before the lines that carry the information.