How Many Voices Does an Audio Drama Need?

August 9, 2026

Answer first. Most audio dramas need fewer distinct voices than the character list suggests. Count the speakers who share a scene rather than the names in the script, and a cast of twenty names usually resolves to somewhere around six to ten voices, with the rest doubled.

That falls out of how listening works. A scene almost always runs on two or three people talking, and the listener only has to tell those apart in the moment. For the production side of this, how to make an audio drama with AI covers casting and generation end to end. This post is only about the count.

Count speakers per scene, not names in the script

Go through your script scene by scene and write down how many characters actually speak in each one.

Most scenes come back with two or three. A busy one might hit five. The number you want is the largest of those, because that is the moment where a listener has to hold the most voices apart at once.

Names that never share a scene are not competing for anything. Two characters who appear in different chapters can carry the same voice and nobody notices. That is where most of the savings live.

Do this before you start auditioning. A count you arrive at by scrolling the cast list is a guess, and it is usually high.

Doubling: one voice, several small parts

Doubling is one voice covering more than one part. Stage and radio have done it forever, and it is the reason a cast list is longer than a voice list.

In AudioProducer.ai you assign a voice to each character, so doubling is just assigning the same voice twice. Auto-Assign Characters tags the lines by speaker first, and you keep those tags exactly as they are. Only the voice behind two of them is shared.

The rule that keeps it invisible: separate the doubled parts across scenes, never inside one. A voice that plays the innkeeper in chapter two and the harbor master in chapter nine reads as two people. The same voice answering itself across a table reads as a mistake.

Give the doubled parts different weight where you can. A short functional part and a recurring one sit further apart in a listener's memory than two parts of the same size and the same energy.

Does your drama need a narrator?

The narrator is a casting decision before it is a stylistic one, because it is usually the voice with the most lines in the project. Some dramas keep one to carry description between scenes. Others drop it and let dialogue and sound do that work, the way radio plays did. Audio drama vs audiobook walks through which format that choice lands you in.

Dropping the narrator does not lower the count. It moves the load. Everything the narrator would have said has to arrive as dialogue or as a sound cue, so scenes grow speaking parts and the cues carry more meaning. Adding sound design to an audio drama covers what a cue can reasonably say on its own.

Keeping a narrator adds one voice and takes pressure off the rest of the cast. It is the easier choice when your source is prose rather than a script, since the description is already written.

When a separate voice earns its keep

A character earns its own voice when a listener would otherwise lose track of who is speaking. In practice that is the two or three people who carry an argument in a scene, plus anyone whose lines run long enough for the listener to settle into them.

Below that line, writing does the work more cheaply than casting. A distinctive turn of phrase, or one character saying another's name in reply, identifies a speaker as reliably as a new voice and costs nothing in casting effort. The wider trade-off between a lean cast and a full one is laid out in single-voice vs full-cast.

A separate voice will not rescue a character defined only by how loud they are. There is no volume control on a line in AudioProducer, so a whisper or a shout has to be built from wording and from the voice you cast. How to narrate a character who whispers or shouts is the read for that case.

The ceiling is casting effort and listener memory

Every extra voice costs you twice. You audition it once, and then you carry the risk that it lands close to a voice you already cast.

Audition candidates against a real line from the scene they will play rather than neutral prose, and audition them next to the voice they will share that scene with. Two voices that sound distinct on their own can blur when they alternate.

The practical ceiling is the listener rather than the software. Every speaking character in a project can be assigned its own voice; what runs out first is the audience's ability to hold them apart in one exchange. How to narrate dialogue-heavy scenes goes into keeping a fast back-and-forth readable once the cast is set.

If you are unsure, cast small and add. Giving a part its own voice after you hear that it matters is easy. Pulling a voice back out after listeners have learned it is the harder direction.

When the cast is settled, AudioProducer generates the finished audio and gives you the file. Where it goes after that is up to you. The full guide index lists every guide we have published, casting and production guides included.

Frequently asked questions

How many voices does an audio drama need?
Fewer than the character list suggests. Count how many characters speak in each scene, take the largest of those numbers, and cast that many distinct voices. A twenty-name script often resolves to somewhere around six to ten voices once the parts that never share a scene are doubled onto voices you already cast.
Can one voice play more than one character in an audio drama?
Yes. You assign a voice per character in AudioProducer, so the same voice can be assigned to two parts. Keep the doubled parts in separate scenes rather than in the same exchange, and give them different weight where you can, so the listener hears two people instead of one voice repeating itself.
Does an audio drama need a narrator?
No. Some dramas keep a narrator to carry description between scenes and others let dialogue and sound cues do that work. Dropping the narrator does not reduce the cast; it pushes description into speaking parts and into your sound cues, so scenes tend to gain speakers rather than lose them.

Related posts