How to Make an Anime Fandub With AI Voices
A fandub is the opposite problem from most voice work. Normally the audio sets the pace and everything else follows it. In a fandub the picture is already finished, already cut, and already the exact length it is going to be -- so every line you write has a window it has to live inside, and the window does not negotiate.
This is a walkthrough of working inside that constraint: scripting to timecode, fitting lines to windows you cannot move, holding a cast together across a season, and sitting your dialogue against a soundtrack that already exists. If what you actually have is a comic rather than animation, the panels wait for you and the whole method inverts -- read how to make a comic dub instead, and the full article index has the rest of the cluster.
What a fandub asks that other voice work does not
A fandub is a new vocal track laid over animation that already has one, or over animation that shipped without dialogue in your language. The art is locked. The cuts are locked. The runtime is locked to the frame. What you are producing is a track that has to land in the gaps the animation left.
That single fact reorganizes the job. You do not decide how long a beat is; you measure how long it already is. You do not pace a scene; you fill one. A line that runs four seconds in a three-second window is not a stylistic choice you can defend -- it is out of sync, and viewers hear it immediately even when they cannot say why. Most of the craft below is about finding out how much room you have before you write anything to put in it.
Rights, said plainly
Fandubs of licensed shows are fan work. That is not a technicality to route around, and this post is not going to pretend otherwise. What follows is the craft -- how to make a dub that sounds produced. Whether you can redistribute somebody else's animation with your track on it is a question about that animation, that rights holder, and where you intend to post it, and it is your call to make with your eyes open.
Three cases are simply clean, and they are worth knowing before you pick a project. Animation you made yourself. Animation released under a license that permits derivative works. And a dub made with the studio's or the animator's explicit permission, which independent creators grant more often than people expect when you ask first and say where it will be posted. Everything below works identically in all three, which is a decent argument for starting there.
Build a timecode script before you cast anything
Open the episode in whatever video editor you use and go through it once with a text file beside you. For every line of dialogue, write down three things: who says it, when it starts, and when it ends. You end up with something that reads like a play with stopwatch readings in the margin.
Do this pass before you write a word of dialogue, not after. The windows are the constraint, and writing to them is a different activity from writing and then discovering they do not fit. It is also where you find the free space -- lines delivered off-screen, from behind, at a distance, or under a cut away from the speaker's face. Those have no lip movement to contradict, so they can run long or short without anyone noticing. In a typical episode that is a surprising fraction of the dialogue, and it is where you spend the room you do not have elsewhere.
Fitting a line into a window that will not move
Speech runs at roughly two to three words a second at a normal conversational clip. A two-second window is about five words. Once you have the timecode script, that arithmetic does most of your rewriting for you: if the window is two seconds and your translation is twelve words long, the translation is wrong for this medium, however faithful it is.
So you cut. Drop the address term, drop the hedge, drop the second adjective, turn the subordinate clause into a second sentence you can drop entirely if it still overruns. The discipline is closer to subtitling than to prose -- you are preserving the meaning of the line, not its shape. When a line runs short instead, an inline pause placed at the start holds the delivery until the mouth actually moves, which is how you use silence as placement rather than as pacing.
The four pause levels are worth knowing here, because a fandub uses them differently from an audiobook. Inline pauses place individual lines against picture. The project-wide default paragraph pause, which normally does the pacing work, you usually want set low or near zero -- the animation supplies the gaps already, and a default pause stacked on top of an existing beat is the most common way a first fandub drifts a second late and never recovers.
Mouth flaps, and how close is close enough
Nobody hits every flap. Professional dubs do not either, and the ones that sound right are matching three things rather than thirty: the line starts when the mouth starts, the line ends when the mouth stops, and the stressed syllable lands somewhere near the biggest mouth movement. Get those and the brain fills in the rest.
Emotion tags do more work here than in almost any other format, because the picture is already telling the viewer what the character feels. A face mid-shout over a flat, level read is actively worse than a plain delivery with no picture at all -- the mismatch is what the viewer notices. Tag the line angry, or afraid, or calm, and let the delivery agree with the frame instead of fighting it.
One voice per character, held across a season
Give every speaking character its own voice, separate from the narrator, chosen from the voice library and previewed before you commit. The library runs to well over a hundred voices and keeps growing, with a usable spread of ages, accents and character-flavored styles. Budget real listening time for a cast of six -- picking by name and hoping is how two characters end up indistinguishable in episode three.
Consistency across episodes is the part that separates a dub project from a one-off. Make one project per episode, then pull the character list, voices and settings included, from the previous episode's project rather than recasting from memory. A series that quietly re-voices its lead halfway through is the single fastest way to lose the audience that stayed for it. And if you want to voice one character yourself and let the platform handle the rest of the cast, clone your voice once and use it like any other voice in the library.
Sitting your dialogue against a track that already exists
The episode already has music and effects, mixed to sit under a vocal track that you are now replacing. If you can get a clean music-and-effects version -- some releases ship one -- use it and the job gets much easier. Usually you cannot, and you are working over a full mix with the original dialogue still in it.
The honest answer is that this is a video editing problem more than an audio generation one: duck the original under your lines, keep the existing music and effects where they are, and resist the urge to add your own layer on top. Automatic sound assignment is built for scenes with nothing under them yet, which is exactly not this. Generate dialogue, download each episode as its own file, and do the balancing in your editor against the track that is already there.
The last pass, against picture
Lay the finished audio against the episode and watch it straight through without stopping. You are looking for three failures specifically: a line that starts before or after the mouth does, a character whose voice has drifted from earlier episodes, and a delivery that disagrees with the face on screen. Fix those and stop. A fandub that is tight on entrances and consistent in casting reads as produced even where individual syllables miss.
Then watch it once more with the picture minimized and only the audio playing. If you can follow who is speaking without seeing them, the casting is doing its job.
The cleanest place to try this is animation you own, or animation you made. Take one scene, build the timecode script, cast it, and hear whether the timing method holds before you commit to a whole episode. The free account gives you 1,200 words a month, no credit card, which is enough for one scene end to end. Start a project and find out on two minutes of footage rather than twenty-two.
Frequently asked questions
- How do I know how long each line can be?
- Measure before you write. Go through the episode in a video editor and note, for every line, who says it and the timecodes where the mouth starts and stops. Speech runs roughly two to three words a second, so a two-second window is about five words. Write the translation to fit the window rather than translating first and discovering it overruns.
- Is it legal to fandub an anime I do not own?
- A fandub of licensed animation is fan work, and redistributing someone else's animation with your track on it is a call you have to make with the rights holder and the platform in mind. Three cases are clean: animation you made yourself, animation released under a license permitting derivative works, and a dub made with the creator's explicit permission. The method is identical in all of them.
- How do I keep a character sounding the same across a whole series?
- Make one project per episode and import the character list, voices and settings included, from the previous episode instead of recasting from memory. A series that quietly re-voices its lead halfway through loses the audience that stayed for it. You can also clone a voice if you want to perform one character yourself and let the platform cast the rest.
- What do I do about the music and sound effects already in the episode?
- Leave them where they are and do the balancing in your video editor -- duck the original mix under your lines rather than rebuilding the soundscape. Automatic sound assignment is designed for scenes with nothing underneath them yet, which is the opposite of a fandub. Generate the dialogue, download each episode as its own file, and mix against the existing track.