How to Handle Foreign Words and Phrases in AI Narration
A single French endearment in a romance manuscript. Two lines of Spanish in a thriller. A whole invented tongue in epic fantasy. All three arrive at the narration engine as plain text, and the engine has to work out what language that text is in before it can say any of it out loud.
AudioProducer.ai auto-detects the language of your source text and generates audio in the language it detects. English is the language we officially support. Other languages do work, we do not guarantee the quality outside English, and the detection itself can get things wrong. That last part is what actually bites authors, because a short foreign phrase sitting inside an English paragraph is the hardest case a detector ever sees. There is almost nothing to go on.
The useful question, then, is what happens to a phrase that short, and how much of the result you can steer.
What actually happens when narration hits a non-English word
It depends almost entirely on length. A full paragraph in Spanish gives the detector plenty of signal, and you will usually get Spanish audio. Two words of French inside an English sentence usually get read by the English voice, applying roughly English phonics, which is often exactly what you wanted. An English-speaking narrator reading mon amour to an English-speaking audience does not need a Parisian accent.
The case to listen for is the middle one. A sentence or two of another language can go either way, and it can flip partway through a chapter, so the same phrase lands one way in chapter two and another way in chapter nine. Nothing is broken when that happens. The detector simply had more context in one place than the other.
The voice library includes voices tagged for multilingual use, alongside English voices carrying Spanish, Italian, Irish, Indian and other accents. Browse and preview them on the Voices page on your home screen, and treat a non-English result as something to audition rather than assume.
Respelling the source text, and where it stops working
The main lever you have is the text itself. There is no pronunciation dictionary in the editor and no field for phonetic notation, so what you type is what gets read. If you want the English voice to approximate French rather than flatten it, write the approximation: mon ah-MOOR instead of mon amour. This is the same technique that works on invented proper nouns, and we covered it at length in our guide to building a pronunciation guide for character and place names.
Two limits are worth knowing before you lean on it. Respelling changes what a reader sees, so keep a separate export copy for audio and leave your print file alone. And it only ever reaches an approximation: English letter combinations cannot produce a French uvular r or a properly rolled Spanish double r, and no amount of creative spelling will get you there.
Respell, generate the passage, and listen. The spelling is a suggestion to the engine rather than a phonetic instruction it must obey. The same loop applies to numbers, dates and abbreviations, which fail in a similar way for a similar reason.
When a passage is long enough to deserve its own voice
Past a certain length, respelling is the wrong tool. If a character speaks four lines of Spanish, you are no longer fixing a word. You are casting a performance.
Assign that character their own voice and tag their lines to that speaker in the editor. Auto-Assign Characters will give you a starting point, and it works from dialogue attribution, so it reliably misses unattributed lines and can drift inside a long back-and-forth where nobody is named for half a page. Check its work by hand on any passage that matters. Per-line speaker tagging exists precisely so you can override it.
Choosing which voice is a craft decision rather than a technical one, and it overlaps heavily with giving each character a distinct voice and accent and with handling dialects without tipping into caricature. Match the speaker rather than the language. A Mexican character and a Castilian character are not interchangeable just because both passages happen to be Spanish.
Invented languages, where you set the pronunciation because nobody else can
Constructed languages are the easy case, which surprises most fantasy authors. There is no correct pronunciation to get wrong. The detector has nothing to latch onto and will generally read your elvish as English-flavoured sound, and since you invented the words, that reading is as legitimate as any other.
What matters here is consistency across a long series. Decide once how each recurring term sounds, write it phonetically in your audio export copy, and keep the list somewhere you will still find it two books later. Readers who love a constructed language notice when book three stops matching book one.
Italics do not survive the trip to audio
Print marks a foreign phrase with italics, and that convention carries a lot of quiet information: this is another language, hold it slightly apart from the sentence around it. Audio has no italics. The formatting is gone the moment the text becomes sound.
What replaces it is timing. A small pause on either side of a phrase does more to mark it as foreign than anything you can do to the word itself. AudioProducer.ai gives you pause control at several levels: a project-wide default under Edit Project, a per-paragraph override, and inline pauses for a single moment inside a line. Per-line dialogue emotion helps when the phrase is doing emotional work rather than sitting there as texture. For the wider view on shaping delivery this way, directing AI narration like a producer covers the whole toolkit.
Audition the phrase before you commit the book
Every decision above is cheap to test and expensive to guess at. Preview voices on the Voices page, then generate one short section that actually contains the phrase and listen to the MP3 on the speakers your readers will use. The free tier gives you 1,200 words with no card, more than enough to hear how a paragraph of French lands before you render three hundred pages around it.
You get an MP3 file to download and keep. We do not distribute, list or upload it anywhere, so where the finished audiobook goes afterwards stays entirely your call. And if the non-English content is the point of the book rather than a garnish in it, building a language-learning audiobook is a different job with a different setup, and worth reading before you start.
The whole of it comes down to this. Short phrases usually get read by your English voice, and that is normally fine. Respell anything you want said differently, because the source text is the control you have. Give a genuinely bilingual character their own assigned voice instead of fighting the spelling one line at a time. Let pauses do the work italics used to do. Then listen to a sample before you commit, because the detector will occasionally surprise you, and being surprised by one paragraph beats being surprised by a finished book.
A phrase in another language is one of several things on the page that is not plain prose. Narrating footnotes and endnotes works through the same problem from the reference side.
Frequently asked questions
- Will AudioProducer.ai read a French phrase in French?
- It might. The engine auto-detects the language of your source text and generates audio in the language it detects, so a long passage usually comes out in that language while a two-word phrase inside an English sentence often gets read by your English voice. English is the language we officially support and we do not guarantee quality outside it, so generate the passage and listen to it before committing a whole book.
- Can I set the pronunciation of a single word?
- There is no pronunciation dictionary or phonetic notation field in the editor. The control is the source text itself: respell the word the way you want it said, generate that section, and listen. This is the same approach that works for invented character and place names, and it reaches a good approximation rather than a perfect native accent.
- What if a whole character speaks another language?
- Assign that character their own voice rather than trying to fix the spelling line by line. Tag their lines to that speaker in the editor and pick a voice from the library that fits the character. Auto-Assign Characters gives you a starting point, and because it works from dialogue attribution it will miss unattributed lines, so check the tagging by hand before you render.