GetMeAudio GetMeAudio

Learn: Text to Speech and Voiceover Terms

Short, plain definitions of the terms you run into when making a voiceover, each with how it works in GetMeAudio.

What Is Text to Speech?Text to speech (TTS) is technology that converts written text into spoken audio. You give it a script and a voice, and it returns a recording of that voice reading the script aloud.What Is an AI Voice Generator?An AI voice generator is a tool that turns written text into natural-sounding speech using machine-learning models. It is the everyday name for neural text to speech packaged as an app or website.What Is Neural Text to Speech?Neural text to speech is speech synthesis that uses deep neural networks to generate audio directly from text. It usually sounds smoother and more natural than older rule-based or stitched-together methods.What Is SSML (Speech Synthesis Markup Language)?SSML, or Speech Synthesis Markup Language, is an XML-based markup language that tells a text-to-speech system how to speak: where to pause, how to pronounce a word, and how to change rate, pitch or volume.What Is Auto-Ducking in Audio?Auto-ducking is an audio technique that automatically lowers the volume of one sound, usually background music, whenever another sound, usually a voice, is present, then brings it back up in the gaps.What Is a Pronunciation Dictionary (Lexicon) in Text to Speech?A pronunciation dictionary, also called a lexicon, is a list of words with instructions for how a text-to-speech voice should say each one. It fixes names, acronyms, brand terms and unusual words the voice would otherwise get wrong.What Is Pay-As-You-Go Text to Speech?Pay-as-you-go text to speech means you pay for the audio you actually generate or download, instead of paying a fixed monthly or annual subscription.What Is a Voiceover?A voiceover is a recorded voice that is heard over video, slides, animation or other content without the speaker being seen. The audience hears the narration while watching something else.MP3 vs WAV: What Is the Difference for Voiceovers?MP3 is a compressed (lossy) audio format that makes small files, and WAV is an uncompressed format that keeps all the audio data and makes much larger files. For speech, MP3 is the practical everyday choice and WAV is the choice for editing.What Are SRT and VTT Subtitle Files?SRT and VTT are plain-text subtitle formats. Each one lists subtitle lines together with the start and end time at which each line should appear on screen.What Is Text Normalization in Text to Speech?Text normalization is the step in text to speech that converts things that are not plain words, such as numbers, dates, currencies and abbreviations, into the words that should actually be spoken.What Is Speech Rate? Words Per Minute for VoiceoversSpeech rate is how fast speech is delivered, usually measured in words per minute (wpm). Conversational English is often quoted at about 120 to 160 wpm, and narration commonly sits around 150 wpm.

Looking for a specific answer? See questions and answers.