It is late, the laptop is open, and your favourite scene is paused mid-line. You hit record and say the line yourself: same rhythm, same crack in the voice, same pause before the punchline. Then you play it back and laugh loud enough that someone in the next room asks if you are alright.
Dubbing in your own voice means taking an existing scene and saying the lines again yourself, matching the timing and the delivery of the original. It looks like a five-minute joke. It is a five-minute joke. It also has a side effect: hearing speech and repeating it out loud right away is a language-teaching method called shadowing (you follow the voice like a shadow), and interpreters have trained with it for decades.
At Awe Reads we gathered what the research actually says, then built a browser game where you can try it yourself - Awe Voice. In order: why the first playback almost always makes people laugh, what it does to speech and memory, and why the game refuses to let you hear yourself until the very end.
What dubbing in your own voice actually is
It is re-voicing existing footage with your own voice instead of the original one. You are not writing a script or inventing a parody: the lines already exist, and the job is to land in them. In the tempo, in the pauses, in the emotion.
Amateur re-voicing has its own name: fandub, short for fan dubbing. Anime scenes, sitcom clips and memes get redubbed at home, for free, because someone felt like it. Whole communities have run on that for years.
Professional dubbing works differently. In a studio the actor reads an adapted script where every line is measured against lip movement, with a director and a sound engineer in the room. At home you have none of that. You have a scene, your voice, and an attempt to catch somebody else's rhythm. The gap between those two is exactly where the comedy lives.
And the biggest laugh usually arrives in the first second of playback, for a reason that has nothing to do with your acting.
Why your recorded voice sounds like a stranger
Because a recording drops half of what you normally hear. When you speak, the sound reaches your inner ear along two routes. One is air conduction: your voice leaves your mouth and comes back to you through the air, like any other sound. The other is bone conduction: vibration travels from your vocal folds straight through the bones of your skull.
A microphone can only capture the first route. The second one is yours alone, and you have spent your whole life counting it as part of your voice.
Swedish engineers from Chalmers University of Technology and Linkoping University measured both routes in 20 people and published the results in the Journal of the Acoustical Society of America (2010). Between roughly 1 and 2 kHz, the bone route dominated for most speech sounds - a real share of what you think your voice sounds like never travels through air at all.
That is why recordings sound thinner and higher than the voice in your head. Nothing is wrong with the microphone: everyone else has always heard the recorded version. The longer breakdown, with more studies, is here: why your recorded voice sounds wrong to you.
That explains the laughing. It does not explain why ten minutes of this changes how your mouth handles a language.
Dubbing is shadowing, only funnier
Shadowing means repeating speech almost simultaneously with the speaker, without waiting for the sentence to finish. The textbook definition calls it a paced tracking task: you vocalise what you hear the moment you hear it, at the speed of the original.
Yo Hamada, a language teacher and researcher at Akita University in Japan, spells out the difference between shadowing and ordinary repetition (Contact Magazine, TESL Ontario, 2018). In repetition you hear a chunk, understand it, then say it. In shadowing there is no time to understand anything - almost all your attention goes to the sound itself. The method came out of interpreter training, where nothing slower would work.
Dubbing a scene runs on the same machinery, with a picture attached and no textbook. You hear the line, catch its melody, and produce your own version on top. The difference is motivation: shadowing a coursebook recording gets old in four minutes, while re-voicing a character yelling at a dragon makes people go again, and again.
What research says about dubbing and speech
Small classroom studies keep finding gains in pronunciation and listening, with modest claims because the samples are modest. Three studies, three continents.
- Japan. Yo Hamada (Language Teaching Research, 2016) ran nine shadowing lessons with 43 university learners of English. Perception of individual speech sounds improved across proficiency levels, while listening comprehension gains showed up clearly in the lower-proficiency group.
- Vietnam. Phan, Nguyen and Nguyen (TESL-EJ, 2025) had 30 sixth-graders dub video three times a week, 45 minutes a session, for ten weeks. Overall pronunciation scores rose from 14.86 to 16.50, with intonation improving most: 3.36 to 4.30 (p < 0.01).
- China. Mu and Wasuntarasophit (LEARN Journal, 2025) ran eight weeks of shadowing with 30 senior high school students. Mean listening scores went from 12.1 to 15.98 (p < 0.05), and 27 of the 30 wrote in their logs that holding attention through long recordings got easier.
The honest caveats: the Vietnamese and Chinese studies had no control group, so they compared the same people before and after, and some of the gain could come from simply getting used to the test format. Both samples were 30 people. This is not proof that dubbing replaces language classes. It is a reasonable signal that saying somebody else's words out loud, regularly, shows up in the numbers.
And the effect does not seem to rest on listening alone. Part of it comes from the plain act of opening your mouth.
Why saying things out loud sticks better
Words spoken aloud are remembered better than words read silently. Memory researchers call this the production effect.
Colin MacLeod and Glen Bodner, both memory researchers in Canada, reviewed the evidence for Current Directions in Psychological Science (2017). Their summary: producing an item in any simple way - saying it, writing it, typing it - yields substantial memory gains over silent reading. The leading explanation is distinctiveness: a produced item leaves a more specific trace, which makes it easier to retrieve later.
That is the practical payoff of the game. Dubbing a scene is not listening to a line. It is producing one: articulating it, shaping the intonation, holding the beat. For memory those are different operations, and the second one is stronger.
Which raises a fair question: does the rest of the world even watch films the way you do?
Dubbing countries and subtitle countries
The world splits roughly in half between viewers who expect to hear a film in their own language and viewers who expect to read it. Morning Consult surveyed adults in 15 countries (March 2022), and the preferences fall out like this:
- Germany, Italy, Spain, France - dubbing wins, backed by a long tradition of re-voicing foreign releases for cinemas.
- United States - close to a split decision: 43% prefer subtitles, 36% prefer dubbing.
- China and South Korea - about 70% prefer subtitles; Japan and India lean the same way.
- Mexico and Brazil - firmly dubbing territory, especially for animation and family films.
- Russia and Russian-speaking countries - 86% prefer dubbing, one of the highest figures in the survey.
There is a practical upside either way. Grow up on dubbing and your ear has spent years hearing one voice ride somebody else's mouth movements, so landing in a line comes more naturally than you would expect. Grow up on subtitles and you have heard far more original speech, which makes foreign intonation easier to copy.
How to try it in five minutes
You need three things: a browser, a microphone and your own video. Awe Voice is a free dubbing game in the Awe Reads games section, with no sign-up and nothing to install.
The flow:
- Bring a video. Your own file from a phone, a computer or a cloud drive, in MP4, MOV or WEBM. Links to other people's players are not accepted: a browser cannot pull the picture out of an external player, and without the picture there is nothing to build a file from.
- Mark the lines. Play the video and mark the short chunks you want to voice. Up to five lines, up to fifteen seconds each.
- Record. Listen to the original, then say it yourself. Hated it? Re-record as many times as you like. A level meter next to the original shows whether you are reaching the same loudness.
- Build and download. At the end the game lays your voice over the picture and hands you the finished file.
Headphones are not a nicety here. Without them the microphone picks up your speakers, and the original soundtrack ends up baked on top of your take.
Now the promised part. You do not hear your own recording until the very end - not after the first line, not after the third. That is deliberate. The comedy comes from the gap between what you performed and what actually came out, and letting people replay each take turns the game into a voice recorder, where everyone polishes intonation instead of playing.
Where your voice goes
Nowhere. Neither the video nor the recording reaches a server: everything runs inside your browser tab and disappears when you close it.
That is a property of the technology, not a promise on a marketing page. The browser recording interface, MediaRecorder (the standard way a site captures microphone audio), hands data straight to the page as a file in memory. Nothing is transmitted by default - a site has to explicitly send it somewhere. Awe Voice does not, which is also why there is no share link: the result exists only as your file.
The contrast matters more than it used to. Modern voice cloning needs only a few seconds of audio to build a copy of how you sound, as covered here: how many seconds it takes to clone a voice. A game that never collects the recording is not a small detail in that world.
How to pick a scene that works first time
Take a short clip with strong emotion and clear articulation. Four traits separate an easy first attempt from a frustrating one.
- One speaker, not an argument. Three lines of back-and-forth sound tempting, but you end up chasing somebody else's pauses and arriving late. A ten-second monologue forgives far more.
- A close-up. When the face fills the frame you can see the mouth move and adjust almost without thinking. On a wide shot there is nothing to aim at, and timing becomes pure listening.
- Speech without a music bed. If loud score sits under the line, your voice lands on top of it and the result sounds muddy. A plain dialogue scene wins.
- A line you already know by heart. A familiar line frees your head: instead of recalling words you can play the intonation. Shadowing classes exploit the same trick by using familiar texts.
Animation is the friendliest starting point. The dialogue was recorded separately from the picture and performed with deliberate exaggeration, and drawn mouths forgive the small misses a live actor would expose.
Four beginner mistakes
Nearly everyone trips over the same things, and nearly all of them take a minute to fix.
- Picking a line that is too long. Fifteen unbroken seconds at somebody else's pace is a lot. Start with five to seven seconds: a short accurate take beats a long one you fumbled twice.
- Recording too quietly. The most common technical problem. A laptop microphone sits low and far away, so a normal speaking voice arrives barely audible. Get closer, push a little louder than feels natural, and watch the level meter.
- Copying the timbre. Copy the rhythm, the pauses and the melody of the line, not the actor's voice. Your timbre is yours, and it is the reason the result is funny at all.
- Performing at half power. A cautious take sounds flat even when the timing is perfect. Voice actors overplay on purpose: without a face on screen, the voice has to carry the whole emotion alone.
One more thing people forget: closing the tab wipes everything. Neither the video nor the recorded lines survive between sessions, so download the result before you close the browser.
What else Awe Reads has besides the game
Awe Reads is a publication that explains complicated things in plain language, plus a set of free tools built around the articles. The game did not appear next to the texts, it came out of them: first a breakdown of how people hear their own voice, then a way to feel that effect yourself.
What you can try right now:
- Burnout test - 13 questions, three minutes. Built on the Copenhagen Burnout Inventory, a scale published by Kristensen and colleagues in Work & Stress in 2005 and released for free use. Unlike most quizzes online, this one has measured reliability: independent validations put internal consistency around 0.9 out of 1.
- IQ test - matrices, number series, mental rotation and analogies. Items are generated fresh for every attempt, and the score arrives with a confidence interval (the range your true score probably falls in, because every measure of ability carries error).
- Games - the section just opened and Awe Voice is the first entry. Anything added later follows the same rule: it runs on your device and your files stay there.
- Articles in two languages. The English and Russian versions are written in parallel rather than machine-translated from one another.
Who makes all this, how facts get checked and what the editorial rules are: see the about page.
What to do tonight
Open your phone and find fifteen seconds you know by heart: a cartoon scene, a chunk of an interview, one line from a show. One line, no more. Record it in your own voice, then watch the result all the way through instead of stopping halfway.
You will probably laugh: that is the voice without its bone-conducted half, the one everybody except you has always heard. And you will have done exactly what interpreters train with - heard speech and produced it out loud. Five minutes of an evening, and your laptop learns nothing about you.
Sources
- Reinfeldt S., Ostli P., Hakansson B., Stenfelt S. Hearing one's own voice during phoneme vocalization - transmission by air and bone conduction. Journal of the Acoustical Society of America (2010)
- Hamada Y. Shadowing: Who benefits and how? Language Teaching Research (2016)
- Hamada Y. Shadowing for Language Teaching. Contact Magazine, TESL Ontario (2018)
- Phan V.T.T., Nguyen T.Q., Nguyen K.D. The Effects of Video Dubbing on EFL Learners' Pronunciation. TESL-EJ (2025)
- Mu Y., Wasuntarasophit S. Effects of the Shadowing Technique on English Listening Comprehension. LEARN Journal (2025)
- MacLeod C.M., Bodner G.E. The Production Effect in Memory. Current Directions in Psychological Science (2017)
- Morning Consult. Subtitles vs. dubbing: viewers in 15 countries (2022)
- Validity and Reliability of the Copenhagen Burnout Inventory (2021), on the Kristensen et al. scale, Work & Stress (2005)
- MDN Web Docs. MediaRecorder API - recording media in the browser







Те самые голоса можно сделать?