You play a recording in Chinese. The speaker begins, and about half a second later you start speaking too — not waiting for a pause, not repeating after them, but talking over the recording, following just behind like a shadow.
That is shadowing. It sounds simple. Try thirty seconds of it and you will find it surprisingly hard — and the difficulty is exactly where the benefit lives.
It did not come from a language classroom
Shadowing arrived in language learning late. It began as a research instrument.
In 1953 Colin Cherry used shadowing to study attention: he played different speech into each ear and asked people to repeat one of them. That is the experiment that gave us the cocktail party effect, the ability to pull one voice out of noise.
Twenty years later William Marslen-Wilson found something striking. Some people could shadow at a lag of roughly a quarter of a second — and while doing it, they spontaneously corrected grammatical errors in the original. At that latency they were plainly not echoing sound mechanically. They were understanding, predicting and processing meaning, all inside 250 milliseconds.
From the laboratory it moved into conference-interpreter training, where listening and speaking at once is the job. Only later did it become a widespread self-study technique, especially in Japan, where the researcher Shuhei Kadota has studied it systematically for years.
So what does it train?
The most important thing to understand: shadowing is not a speaking exercise. It is a listening exercise.
When you have to reproduce a stream of sound half a second after hearing it, there is no time to translate, analyse or look anything up. You are forced to process the audio at real speed. What gets trained is sound decoding — the stage where most learners are stuck.
Kadota places shadowing inside Alan Baddeley's model of working memory, specifically the phonological loop — the mechanism that holds sound in mind for a few seconds by silently rehearsing it. Shadowing essentially takes that loop, moves it outside the head, and puts it under time pressure. With practice, sound decoding becomes automatic, and once it is automatic, cognitive resources are freed for meaning.
That is why someone can know enough vocabulary and still fail to keep up in conversation: the bottleneck is not knowledge, it is processing speed.
What the research shows, and how strongly
It is worth being straight about the strength of the evidence.
Yo Hamada reviewed the studies and drew a notable conclusion: the clearest effect is on listening, and it is clearest for lower and intermediate learners. For advanced learners the benefit narrows — presumably because their sound decoding is already automatic enough.
For pronunciation the evidence is also positive: Foote and McDonough found that shadowing on a mobile phone improved learners' comprehensibility and fluency.
Keep it in proportion, though. Most of these studies are small, run over weeks or months, and cluster in Japanese learners of English. This is a technique with a reasonable basis, not one proven across every setting. Treat it as a good tool for a specific job rather than a miracle.
What shadowing cannot do
This is the part most often skipped, and the reason many people practise for months without progress.
It will not teach you vocabulary. If you do not know a word, saying it fifty times gives you a sound, not a meaning. Words have to be learned somewhere else.
It will not teach you grammar. You can reproduce a structure perfectly and still not know why it works, or be able to build a new sentence with it.
It does not replace real speaking. Following someone else's script does not train the ability to produce your own — which only comes from having to express your own meaning, in real time, to someone waiting for an answer.
And if the material is too hard, shadowing collapses into parroting. You make the noises without understanding. Your mouth gets exercise; the language processing that matters never switches on. This is the commonest mistake.
How to practise it properly
Choose material you already almost fully understand. This is the most important rule and the most often broken. If you have to stop and work out the meaning, the passage is too hard to shadow. Drop a level.
Short passages, many repetitions. Thirty to sixty seconds repeated ten times beats ten minutes played once. The goal is automaticity, and automaticity only comes from repetition.
Work through four stages. Understand it by listening first. Then read aloud along with the text (chorus reading — not yet shadowing). Then shadow with the text in front of you. Finally shadow without it. Skipping a stage is how parroting creeps in.
Record yourself. You essentially cannot hear your own errors while producing speech. Playing back your recording against the original is where most pronunciation progress actually happens.
Chase the rhythm before the sounds. Early on, prioritise matching the intonation, the pauses and the stress, even if individual sounds slip. Wrong rhythm hurts a listener more than one imperfect vowel.
Fifteen minutes a day. This is concentration-heavy work. Short and regular beats long and occasional.
What it fixes for a Vietnamese speaker specifically
For English: rhythm. Vietnamese is syllable-timed — each syllable gets roughly equal weight. English is stress-timed: stressed syllables stretch, and unstressed ones collapse into a faint /ə/. This is why a Vietnamese speaker can pronounce every word correctly and still be hard to follow. Shadowing is one of the few ways to train it, because matching someone's timing is matching their rhythm.
For Chinese: tones in the stream. Vietnamese speakers start with a real advantage, being used to a tonal language. But Mandarin tones in running speech are not the tones in the textbook table: there is sandhi (two third tones in a row shift), there are neutral tones, and everything compresses with speed. Pinyin on the page shows none of it. Shadowing teaches the melodic line of a whole sentence rather than tones one at a time.
For any language: what writing hides. Where speakers link words, where they swallow sounds, where they hesitate. No written material teaches this. Only the ear does.
Where to start
Shadowing needs only one thing: a natural recording, short enough, with a transcript to check against.
The A1 dialogues in the Greek section of this site are built exactly that way — every line has its own audio, its transcript, a transliteration and a translation, plus a 0.75× slow toggle for the early stages. Listen for meaning first, then read along, then shadow line by line, then run the whole conversation.
And it is worth remembering why the technique works at all. Listening and speaking are biologically primary — machinery the species evolved to use. Shadowing teaches you nothing new about the language. It simply runs that existing machinery at the speed it was built for.