UNTOLD · Mind · NO. M01

The Syllable That Isn't There

A dubbing accident from 1976 revealed that your eyes quietly edit every word you hear.

Share
The Syllable That Isn't There

Play a short clip of a face speaking. Once with your eyes open, once with them closed. The audio file does not change. Not by a single millisecond, not by a single sample. And yet the syllable you hear shifts depending on whether you are looking.

With your eyes shut, you hear a clear “ba.” Open them, watch the mouth, and the same sound becomes “da.” Same speaker, same volume, same waveform pressing against your eardrum. Something between the ear and the awareness has quietly rewritten the signal.

Most of us carry an intuition about hearing that turns out to be wrong. We imagine it as a sealed channel: sound enters through the ears, travels inward, and arrives at consciousness untouched. The eyes, in this picture, are spectators. They have nothing to do with what we hear. But they do. Under the right conditions, watching a mouth can override the incoming sound completely, and it can do so without your permission and without your awareness. The phenomenon has a name. It is called the McGurk effect, and its most unsettling feature is this: once you know it exists, you still cannot turn it off.

An accident in a dubbing room

The discovery was not planned. In the mid-1970s, a developmental psychologist named Harry McGurk was studying how infants learn to link faces with voices, how a baby comes to understand that the moving mouth in front of it belongs to the sound it is hearing. To probe this, McGurk and his research assistant John MacDonald wanted to test what happened when the two cues disagreed. They needed a video where the visible mouth movement and the audible speech did not match.

So they asked a technician to do something simple. Take a film of a person saying one syllable and dub over it the recording of a different syllable. The face on screen would articulate “ga.” The sound played would be “ba.” A mismatch, nothing more, meant only as an experimental manipulation, a control condition on the way to studying babies.

When the finished tape came back and they played it, nobody in the room heard what was on either track. They did not hear the “ba” from the audio. They did not hear the “ga” from the video. Every listener reported a third syllable, one that appeared on neither the tape nor the soundtrack: “da.” A phantom consonant, conjured out of the disagreement between the two senses.

McGurk, according to the accounts of the episode, did the obvious test. He closed his eyes. The illusion vanished at once and the sound reverted to its true form, “ba.” He opened them again and the phantom “da” snapped back instantly. The stimulus on the tape was fixed and unchanging. The only variable was whether his eyes were open. The accident had turned into the finding.

Hearing lips, seeing voices

McGurk and MacDonald recognised that they had stumbled onto something larger than a quirk of one dubbed tape, and they ran it as a proper experiment. They built sets of these mismatched clips and tested them on children and adults, asking each person to report what they heard 1. The results were striking in both their size and their strange developmental slope.

The overwhelming majority of adults experienced the fusion. In their original work the illusion captured nearly all grown listeners, who reported the phantom syllable rather than the true audio. What surprised the researchers was that the effect was weaker in young children and stronger in adults. Ordinarily we expect illusions to fade with maturity, to be the kind of thing a sophisticated adult brain sees through. Here the opposite held. The more years a person had spent watching human mouths make sounds, the more thoroughly their brain leaned on the visual cue, and the harder the illusion was to resist. Experience did not inoculate against the effect. Experience deepened it.

The paper appeared in Nature in December 1976 under a title that has since become almost as famous as the finding itself: “Hearing lips and seeing voices” 1. It went on to become one of the most cited results in the science of perception, reproduced in laboratories around the world and across many languages, though its strength varies from one language and one listener to another. Fifty years on, it remains a standard demonstration that perception is not a passive recording of the world.

Why the brain builds a compromise

To understand why the brain does this, it helps to abandon the idea that speech is fundamentally a sound. Speech is, first, a physical action. When a person speaks, they are moving a set of soft and hard structures inside the mouth in precise ways. The lips press together and release. The tongue rises to touch the ridge behind the teeth, or draws back toward the throat. The result is a stream of air shaped into consonants and vowels. Sound is the product of that action, its acoustic shadow, but the action itself is visible.

Your brain has spent your entire life watching mouths perform this choreography while listening to the sounds it makes. Over years of face-to-face conversation it has absorbed the statistical rules that tie one to the other. It learned that a closed-lip pop, a sudden release of pressure from sealed lips, produces “ba” or “pa.” It learned that a sound formed with an open mouth and constriction at the back of the throat produces “ga” or “ka.” These associations are so overlearned that they operate beneath thought. When you watch someone talk, you are not merely listening. You are also, unconsciously, reading the mouth.

Now imagine the two sources of information disagree. The ears deliver an unambiguous “ba,” a sound that could only come from lips closing. But the eyes report that the lips never closed. They saw an open mouth, a back-of-the-throat gesture. The brain is now holding two facts that cannot both be true. Lips cannot be simultaneously sealed and open. Faced with this contradiction, the brain does not simply pick a winner and discard the loser. It looks for a single interpretation that could plausibly satisfy both witnesses.

And there is one. The syllable “da” is articulated with the tongue against the ridge behind the teeth, a place roughly between the lips and the throat. It is a compromise consonant, acoustically related to the “ba” the ears report and visually compatible with the open-mouthed gesture the eyes saw. So the brain settles on it. It manufactures a percept that neither sense delivered on its own, because that invented percept is the best available explanation for the conflicting evidence. What you experience as hearing is, in this moment, an act of inference.

The most humbling detail is the speed. The fusion happens within roughly the first couple of hundred milliseconds of processing, faster than conscious attention can intervene. By the time the syllable arrives in your awareness, the negotiation is already over. You never witness the raw “ba.” You are handed only the verdict.

The fold in the brain that edits sound

If hearing were the sealed channel we imagine, there would be no place in the brain for the eyes to interfere with it. But there is such a place, and it has been located with reasonable precision. Much of the evidence points to a region called the superior temporal sulcus, a long fold running along the side of each hemisphere. It is one of the brain’s integration hubs, a site where information from different senses converges and is stitched into a unified experience 2.

When researchers placed people inside MRI scanners and played them McGurk clips, watching to see which parts of the brain responded, the superior temporal sulcus was consistently among the regions that lit up during the fusion, particularly when the audio and visual cues conflicted and the illusion took hold 23. The pattern of activity there tracked whether a given person actually experienced the effect. Studies have also shown that briefly disrupting this region can weaken the illusion, which suggests it is not merely a bystander but part of the machinery producing the fused percept 3.

The implication reorders our folk model of perception. This region does not wait for the auditory system to finish decoding a sound and then politely add a visual footnote. It takes the sight of the moving mouth and the incoming sound and treats them as a single event to be interpreted together. The visual signal, which travels along its own pathways, is available to shape the outcome before the sound has been fully resolved into a syllable. In other words, the mouth you see is folded into the sound before the sound reaches you. What you call hearing is the output of that blend, not the input to it.

This is why the effect cannot be switched off by knowing about it. The fusion happens in a subsystem that runs automatically, below the level where conscious knowledge lives. You can understand the trick completely, explain it to a friend, and still hear “da” the instant you open your eyes. Your higher reasoning has no veto over what your perceptual system hands it. It only gets to comment afterward.

Your best guess, not the sound

Here is the quiet reversal at the centre of all this. You are not, in the strict sense, hearing sound. You are hearing your brain’s best guess about what produced the sound, a guess assembled from every sense that has an opinion, with vision often getting a louder vote than we ever suspected. Ordinarily this system is invisible because the senses agree. The eyes and ears corroborate each other, the guess matches the input, and the whole process feels like simple, direct perception. The McGurk effect is valuable precisely because it forces a disagreement and drags the machinery into the light.

Once you see this, familiar experiences start to make new sense. Consider the way you behave in a loud restaurant. When the noise rises and speech becomes hard to follow, you lean in and fix your gaze on the speaker’s lips. You may feel as though you are simply concentrating harder on the sound, but that is not the whole story. You are recruiting your eyes to help decode the speech, using the visible articulation to fill in what the ears cannot resolve. This visual assist is measurable: seeing a talker’s face can improve comprehension of degraded or noisy speech substantially, an everyday version of the same integration that, when the cues clash, produces the illusion 4.

The same principle explains a subtler discomfort. Dubbed films, where an actor’s mouth forms one language while the soundtrack speaks another, often feel faintly, indescribably wrong even to viewers who cannot say why. The eyes are reading one set of articulations while the ears receive another, and the brain registers the mismatch it is constantly trying to reconcile. Video calls carry a version of the same problem. When audio and video drift out of sync by even a fraction of a second, the mouth and the sound no longer belong to the same instant, and the disagreement can quietly degrade how well you understand the speaker. In each case the culprit is the same fact we tend to forget: hearing was never sealed off from sight.

There is also a developmental thread worth pulling. The reason McGurk began this work at all was to study how infants learn to bind faces to voices. The illusion he found by accident turns out to be a window onto that very process. A baby’s growing sensitivity to the correspondence between seen mouths and heard sounds is part of how it learns language in the first place, matching the shapes it sees to the sounds it hears until the two become one perception. The adult who cannot escape the McGurk illusion is, in a sense, paying the price of having learned this lesson too well. The fusion is not a bug in an otherwise clean system. It is the visible edge of a talent so useful the brain performs it constantly, on every face we have ever watched speak.

Coda

We tend to trust our senses as reporters, faithful witnesses relaying the world as it is. The McGurk effect is a reminder that they are something closer to editors, taking in raw material from several sources and returning a single, coherent, sometimes invented account. The sound on the tape never changed. Only the story your brain told about it did, and it told that story before you had any say. The next time you watch a mouth move as someone speaks, it is worth remembering what is really happening in that ordinary moment. You are not simply receiving their words. Somewhere behind your awareness, faster than thought, your brain is deciding what you hear.

Watch the companion essay on YouTube
— Companion videoThe same essay, told visually. About seven minutes.

Sources

  1. McGurk, H. & MacDonald, J., “Hearing lips and seeing voices,” Nature, 1976. — https://www.nature.com/articles/264746a0
  2. Calvert, G. A. et al., “Activation of auditory cortex during silent lipreading,” Science, 1997. — https://www.science.org/doi/10.1126/science.276.5312.593
  3. Beauchamp, M. S. et al., “fMRI-Guided Transcranial Magnetic Stimulation Reveals That the Superior Temporal Sulcus Is a Cortical Locus of the McGurk Effect,” Journal of Neuroscience, 2010. — https://www.jneurosci.org/content/30/7/2414
  4. Sumby, W. H. & Pollack, I., “Visual Contribution to Speech Intelligibility in Noise,” Journal of the Acoustical Society of America, 1954. — https://pubs.aip.org/asa/jasa/article/26/2/212/730552
  5. Rosenblum, L. D., See What I’m Saying: The Extraordinary Powers of Our Five Senses, W. W. Norton, 2010. — https://wwnorton.com/books/9780393339604
  6. Tiippana, K., “What is the McGurk effect?” Frontiers in Psychology, 2014. — https://www.frontiersin.org/articles/10.3389/fpsyg.2014.00725/full

Related reading

More from the Mind edition →