McGurk Effect

The McGurk effect is a perceptual illusion in which what you see a speaker’s mouth doing changes what you hear them say. When the lips shape one syllable but the audio plays another, most people perceive a third syllable that was in neither channel (McGurk & MacDonald, 1976).

In the classic demonstration, a voice saying “ba” is played over a video of lips saying “ga”. Most people hear “da”, a sound that appeared in neither channel.

Closing your eyes restores the true “ba”. Opening them changes the sound again, even when you know the trick.

A poor audio signal combined with clear visual information can produce the effect. Many other factors also influence its strength.

Key Takeaways

  • Definition: The McGurk effect is an illusion in which what you see a speaker’s lips do changes what you hear. Mismatched lips and sounds fuse into a third syllable.
  • Classic Example: Audio “ba” paired with lips saying “ga” is usually heard as “da”. Closing your eyes restores the true “ba”.
  • Discovery: Harry McGurk and John MacDonald reported the effect in 1976 after noticing it by accident during an infant speech study.
  • External Factors: Visual distraction, tactile diversion, familiarity with the speaker, and syllable structure can all change the strength of the effect.
  • Internal Factors: Brain damage, Alzheimer’s disease, Specific Language Impairment, and aphasia can alter how sight and sound combine.
  • Languages: Older studies found weaker effects in Chinese and Japanese listeners, but a large 2025 study found similar rates in Mandarin and English speakers.
  • Individual Differences: Some people almost always experience the illusion and others almost never do. Each person’s pattern is stable over time.
a woman's mouth blowing hand drawn icons and symbols close up

Origin and History

Cognitive psychologists Harry McGurk and John MacDonald introduced the effect in 1976. Their paper was titled Hearing Lips and Seeing Voices (McGurk & MacDonald, 1976).

The discovery was accidental. McGurk and his assistant, MacDonald, were studying how infants perceive language at different developmental stages. They played the video of a mother speaking in one location and the sound of her voice in another.

They then asked a technician to dub the audio syllable “ba” onto a video of a speaker saying “ga”. When the dubbed tape played, both researchers heard “da”. This phoneme matched neither the audio nor the video.

Both McGurk and MacDonald were confused. Eventually, however, they realized that the phenomenon stemmed not from a mistake by the technician but from an idiosyncrasy in human perception.

The 1976 Experiment: Fusion and Combination

McGurk and MacDonald then tested the effect systematically. Their key comparison was audio alone against audio plus a visible face.

  • Aim: To test how conflicting lip movements and sounds change what children and adults report hearing.
  • Method: Audio of one syllable was dubbed onto video of another. Pre-school children, school-age children and adults reported what they heard, first watching the face and then with audio alone.
  • Results: With audio alone, listeners identified the syllables almost perfectly. With the face visible, about 98% of adults heard “da” for audio “ba” plus lips “ga”, but only around half of the children did.
  • Conclusion: The illusion shows the brain blends conflicting sight and sound into one experience. Speech perception therefore includes a visual contribution as a matter of course.

The reverse pairing gave a different outcome. Audio “ga” with lips “ba” produced combination responses such as “bga” or “gba”, in which both consonants are heard in sequence.

The fused “da” follows from where each sound is made. “Ba” is bilabial (made with the closed lips), “ga” is velar (made at the back of the mouth), and “da” is alveolar (made in between).

The visible lips do not close, which rules out “ba”. The brain settles on an in-between place of articulation that fits both streams.

Later work confirmed the illusion is robust. It survives even when a female face is dubbed with a male voice (Green et al., 1991). The stimuli remain artificial, though, because mismatched cues almost never occur in natural speech.

How the Brain Combines Sight and Sound

The McGurk effect makes the fusion of sight and sound easy to notice. Explaining it has produced a clear neural target and several competing theories.

Why the Illusion Is Automatic

The illusion persists even when observers know the audio and video are mismatched and try to report only the sound. This suggests the brain combines the two streams early and automatically, not as a late, deliberate judgement.

This echoes the constructivist view in theories of perception: perception is an active construction, not a passive readout of the senses. Three further points follow:

  • Everyday Benefit: Lipreading, understanding speech from lip movements, improves comprehension, especially in noise, because the mouth shows where sounds are made.
  • The Cost: The McGurk effect is the price of that benefit. When the streams conflict, the brain still combines them.
  • Timing: Visual articulation usually begins slightly before the sound, and integration is strongest when the streams are roughly in sync. The effect weakens as they are pulled apart.

The Superior Temporal Sulcus

Brain imaging locates audiovisual speech integration in the superior temporal sulcus (STS). This strip of cortex on the side of the temporal lobe is where auditory and visual speech streams converge. The left STS is especially involved in speech.

Nath and Beauchamp (2012) asked why some people reliably perceive the illusion and others almost never do.

  • Aim: To find a brain region whose activity tracks both the stimulus and how likely a person is to perceive the illusion.
  • Method: Functional MRI (fMRI, a scan of brain activity) recorded McGurk perceivers and non-perceivers viewing congruent, McGurk and other incongruent audiovisual syllables.
  • Results: Only the left superior temporal sulcus responded to both the stimulus and the susceptibility group. The size of its response predicted whether a person perceived the illusion.
  • Conclusion: The left STS is a key site of individual differences in audiovisual speech perception.

This design separates stimulus-driven from perceiver-driven activity. As correlational fMRI, however, it cannot show that STS activity causes the percept.

Beauchamp, Nath and Pasalar (2010) addressed that limitation directly.

  • Aim: To test whether the STS is necessary for the McGurk effect, not merely correlated with it.
  • Method: fMRI first located each participant’s own STS speech region. Single magnetic pulses (transcranial magnetic stimulation, TMS) then briefly disrupted it while participants rated McGurk and congruent syllables.
  • Results: TMS to the STS reduced the McGurk percept but left congruent speech intact. It worked only within about 100 ms of the sound’s onset, and a control site had no effect.
  • Conclusion: The STS plays a causal, time-specific role in combining sight and sound.

Combining individual localisation with TMS upgrades the STS from correlate to cause. The samples are small, and TMS disrupts processing rather than reading it out, so the finding shows necessity, not the full computation.

The STS does not work alone. Visual speech also shortens the latency of two early brain-wave components, the auditory N1 and P2 (Alsius et al., 2014). This early marker of integration shrinks when attention is withdrawn, so the STS is best seen as the hub of a wider network.

Competing Explanations of Integration

Several theories explain why vision changes what is heard. Each answers a different question about the same phenomenon.

  • Amodal Speech: Rosenblum (2008) argues the brain recovers a talker’s intended gestures from any sense. The effect arises when two channels specify incompatible gestures.
  • Fuzzy Logical Model: Each sense gives graded evidence for candidate sounds, and a multiplicative rule combines them (Massaro & Cohen, 1983). The seen lips rule out “ba”, so “da” gets the highest combined support.
  • Bayesian Causal Inference: The perceiver first estimates whether sound and sight share one cause, and fuses them only when that probability is high (Körding et al., 2007; Magnotti & Beauchamp, 2017).
  • Unity Assumption: Signals that seem to come from one event bind more tightly (Welch & Warren, 1980; Vatakis & Spence, 2007).
  • Early Versus Late: Seeing a face lowers the detection threshold for noisy speech, which favours early integration over a late decision stage (Schwartz et al., 2004).
  • Perceptual Constraints: Boersma (2012) proposes that the fused percept is the best-satisfying candidate under a set of ranked perceptual constraints.

The two computational accounts differ in what they assume. The fuzzy logical model takes integration for granted and asks how the streams are weighted. Causal inference adds a prior decision. It asks whether to integrate at all.

This reframes the effect’s variability as differences in the estimated probability of a common cause. The unity assumption fits the illusion’s robustness, since matching articulation appears enough for perceived unity even when gender cues conflict.

The early-versus-late debate is unresolved. The views are compatible, because a system could combine the streams early for detection and again at a decision stage for categorisation.

Real-World Applications

The same integration machinery matters well beyond the laboratory. It shapes how people cope with hearing loss, masks, technology and language learning.

  • Hearing Loss: Integration matters more as hearing declines. Cochlear-implant users rely heavily on visual speech, so rehabilitation increasingly trains audiovisual listening.
  • Face Coverings: Opaque masks remove visual speech and measurably reduce comprehension in noise. The burden falls hardest on hearing-impaired listeners.
  • Speech Recognition: Massaro and Stork (1998) showed machines can combine sound and lip information using reliability-weighted logic. Modern lip-reading systems descend from this insight.
  • Dubbing and Avatars: Mismatched lips in dubbed film, laggy video calls or animated characters trigger the same machinery. The result feels effortful and overlaps with the uncanny valley.
  • Language Learning: Visual cues help listeners perceive non-native and accented speech, so seeing a face is a genuine aid to language acquisition.
  • Clinical Research: The effect probes crossmodal integrity in conditions such as Alzheimer’s disease. Only large, stimulus-controlled measurements can support claims of group differences.

External Factors Impacting the McGurk Effect

Visual Distraction

A study of the role of optic attention in audiovisual speech perception observed the McGurk effect in two different situations (Tiippana, Andersen & Sams, 2004).

Listeners attended to the speaker’s face in the first condition. In the second, they ignored the face and watched a leaf moving across it.

The McGurk effect was weaker in the second condition. The authors attributed this to visual attention modulating audiovisual speech perception.

The modulation could occur early, while each sense is still processed separately. Alternatively, attention could alter the later stage where vision and hearing are combined.

Tactile Diversion

Research has recently challenged the popular assumption that audiovisual pairing transpires in an attention-free mode (Alsius, Navarra & Soto-Faraco, 2007).

One study asked whether audiovisual speech integration draws on attention. It imposed attentional demands on the tactile domain, which is not directly related to speech perception.

The McGurk effect was measured in a dual-task design with a demanding tactile task. As attention was deflected to touch, visually swayed responses fell.

This outcome was attributed to a limit on supramodal attention, meaning attention shared across the senses. Touch drew resources away from binding sight and sound.

These findings provide a glimpse of the dynamism and the extensiveness of the interactions between crossmodal binding mechanisms and the attentional system.

Familiarity

An experiment examined the claims for the independence of facial speech processing and facial identity (Walker, Bruce & O’Malley, 1995).

The researchers manipulated the faces so that some participants knew them and others did not. The voices and faces were either congruent (from the same person) or incongruent (from different people).

Familiarity mattered. When the voices and faces were incongruent, participants who knew the faces were less susceptible than those who did not.

This outcome implies that facial speech and facial identity are far from independent. Those familiar with a speaker’s face are less likely to be swayed by the McGurk effect.

Syllable Structure

One research study conducted four experiments to decide whether optic information influences the judgments of acoustically-specified nonspeech events and speech events (Brancazio, Best & Fowler, 2006).

The study employed click sounds perceived as nonspeech by many English listeners but function as consonants in certain African languages.

The results demonstrated a significant McGurk effect for isolated clicks. This effect, however, was notably smaller than that for the stop-consonant-vowel syllables.

Moreover, strong McGurk effects were discovered for click-vowel syllables, which were similar to those for English syllables.

On the other hand, weak McGurk effects were found for excised release bursts of stop consonants in isolation; these were similar to the effects for isolated clicks.

This outcome shows that the McGurk effect may occur even in non-speech settings.

Moreover, while phonological significance is not a prerequisite for the McGurk effect, it does seem to intensify it.

Expectation and Sentence Context

The effect is not fully automatic. Windmann (2004) found more McGurk responses when the illusory word fitted the meaning of the carrier sentence than when it did not.

This shows that top-down expectations about words and meaning feed into what people hear. If integration were purely stimulus-driven, sentence meaning should not matter.

Attention matters in the same way. Attentional load reduces both the behavioural effect and the early brain response linked to integration (Alsius et al., 2014).

These modulations are real but partial. The effect survives gender-mismatched talkers (Green et al., 1991) and wide variation in viewing conditions. Integration is therefore strongly bottom-up, yet open to influence from attention and expectation.

Internal Factors Impacting the McGurk Effect

Individual Differences in Susceptibility

Susceptibility ranges from 0% to 100% across neurologically typical adults. Some people perceive the illusion on almost every trial. Others, with normal hearing and vision, almost never do (Nath & Beauchamp, 2012).

Two factors help explain why people differ:

  • Lipreading Skill: Strand et al. (2014) found susceptibility correlated with lipreading ability and with detecting audiovisual incongruity, and was highly stable on retest.
  • Where People Look: Gurler et al. (2015) used eye-tracking and found frequent perceivers fixate the talker’s mouth more than infrequent perceivers do.

Both factors act at the perceptual front end, in how much visual speech a person picks up. The rest depends on how that evidence is weighed against the sound.

Susceptibility also changes across the lifespan. Adults are more susceptible than young children (McGurk & MacDonald, 1976). Visual influence keeps rising with age in English-learning children (Sekiyama & Burnham, 2008).

Older adults often show a stronger effect. Hearing loss plausibly increases their reliance on the visual channel. Consistent with this, estimated sensory noise rose with age in native Japanese speakers (Magnotti et al., 2024).

Brain Damage

The hemispheres of the brain cooperate to integrate speech information received via the optic and aural senses (Baynes, Funnell & Fowler, 1994).

Handedness plays a role. In right-handed individuals, words have privileged access to the left hemisphere and faces to the right hemisphere. These individuals are more likely to experience a McGurk effect.

The effect is still present after callosotomy, surgery that cuts the link between the hemispheres. It is, however, significantly slower.

Left-hemisphere lesions can strengthen the effect. Visual stimuli strongly influence speech perception in these individuals, producing a larger McGurk effect than in the average person (Schmid, Thielmann & Ziegler, 2009).

The effect would weaken, however, if the damage had impaired their visual speech perception.

Right-hemisphere damage has a different profile. A case study reported impaired perception of both auditory and visual speech prosody (Nicholson, Baum, Cuddy & Munhall, 2002). More broadly, such damage is reported to impair visual-only as well as audio-visual integration.

These individuals can still show a McGurk effect. Integration appears only when visual cues are used to boost performance while the audio signal is poor.

Their McGurk effect is therefore weaker than in a typical group.

Alzheimer’s Disease

One study used a crossmodal effect to investigate brain connectivity in Alzheimer’s disease (AD). It compared the McGurk effect in people with AD and matched control participants (Delbeuck, Collette & Van der Linden, 2007).

The results showed impaired crossmodal integration of speech in AD. Yet the processing of the visual and auditory speech signals separately was not disrupted. Integration itself had failed.

This fits the view that Alzheimer’s is a disconnection syndrome. Impaired communication between distant brain regions may degrade the binding of information across the senses.

Specific Language Impairment

A study compared the auditory-visual integration of 28 preschoolers with Specific Language Impairment (SLI) and 28 preschoolers without SLI (Norrix, Plante, Vance & Boliek, 2007).

Both groups performed equally well with congruent audio-visual speech. With incongruent speech, however, children with SLI showed a weaker McGurk effect than the other group.

Children with SLI may pay less attention to articulatory gestures and use less visual information in perceiving speech. Perceiving aural cues alone, however, poses no significant difficulty.

Aphasia

One study tested a person with mild aphasia, a language disorder usually caused by brain damage, on speech tokens in visual-only, auditory-only and audio-visual conditions (Youse, Cienkowski & Coelho, 2004).

The hypothesis was that performance would be best in the bimodal condition and that the McGurk effect would show integration of speech information.

The results did not support the hypotheses. They suggested that a perseverative response pattern, repeating the same response, limits the integration of audio-visual speech information. Use of bisensory speech stimuli may therefore be compromised in adults with aphasia.

These findings come from small samples and single cases. They show that brain pathology can disrupt the effect, but they cannot support diagnostic claims.

The McGurk Effect in Different Languages

Regardless of the language being used, listeners generally depend, to some degree, on visual information in speech perception. However, the intensity of the McGurk effect varies across languages.

Earlier studies found a stronger McGurk effect in Spanish, Italian, Turkish, English, Dutch and German listeners (Sekiyama, 1997; Bovo, Ciorba, Prosser & Martini, 2009; Erdener, 2015). Chinese and Japanese listeners showed a weaker one.

A large 2025 study challenges this picture. Native Mandarin and English speakers showed similar McGurk rates (Magnotti et al., 2025).

Explanations for Language Differences

The cultural practice of avoiding eye contact might account for the diminished effect among Japanese and Chinese listeners. Tonic and syllabic linguistic structures may also contribute.

Japanese children differ from English children in development. They show no advance in visual influence after age six (Sekiyama & Burnham, 2008; Hisanaga, Sekiyama, Igasaki & Murayama, 2009).

However, Japanese listeners detect the mismatch between sound and sight better than English listeners, perhaps because Japanese lacks consonant clusters (Sekiyama & Tohkura, 1991). Noise changed this pattern. The Japanese effect grew when noise was added to the audio (Sekiyama & Tohkura, 1991).

Notwithstanding the manifest differences, listeners of all languages are compelled to rely on optic stimuli when audio stimuli are unintelligible. When this occurs, variation across languages disappears, and the McGurk effect is applied equally.

Critical Evaluation of the McGurk Effect

The McGurk effect is a landmark demonstration that perception builds a single experience from several senses. Its value now lies in what its variability reveals.

Strengths of the McGurk Effect

Four features make the effect one of the most useful tools in perception research:

  • Theoretically Decisive: It shows clearly that speech perception is audiovisual. It retired the assumption that heard speech is decoded from sound alone.
  • Robust and Easily Produced: It replicates readily and survives gender-mismatched talkers (Green et al., 1991). Anyone can experience it in seconds.
  • Convergent Support: Behavioural tests, fMRI (Nath & Beauchamp, 2012), TMS (Beauchamp et al., 2010) and brain-wave recordings (Alsius et al., 2014) all point to the same STS-centred network.
  • Generative: It has driven productive research on individual differences, computational modelling and clinical assessment.

Together, these features explain why the effect became a standard tool. It makes an invisible process visible to any observer, and it links behaviour to a specific brain network.

Limitations of the McGurk Effect

Six limitations qualify how far the illusion can be generalised:

  • Artificial Stimuli: The illusion depends on deliberately mismatched cues that almost never occur in natural speech. Its ecological validity, or match to everyday perception, is debatable.
  • Not One Thing: Susceptibility ranges from 0% to 100% across people and varies across stimuli, tasks and languages, so a single McGurk rate misleads (Basu Mallick et al., 2015).
  • Inconsistent Scoring: Laboratories differ on which responses count and on open versus forced-choice formats, which undermines comparison (Tiippana, 2014).
  • Not Fully Automatic: Attention and expectation modulate the effect (Tiippana et al., 2004; Alsius et al., 2007; Windmann, 2004).
  • Uncertain Real-World Link: It is not settled that susceptibility indexes the integration behind the everyday benefit of seeing a face, so generalise cautiously.
  • Thin Clinical Evidence: Much neuropsychological work rests on small samples and single cases, which cannot support strong diagnostic claims.

Contemporary Research

Recent work asks how the effect’s striking variability should be measured and explained. The effect is not a fixed rate. It is the output of a noisy inference process that differs across people and stimuli.

Measuring it properly requires many stimuli, large samples and models that separate the two sources of variation. Four strands stand out: the size of individual differences, a model to explain them, how to define the effect, and whether group differences hold up.

Variability and Stability

Basu Mallick, Magnotti and Beauchamp (2015) ran the field’s most rigorous test of how the effect behaves across people and time.

  • Aim: To describe how McGurk perception varies across people, stimuli and time, and to test whether the mean fusion rate is a valid summary.
  • Method: In Experiment 1, 165 English-speaking adults viewed 12 McGurk stimuli, and 40 were retested a year later. Experiment 2 used 8 new stimuli and compared open-choice with forced-choice responses.
  • Results: Perception varied from 0% to 100% across people and from 17% to 58% across stimuli. About 77% of participants almost never or almost always perceived the effect, and individual susceptibility was highly stable (test-retest r = 0.91).
  • Conclusion: Individual differences are large but stable. The mean fusion rate is a poor, statistically invalid summary of such two-humped data.

Forced-choice responding also inflated the effect by roughly 18% relative to open choice.

The study sits near the top of the evidence hierarchy, with a large sample, an internal replication and a one-year retest.

Its scope is methodological, not mechanistic, and it tested English speakers only.

Modelling the Variability

The noisy encoding of disparity (NED) model tackles this variability. It describes each stimulus by the disparity between its audio and video channels.

It describes each person by how noisily they encode that disparity and by their threshold for perceiving the illusion (Magnotti & Beauchamp, 2015).

The model reproduced perception accurately. It also allowed groups to be compared without the confound of stimulus differences. Frequent mouth-lookers showed lower sensory noise and higher disparity thresholds (Gurler et al., 2015).

What Counts as a McGurk Response?

Researchers do not all mean the same thing by the effect (Tiippana, 2014). Three choices matter:

  • Which Percepts Count: A strict definition counts only fusions. A broader one counts any visually influenced response. Tiippana argues for defining the effect as a change in the auditory percept caused by visual speech.
  • Response Format: Free reporting and forced choice give different rates, because a fixed list nudges listeners toward the illusory option.
  • Number of Stimuli: Most early studies used one video, which confounds how compelling the stimulus is with how susceptible the person is.

Reappraising Group Differences

Large samples and stimulus-controlled models have revisited claims of group differences. Magnotti and colleagues (2025) compared Mandarin and English speakers directly.

  • Aim: To test whether native Mandarin and American English speakers differ in McGurk perception, using large samples and many stimuli.
  • Method: 307 native Mandarin and American English speakers viewed nine McGurk stimuli. The design separated variation across participants, stimuli and language group.
  • Results: McGurk frequencies were similar (48% versus 44%). Language group explained about 0.2% of the variance, against 0-100% variation across participants and 14-83% across stimuli.
  • Conclusion: Earlier reports of a large East Asian versus Western gap were probably overstated by small, single-stimulus designs.

Magnotti, Lado and Beauchamp (2024) applied the NED model to native Japanese speakers. They again found high variability across people and stimuli. Unvoiced “pa/ka” pairings evoked the illusion more strongly than voiced “ba/ga” pairings, a regularity that a single-token study would miss.

Within-group variability therefore dwarfs any between-group difference. Credible claims about who differs from whom need large, multi-stimulus, model-based designs.

Is the McGurk Effect an Illusion?

Yes, the McGurk effect is a perceptual illusion. The brain automatically combines what it sees on a speaker’s lips with what it hears. When the two conflict, most people hear a third sound that was in neither channel. It persists even when you know the trick.

References

Alsius, A., Möttönen, R., Sams, M. E., Soto-Faraco, S., & Tiippana, K. (2014). Effect of attentional load on audiovisual speech perception: Evidence from ERPs. Frontiers in Psychology, 5, 727. https://doi.org/10.3389/fpsyg.2014.00727

Alsius, A., Navarra, J., & Soto-Faraco, S. (2007). Attention to touch weakens audiovisual speech integration.  Experimental Brain Research,  183 (3), 399-404.

Basu Mallick, D., Magnotti, J. F., & Beauchamp, M. S. (2015). Variability and stability in the McGurk effect: Contributions of participants, stimuli, time, and response type. Psychonomic Bulletin & Review, 22(5), 1299-1307. https://doi.org/10.3758/s13423-015-0817-4

Baynes, K., Funnell, M. G., & Fowler, C. A. (1994). Hemispheric contributions to the integration of visual and auditory information in speech perception.  Perception & Psychophysics,  55 (6), 633-641.

Beauchamp, M. S., Nath, A. R., & Pasalar, S. (2010). fMRI-guided transcranial magnetic stimulation reveals that the superior temporal sulcus is a cortical locus of the McGurk effect. The Journal of Neuroscience, 30(7), 2414-2417. https://doi.org/10.1523/JNEUROSCI.4865-09.2010

Boersma, P. (2012). A constraint-based explanation of the McGurk effect.  Phonological Architecture: Empirical, Theoretical and Conceptual Issues, 299-312.

Bovo, R., Ciorba, A., Prosser, S., & Martini, A. (2009). The McGurk phenomenon in Italian listeners. Acta Otorhinolaryngologica Italica, 29(4), 203-208.

Brancazio, L., Best, C. T., & Fowler, C. A. (2006). Visual influences on perception of speech and nonspeech vocal-tract events.  Language and speech,  49 (1), 21-53.

Delbeuck, X., Collette, F., & Van der Linden, M. (2007). Is Alzheimer’s disease a disconnection syndrome?: Evidence from a crossmodal audio-visual illusory experiment.  Neuropsychologia,  45 (14), 3315-3323.

Erdener, D. (2015). The McGurk illusion in Turkish. Turkish Journal of Psychology. 30 (76): 19–31.

Green, K. P., Kuhl, P. K., Meltzoff, A. N., & Stevens, E. B. (1991). Integrating speech information across talkers, gender, and sensory modality: Female faces and male voices in the McGurk effect. Perception & Psychophysics, 50(6), 524-536. https://doi.org/10.3758/BF03207536

Gurler, D., Doyle, N., Walker, E., Magnotti, J., & Beauchamp, M. (2015). A link between individual differences in multisensory speech perception and eye movements. Attention, Perception, & Psychophysics, 77(4), 1333-1341. https://doi.org/10.3758/s13414-014-0821-1

Hisanaga, S., Sekiyama, K., Igasaki, T., & Murayama, N. (2009). Audiovisual speech perception in Japanese and English: inter-language differences examined by event-related potentials. In  AVSP  (pp. 38-42).

Körding, K. P., Beierholm, U., Ma, W. J., Quartz, S., Tenenbaum, J. B., & Shams, L. (2007). Causal inference in multisensory perception. PLoS ONE, 2(9), e943. https://doi.org/10.1371/journal.pone.0000943

Magnotti, J. F., Basu Mallick, D., Feng, G., Zhou, B., Zhou, W., & Beauchamp, M. S. (2025). The McGurk effect is similar in native Mandarin Chinese and American English speakers. Frontiers in Psychology, 16, 1531566. https://doi.org/10.3389/fpsyg.2025.1531566

Magnotti, J. F., & Beauchamp, M. S. (2015). The noisy encoding of disparity model of the McGurk effect. Psychonomic Bulletin & Review, 22(3), 701-709. https://doi.org/10.3758/s13423-014-0722-2

Magnotti, J. F., & Beauchamp, M. S. (2017). A causal inference model explains perception of the McGurk effect and other incongruent audiovisual speech. PLOS Computational Biology, 13(2), e1005229. https://doi.org/10.1371/journal.pcbi.1005229

Magnotti, J. F., Lado, A., & Beauchamp, M. S. (2024). The noisy encoding of disparity model predicts perception of the McGurk effect in native Japanese speakers. Frontiers in Neuroscience, 18, 1421713. https://doi.org/10.3389/fnins.2024.1421713

Massaro, D. W., & Cohen, M. M. (1983). Evaluation and integration of visual and auditory information in speech perception. Journal of Experimental Psychology: Human Perception and Performance, 9(5), 753-771. https://doi.org/10.1037/0096-1523.9.5.753

Massaro, D. W., & Stork, D. G. (1998). Speech recognition and sensory integration: a 240-year-old theorem helps explain how people and machines can integrate auditory and visual information to understand speech.  American Scientist,  86 (3), 236-244.

McGurk, H., & MacDonald, J. (1976). Hearing lips and seeing voices.  Nature,  264 (5588), 746-748.

Nath, A. R., & Beauchamp, M. S. (2012). A neural basis for interindividual differences in the McGurk effect, a multisensory speech illusion.  Neuroimage,  59 (1), 781-787.

Nicholson, K. G., Baum, S., Cuddy, L. L., & Munhall, K. G. (2002). A case of impaired auditory and visual speech prosody perception after right hemisphere damage.  Neurocase,  8 (4), 314-322.

Norrix, L. W., Plante, E., Vance, R., & Boliek, C. A. (2007). Auditory-visual integration for speech by children with and without specific language impairment. Journal of Speech, Language, and Hearing Research, 50 (6), 1639–1651.

Rosenblum, L. D. (2008). Speech perception as a multimodal phenomenon. Current Directions in Psychological Science, 17(6), 405-409. https://doi.org/10.1111/j.1467-8721.2008.00615.x

Schmid, G., Thielmann, A., & Ziegler, W. (2009). The influence of visual and auditory information on the perception of speech and non‐speech oral movements in patients with left hemisphere lesions.  Clinical linguistics & phonetics,  23 (3), 208-221.

Schwartz, J.-L., Berthommier, F., & Savariaux, C. (2004). Seeing to hear better: Evidence for early audio-visual interactions in speech identification. Cognition, 93(2), B69-B78. https://doi.org/10.1016/j.cognition.2004.01.006

Sekiyama, K. (1997). Cultural and linguistic factors in audiovisual speech processing: The McGurk effect in Chinese participants.  Perception & psychophysics,  59 (1), 73-80.

Sekiyama, K., & Burnham, D. (2008). Impact of language on development of auditory‐visual speech perception.  Developmental science,  11 (2), 306-320.

Sekiyama, K., & Tohkura, Y. I. (1991). McGurk effect in non‐English listeners: Few visual effects for Japanese participants hearing Japanese syllables of high auditory intelligibility.  The Journal of the Acoustical Society of America,  90 (4), 1797-1805.

Strand, J., Cooperman, A., Rowe, J., & Simenstad, A. (2014). Individual differences in susceptibility to the McGurk effect: Links with lipreading and detecting audiovisual incongruity. Journal of Speech, Language, and Hearing Research, 57(6), 2322-2331. https://doi.org/10.1044/2014_JSLHR-H-14-0059

Tiippana, K. (2014). What is the McGurk effect? Frontiers in Psychology, 5, 725. https://doi.org/10.3389/fpsyg.2014.00725

Tiippana, K., Andersen, T. S., & Sams, M. (2004). Visual attention modulates audiovisual speech perception.  European Journal of Cognitive Psychology,  16 (3), 457-472.

Vatakis, A., & Spence, C. (2007). Crossmodal binding: Evaluating the “unity assumption” using audiovisual speech stimuli. Perception & Psychophysics, 69(5), 744-756. https://doi.org/10.3758/BF03193776

Walker, S., Bruce, V., & O’Malley, C. (1995). Facial identity and facial speech processing: Familiar faces and voices in the McGurk effect.  Perception & Psychophysics,  57 (8), 1124-1133.

Welch, R. B., & Warren, D. H. (1980). Immediate perceptual response to intersensory discrepancy. Psychological Bulletin, 88(3), 638-667.

Windmann, S. (2004). Effects of sentence context and expectation on the McGurk illusion. Journal of Memory and Language, 50(2), 212-230. https://doi.org/10.1016/j.jml.2003.10.001

Youse, K. M., Cienkowski, K. M., & Coelho, C. A. (2004). Auditory-visual speech perception in an adult with aphasia.  Brain injury,  18 (8), 825-834.

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.

Ayesh Perera

Researcher

B.A, MTS, Harvard University

Ayesh Perera, a Harvard graduate, has worked as a researcher in psychology and neuroscience under Dr. Kevin Majeres at Harvard Medical School.