Echoic memory is a type of sensory memory that registers and temporarily holds auditory information (sounds) until it is processed and comprehended (Carlson, 2010). This sensory store can retain a great amount of auditory information for a brief period of 3 to 4 seconds (Clark, 1987).
Following the initial registration, the sound resonates and is briefly replayed in the mind (Radvansky, 2005). This trace is pre-categorical. It represents the sound itself, before it has been identified as a specific word or note.
Ulric Neisser, a German-American psychologist, introduced the term ‘echoic memory’ in 1967, by analogy with the ‘iconic’ store he had already named for vision.
Neisser’s contribution was naming and framing the idea, not running the original experiments himself.
The real measurement problem had just been solved for vision. George Sperling’s partial-report experiments showed that a flashed visual array holds far more information than a viewer can report before it fades (see iconic memory). The question that followed: could the same logic apply to hearing?
Subsequently, more advanced neuropsychological techniques were utilized to estimate the duration, location, and capacity associated with the echoic memory store.
Examples
Echoic memory matters because sound only exists as it unfolds. A brief auditory buffer is what lets the brain make sense of hearing at all. Three everyday examples show it at work:
- Listening to a song: When we listen to music, our brains briefly link each note to the one that follows, so a sequence of notes is heard as a tune.
- Conversing with another person: Echoic memory retains each syllable long enough for the brain to link it to the one before it, turning a stream of sound into words.
- Repeated speech: When something someone says is unclear, we ask them to repeat it. If the repetition matches the echoic trace of the original, we recognise it as familiar.
Take-home Messages
- Definition: Ulric Neisser introduced the term ‘echoic memory’ to signify a type of sensory memory that registers and temporarily holds auditory information (sounds) until it is processed and comprehended.
- Origins: The initial search for echoic memory emulated Sperling’s experiments on iconic memory (the visual sensory register), but subsequent research has used more advanced neuropsychological techniques.
- Brain regions: The brain regions involved in echoic memory include Broca’s area, the dorsal premotor cortex, the posterior parietal cortex, the superior temporal gyrus, and the inferior temporal gyrus.
- Age-related change: Research suggests that echoic memory grows with age until adulthood, and then declines with old age.
Research
Much of what we know about echoic memory comes from a single landmark experiment. Researchers then traced how long the resulting trace lasts, and how it feeds into later memory systems.
The Darwin, Turvey and Crowder (1972) Study
George Sperling’s partial-report method for vision inspired a direct auditory adaptation.
- Aim: To test whether hearing has a brief, high-capacity sensory store like Sperling’s visual icon, and to estimate how long it lasts.
- Method: Using stereo headphones, Darwin, Turvey and Crowder (1972) played three spoken lists seeming to come from the left ear, right ear, and mid-head. Listeners either recalled the whole array (whole report) or, after a delayed cue, just one location (partial report).
- Results: Partial report beat whole report: listeners recalled a higher proportion of the cued location than they could manage for the whole array. This advantage shrank as the cue was delayed, and had disappeared by about four seconds.
- Conclusion: A brief, high-capacity auditory store holds more sound than a listener can report, and persists up to about four seconds (Darwin, Turvey & Crowder, 1972).
The auditory advantage was smaller than Sperling’s visual one, and depended heavily on the spatial cue. This raises a question later work tried to resolve: how much of the effect is a pure acoustic trace, and how much is remembered spatial location?
How Long Does the Echo Last?
No single number captures how long the echo lasts. Some early estimates suggested it might last a second or less, while Johnson and Eriksen found it could take up to 10 seconds (Eriksen & Johnson, 1964).
Nelson Cowan, a psychologist at the University of Missouri, proposed a way to reconcile these numbers: echoic memory is not one store but two (Cowan, 1984).
A very brief, high-fidelity phase lasts around 200 to 300 milliseconds. It preserves fine acoustic detail. A longer, seconds-long phase then holds a coarser, more categorised version, available for comparison and recognition, especially for coarse features like pitch.
So echoic memory is graded, not fixed: sharpest for a few hundred milliseconds, still usefully readable for a few seconds, faintly detectable for longer still.
The Modality Effect and the Suffix Effect
Two further findings backed the case for a distinct auditory store. When people recall a list immediately, the last item or two are usually remembered best, an effect called recency. This boost is larger for lists heard aloud than for lists read silently, called the modality effect.
The effect is sharper still. A redundant spoken “suffix”, which listeners are told to ignore, wipes out the recency advantage entirely.
A non-speech suffix, like a buzzer, does far less damage. This suggests the suffix overwrites the echoic trace of the true final items before they can be reported (Crowder & Morton, 1969).
This pattern became known as precategorical acoustic storage (PAS). It is a brief auditory store, before words are recognised, whose contents a later sound can mask.
Morton, Crowder and Prussin (1971) showed the damage tracks the suffix’s acoustic similarity to the list, not its meaning. A suffix in the same voice and location hurt recall most, while a different voice, location, or tone did far less harm.
Echoic Memory and the Working Memory Model
In 1974, Alan Baddeley and Graham Hitch proposed a working memory model built around a phonological loop for auditory-verbal material (Baddeley & Hitch, 1974; Baddeley, Eysenck & Anderson, 2009).
The loop has two parts. A phonological store briefly holds heard speech in a sound-based code, and an articulatory rehearsal process refreshes the trace using inner speech. This early model, however, could not fully explain how raw auditory input becomes the loop’s contents.
The key distinction is this: the phonological store is attended, capacity-limited and rehearsable, while echoic memory itself is pre-attentive, high-capacity and cannot be rehearsed.
Echoic memory sits upstream of the loop. It is where sound first lands, before any of it is selected for further processing.
Baddeley later added an episodic buffer (Baddeley, 2000) to the model, a further stage for integrating information drawn from multiple sources. That stage sits downstream of both echoic memory and the phonological loop.
Methods for Testing
Whole Reporting and Partial Reporting
George Sperling’s research on iconic memory in the 1960s inspired other researchers to test the same idea in hearing (Darwin, Turvey & Crowder, 1972). In Sperling’s experiments, participants had to repeat the letters they saw.
Echoic memory studies followed the same logic. Listeners repeated sequences of syllables, words or tones that they heard, and partial reporting again beat whole reporting.
The gap between the sounds and the recall cue mattered too. The longer that interval, the less listeners could recall.
ABRM (Auditory Backward Recognition Masking)
ABRM presents a brief target sound, followed after a short interval by a second sound, the mask (Bjork & Bjork, 1996). Lengthening that interval changes how long the echoic trace has to be read out before the mask arrives.
Performance improves as the interval grows to around 250 milliseconds. The mask does not seem to erase the initial registration of the target. Instead, it interferes with the further processing needed to identify it.
Mismatch Negativity
Mismatch negativity (MMN) is the field’s most objective tool. It needs no attention and no verbal report (Näätänen & Escera, 2000).
Näätänen, Gaillard and Mäntysalo (1978) discovered it. They used the oddball paradigm: a stream of identical “standard” tones is interrupted, rarely and unpredictably, by a physically different “deviant” tone.
Subtracting the brain’s response to the standard from its response to the deviant reveals a negative deflection. It peaks around 150 to 200 milliseconds after the deviant.
To register as “deviant” at all, the incoming sound must be compared against a memory representation of the standard (Sabri, Kareken, Dzemidzic, Lowe & Melara, 2004). That comparison is what makes MMN a genuine index of the echoic store, not just a marker of change detection.
Neural Basis of Echoic Memory
Echoic processing is not confined to one region. It draws on a distributed cortical network. Lesion, fMRI and EEG studies all converge on its main components.
Much of the network sits in the prefrontal cortex. This region governs executive control and the direction of attention (Alain, Woods & Knight, 1998).
Alain and colleagues studied patients with focal prefrontal lesions. Damage there reduced the brain’s automatic detection of auditory change.
Beyond the frontal lobe, several regions contribute. Broca’s area, in the ventrolateral prefrontal cortex, handles articulatory processing and verbal rehearsal. The dorsal premotor cortex organises rhythm. The posterior parietal cortex localises sound in space.
The temporal lobe matters too, including the superior and inferior temporal gyri. Schönwiesner and colleagues (2007) studied Heschl’s gyrus, the posterior superior temporal gyrus, and the mid-ventrolateral prefrontal cortex together. Each plays a distinct role in detecting acoustic change, rather than acting as one undifferentiated detector.
Echoic Memory and Age
Echoic memory is not fixed. Using mismatch negativity, which needs no verbal report, Glass, Sachse and von Suchodoletz (2008) tracked auditory sensory memory in young children. They found that the trace’s duration lengthens dramatically between ages 2 and 6, rising from roughly 500ms to 5,000ms.
The gain continues into adulthood. Then, in old age, the trace declines. Because mismatch negativity is recorded automatically, this developmental curve reflects real maturation of the sensory-memory system rather than a change in strategy or motivation.
Critical Evaluation
Echoic memory is a well-established idea, but it carries several real, unresolved difficulties. Most stem from its position on the border between perception and memory.
- Converging methods: Behavioural partial report, the modality and suffix effects, backward masking, EEG and fMRI all point to the same brief, high-capacity, pre-categorical auditory store.
- An objective, attention-free measure: Mismatch negativity indexes the store without requiring report or attention, making it usable across the lifespan and in clinical groups.
How Long It Lasts, and Whether It Is Really ‘Memory’
The most visible weakness is the wide disagreement over how long the echo lasts, from under a second to about ten. Some of this reflects genuinely different phases of storage. Some of it is simply methodological, since different tasks tap different things.
Either way, echoic memory has no single clean duration. Any figure quoted should be tied to the task that produced it.
A deeper worry is whether echoic memory is a genuine memory store at all, or simply the persistence of auditory processing. Perhaps it is just hearing, lingering. Its pre-categorical, involuntary character certainly makes it look that way.
That mismatch negativity compares incoming sound against a stored model, even for omitted sounds, is the strongest evidence it is memory.
The line is genuinely blurry. Where persistence ends and memory begins remains partly a matter of definition.
Finally, the classic paradigms are far removed from everyday listening. Headphone arrays of meaningless letters and digits, and isolated tones in an oddball sequence, are nothing like listening to speech and music in a noisy, meaningful world.
The neat laboratory estimates may not transfer cleanly to real-world hearing. Most of the mechanistic evidence also comes from small samples.
Contemporary Research
Modern work has shifted from inferring the echoic store behaviourally to measuring it directly in the brain. The strongest recent evidence is a meta-analysis, not a single new study.
- Aim: To determine, across the published literature, how reliably mismatch negativity is reduced in schizophrenia.
- Method: A meta-analysis pooling effect sizes from many studies comparing people with schizophrenia, and those at clinical high risk, against healthy controls.
- Results: Mismatch negativity amplitude was substantially reduced in chronic schizophrenia, with a large pooled effect size that grew larger as the illness progressed.
- Conclusion: Reduced mismatch negativity is a robust, replicable marker of impaired auditory sensory memory in schizophrenia, and a candidate index of disease progression (Erickson, Ruffle & Gold, 2016).
As a meta-analysis of many studies, this finding sits near the top of the evidence hierarchy. It echoes Strous and colleagues’ (1995) original behavioural report that people with schizophrenia show degraded echoic memory.
Real-World Applications
Because mismatch negativity can be recorded directly, most practical uses of echoic memory run through it rather than through the classical report-based tasks.
This matters most for people who cannot give a reliable verbal or button-press response at all.
- Newborn screening: Guttorm and colleagues (2005) found that brain responses to speech sounds recorded at birth predicted children’s later language development, and were weaker in infants with a family history of dyslexia.
- Disorders of consciousness: Reviewing the coma literature, Morlet and Fischer (2014) found an MMN-like response in an unresponsive patient to be among the more reliable predictors of eventual awakening.
- Cochlear implants: Kelly, Purdy and Thorne (2005) showed that experienced adult cochlear-implant users’ electrophysiological discrimination and speech-perception scores tracked how faithfully their implant preserved fine acoustic detail.
Across all three uses, the appeal is the same. MMN needs no cooperation beyond passive listening, so it reaches patients that a verbal or attention-demanding test would have to exclude.