Bottom-up processing is perception driven directly by the senses.
Information travels from the raw stimulus, through the sensory receptors, to the brain, which builds a percept from the stimulus’s own properties (Gibson, 1966).
It is also called data-driven processing. It is the mirror image of top-down processing, where the brain uses existing knowledge to interpret what comes in (Gregory, 1970).
Key Takeaways
- Data-Driven: Bottom-up processing is perception driven by the external stimulus itself, interpreted from its own properties without relying on prior knowledge.
- Real-Time: It interprets sensory information as it arrives, in the moment (Gibson, 1966).
- No Prior Learning Needed: It works as sensory receptors take in new information, without drawing on past knowledge or experience.
- Stimulus First: The raw data of the stimulus itself drives perception, not stored expectations.
- The Sequence: New sensory information reaches the receptors, which signal the brain; the brain processes those signals and constructs a perception from them.
The bottom-up process involves information traveling “up” from the stimuli, via the senses, to the brain which then interprets it, relatively passively.
Bottom-up processing is also known as data-driven processing because information processing begins with environmental stimuli, and perceptions are built from sensory input.
Bottom-Up vs. Top-Down Processing
Bottom-up processing begins with the retrieval of sensory information from the external environment, building perceptions from the current input (Gibson, 1966).
Top-down processing is the interpretation of incoming information based on prior knowledge, experiences, and expectations (Gregory, 1970).
Bottom-Up
- Data-driven
- Relies on sensory information
- Takes place in real-time
Top-Down
- Schema driven
- Relies on knowledge and experiences
- Prior knowledge essential
Previous knowledge, experience, and expectations are the driving force in top-down perception (Gregory, 1974).
In bottom-up processing, no learning is required: perception is driven solely by the stimulus currently being experienced (Gibson, 1972).
Sensation vs. Perception
Bottom-up processing is the process of ‘sensation’ and top-down is the process of ‘perception’.
Sensation is the input of sensory information from the external environment, received by our sensory receptors — the bottom-up half of the process.
Perception is how our brains choose, organize, and interpret these sensations.
Perception is unique to each individual as we interpret these sensations based on our individual schemas that are constructed from previous knowledge, experiences, and expectations (Jandt, 2020).
How Bottom-Up Processing Works
Bottom-up processing starts with minute sensory details that are then used to construct larger ideas or perceptions about one’s external environment.
Processing runs in one direction, from the retina to the visual cortex. Each successive stage in the visual pathway carries out an ever more complex analysis of the input.
Bottom-up processing works like this:
- We start with an analysis of sensory inputs such as patterns of light.
- This information is relayed to the retina, where the process of transduction into electrical impulses begins.
- These impulses are passed into the brain where they trigger further responses along the visual pathways until they arrive at the visual cortex for final processing.
Feature Detection: The Neural Evidence
The theory has a physiological basis, too. Recording from single neurons in the cat and monkey visual cortex, Hubel and Wiesel found cells that fire only to specific, simple features of the input (Hubel & Wiesel, 1959).
Simple cells respond to a bar or edge at one orientation in one retinal location. Complex cells respond to the same edge anywhere in a larger area, often preferring one direction of movement.
That hierarchy is the neural embodiment of bottom-up processing.
Because complex cells are driven by groups of simple cells, the visual system builds perception from the bottom up. Receptors feed feature detectors, and feature detectors feed higher detectors, adding detail at each stage.
Bottom-up processing states that we begin to perceive new stimuli through the process of sensation, and the use of our schemas is not required. James J. Gibson (1966) argued that no learning was required to perceive new stimuli.
Gibson looked at perception as more of a ‘what you see is what you get’ kind of situation: perception does not require inference. We experience new stimuli through our sensations and then directly analyze their meaning.
Unlike Gregory’s theory (1970), Gibson believed that the environment holds all of the necessary tools to create accurate perceptions of incoming stimuli.
Gibson’s Theory of Direct Perception
James J. Gibson’s (1966) theory of direct perception, also called the ecological theory, is the purest statement of bottom-up processing.
Gibson argued that no learning is required to perceive new stimuli. The claim is a strong one. The environment holds all the information needed, so perception is direct.
The Optic Array
Gibson rejected the idea that perception begins from meaningless points of light hitting isolated receptors. He looked elsewhere for the starting point.
Instead, he started from the optic array: the structured pattern of light that reaches an observer and carries all the visual information available. The array is not a poor, snapshot-like image waiting to be decoded.
It is already richly structured. As an observer moves, it changes in lawful ways that specify the layout of surfaces. Nothing is missing from it.
This is what Gibson meant by ‘the information is in the light’ (Gibson, 1966).
The visual world does not need to be rebuilt inside the head, because it is already specified in the array. Perception, on this view, is closer to a map-reader’s active sampling than to a camera passively recording. That is the whole point.
Invariants and Texture Gradients
Certain higher-order properties of the optic array stay the same as the observer moves. These are Gibson’s invariants.
Their constancy across changing viewpoints is what makes them informative: an invariant gives unambiguous, directly available information about the enduring layout of the world.
Learning, on this account, is not building sensations into concepts. It is the progressive differentiation of invariants that were there all along.
A key example is the texture gradient. This is the rate at which a textured surface’s density changes from front to back.
Looking across a lawn, the grass nearest you is coarse and detailed, growing steadily finer into the distance. The same texture expands as you approach an object and contracts as it recedes.
No inference is required. The gradient specifies depth directly.
Optic Flow and Affordances
Gibson developed part of his theory while preparing aircraft-landing training films during the Second World War. Optic flow was its centrepiece.
As an observer moves forward, the whole visual field seems to flow outward from the point they are heading toward. That point, the focus of expansion, stays still.
The stationary focus is an invariant that specifies heading directly. The rate of flow specifies speed and, for a landing pilot, altitude.
A related cue is motion parallax. As we move, near objects sweep across the visual field faster than far ones, letting us judge distance directly.
Gibson also gave a bottom-up account of meaning. He called the action possibilities an object offers an affordance.
A surface can be stand-on-able, and an object graspable or edible. Gibson treated affordances as part of the information already in the array, not knowledge the perceiver adds afterward.
The developmental evidence for early, unlearned depth perception comes from the classic visual cliff experiment.
Infants and young animals refused to cross an apparent drop-off. This shows that some depth perception is available early and does not require extensive learning (Gibson & Walk, 1960).
Real-Life Applications
Here is how bottom-up and top-down processing compare in practice.
Stubbing Your Toe
Imagine stubbing your pinky toe on the corner of the bed. The pain receptors in your toe immediately register the injury and send pain signals straight to your brain. That immediate signal, processed with no help from memory, is bottom-up processing.
Afterwards, you remember how painful it was. You start taking care to avoid that corner of the bed. That is a top-down adjustment: it draws on memory, not on the sensory data in front of you.
Blind Food Taste Challenge
A blind taste test isolates one sense, taste, by removing sight and packaging cues. Participants are blindfolded and asked to judge food or drink using taste alone.
Lowengart (2013) used this method on wine. Participants were blindfolded and asked to pick their preferred wine from unlabeled samples, testing whether branding affects consumer choice. The labels stayed hidden.
Because brand and label expectations were removed, judgement rested on taste alone. Bottom-up processing was doing the work, not memory of a preferred brand.
If participants had instead been asked to name the brand, memory would enter the judgement too. That would mix in top-down processing.
Prosopagnosia
Prosopagnosia (phonetically pronounced praa-suh-pag-now-zhuh) is a visual form of agnosia where individuals cannot recognize faces or facial differences (Harris & Aguirre, 2007).
Prosopagnosia, often referred to as face blindness, is a rare condition where patients who are affected cannot recognize whether they have seen someone’s face before or not.
In cases such as these, top-down processing cannot be used to distinguish one face from the next. Individuals must rely on what they see at the moment when analyzing someone’s face.
This is because individuals with prosopagnosia can recognize different facial features but are not able to use their memory to put a name to a face. In essence, individuals with prosopagnosia cannot detect familiar faces because they cannot combine facial features into complete faces that they can then recognize in the future.
“Imagine that every person has a camera inside their head. Every time they meet somebody for the first time, they take a picture with their camera, develop the picture, and file it away for future use. …For me, I take a picture with my camera, but I never store it away” (Lewis, 2013).
Patients with prosopagnosia cannot mentally store the faces of people they know. Building perceptions through top-down processing is therefore impossible: they have no stored memory of the faces they have met before.
Each and every encounter forces individuals with prosopagnosia to place a name with a face with the new sensory information presented within each encounter.
Critical Evaluation
Bottom-up theories such as Gibson’s explain a great deal about everyday perception. But they also run into real problems.
Strengths
Four real strengths stand out, and they help explain why the theory has lasted, even though it also has real limits.
- Explains Everyday Perception: It explains how people perceive accurately and rapidly in good, ordinary viewing conditions, the everyday case illusion-focused theories tend to under-explain.
- Grounded in Physiology: The theory rests on real, measurable mechanisms, transduction and cortical feature detection (Hubel & Wiesel, 1959), rather than on hypothesised inference alone.
- Ahead of Its Time: Gibson insisted perception is normally active and richer in information than lab studies assumed; later research on the visual control of movement, in real environments, bore this out.
- Influential Affordance Concept: The idea that objects directly signal their own uses to a perceiver still shapes ecological psychology and interface design today.
Limitations
Five limitations are documented; three are explained in full below.
- Can’t Explain Illusions: A strictly bottom-up account cannot explain visual illusions, ambiguous figures, or the powerful effects of context and expectation (Gregory, 1970).
- Underestimated the Computation: Gibson underestimated how hard it actually is to extract invariants from the optic array (Marr, 1982).
- Explains “Seeing,” Not “Seeing As”: A pure pick-up account cannot supply the stored, cultural knowledge needed to recognise what an object means (Fodor & Pylyshyn, 1981).
- Methodological Asymmetry: Gibson built his case on optimal, real-world conditions, while constructivists built theirs on brief, impoverished lab displays.
- A False Dichotomy: Mature perception uses both routes together, so bottom-up is really one direction of flow within an interacting system, not a stand-alone theory (Neisser, 1976).
Can’t Explain Illusions
The decisive weakness of a strictly bottom-up theory is its failure to explain illusions and ambiguous figures (Gregory, 1970). That is the sharpest attack on the theory.
Visual illusions such as the Müller-Lyer and Ponzo figures show perception changing even though the sensory input never does. The same is true of the Necker cube, an ambiguous figure that flips between two readings. Nothing in the stimulus changes at all.
Gibson dismissed such effects as artificial ‘laboratory tricks’. He thought they told us little about real perception.
But some illusions resist that dismissal. The hollow-mask illusion and the distorted room both produce effects that feel just like ordinary perception, so they cannot simply be waved away (Gregory, 1970).
This is why the bottom-up account, on its own, cannot be the whole story of perception.
Underestimated the Computation
Gibson underestimated the computational difficulty of his own proposal. This is Marr’s (1982) core objection.
Detecting invariants in the optic array is precisely an information-processing problem, and a hard one. Saying that the information is ‘in the light’ does not explain how the nervous system actually extracts it.
Marr took the problem seriously. His own computational theory treated vision as a genuinely hard problem to solve.
He argued that vision proceeds through a series of increasingly rich representations. It starts with a primal sketch, making the edges and blobs in an image explicit. Object knowledge comes later.
Gibson’s resonance metaphor, an attuned system ‘tuning in’ to invariants like a radio, is evocative. But it is not a mechanism. It does not say how the tuning actually happens.
Methodological Asymmetry
There is a deeper asymmetry behind the whole debate, one that explains why the two sides disagree so completely. Gibson built his case from one kind of evidence.
He used optimal, real-world conditions: a moving observer, rich and unambiguous stimulation. The constructivists built their case from the opposite: brief, static, impoverished laboratory displays.
Each theory looks strongest on its own home ground. Lab demonstrations of top-down effects have been criticised for low ecological validity as a result of this asymmetry.
So which processing route wins depends partly on the conditions being tested. The richer and more natural the viewing conditions, the more bottom-up processing dominates. Context decides the balance. The eventual verdict was that both processes matter, weighted by how much the situation demands.
Contemporary Research
Modern perceptual science has mostly dissolved the either/or framing. The dominant view now treats the brain as a prediction machine.
Higher levels generate expectations about the causes of sensory input, a top-down flow. Lower levels compute the mismatch between that prediction and the actual data. Most of the signal is already expected. Only the surprising, unexplained part is passed back up to revise the model.
This is the ‘Bayesian brain’, or predictive-processing, view. Bottom-up sensory evidence and top-down prior expectation are weighted against each other by how reliable each one is.
Clear input lets bottom-up processing dominate. Noisy or ambiguous input hands more weight to top-down priors instead.
A second line of contemporary evidence comes from neuroscience: two visual streams. The streams work in parallel. A ventral ‘vision-for-perception’ pathway and a dorsal ‘vision-for-action’ pathway operate largely separately.
This vindicates part of Gibson’s insistence on a fast, memory-light, action-guiding form of vision. It does not mean the slower, knowledge-rich recognition system is unnecessary, though.
Both streams matter. The current research studies the balance between data and expectation, rather than trying to crown one winner.
References
Fodor, J. A., & Pylyshyn, Z. W. (1981). How direct is visual perception? Some reflections on Gibson’s “ecological approach”. Cognition, 9(2), 139-196.
Gibson, E. J., & Walk, R. D. (1960). The “visual cliff”. Scientific American, 202(4), 64-71.
Gibson, J. J. (1966). The Senses Considered as Perceptual Systems. Boston: Houghton Mifflin.
Gibson, J. J. (1972). A Theory of Direct Visual Perception. In J. Royce, W. Rozenboom (Eds.). The Psychology of Knowing. New York: Gordon & Breach.
Gregory, R. (1970). The Intelligent Eye. London: Weidenfeld and Nicolson.
Gregory, R. (1974). Concepts and Mechanisms of Perception. London: Duckworth.
Harris, A. & Aguirre, G. (2007). Prosopagnosia. Current Biology, 17 (1), R7-R8.
Hubel, D. H., & Wiesel, T. N. (1959). Receptive fields of single neurones in the cat’s striate cortex. Journal of Physiology, 148(3), 574-591.
Jandt, F. E. (2020). In An Introduction to Intercultural Communication: Identities in a Global Community (10th ed., pp. 68-101). California State University, San Bernardino, California: SAGE Publications.
Lewis, J. G. (2013, September 19). Prosopagnosia: Why Some are Blind to Faces. Retrieved January 10, 2021, from https://www.nature.com/scitable/blog/mind-read/blind_to_faces_the_neuroscience/
Lowengart, O. (2013). The effect of branding on consumer choice through blind and non-blind taste tests. Innovative Marketing, 8(4), 7-18.
Marr, D. (1982). Vision: A computational investigation into the human representation and processing of visual information. W. H. Freeman.
McClelland, J. L., & Rumelhart, D. E. (1981). An interactive activation model of context effects in letter perception: Part 1. An account of basic findings. Psychological Review, 88(5), 375-407.
Neisser, U. (1976). Cognition and reality: Principles and implications of cognitive psychology. W. H. Freeman.
Reicher, G. M. (1969). Perceptual recognition as a function of meaningfulness of stimulus material. Journal of Experimental Psychology, 81(2), 275-280.
BSc (Hons) Psychology, MRes, PhD, University of Manchester
Chartered Psychologist (CPsychol)
Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.
BSc (Hons) Psychology, MSc Psychology of Education
Associate Editor for Simply Psychology
Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.
Victoria Rousay is a current student in the Master of Liberal Arts, Anthropology, and Archeology degree program at Harvard University.