Ecological fallacy refers to a methodological error in which characteristics of a population as a whole are attributed to groups within that population without any real connection between them being demonstrated.
Key Points
- Definition: a mistaken inference about individuals drawn from group-level data.
- Consequences: it can produce false conclusions about social phenomena and the people within them.
- Origin: named and first demonstrated by sociologist William S. Robinson (1950).
- Avoidance: careful study design, statistics that show variation, and awareness of one’s own assumptions all reduce the risk.
Definition
The ecological fallacy is a mistaken conclusion drawn about individuals based on findings from groups to which they belong.
For example, imagine a university administrator finds a strong, positive correlation between studying engineering and scoring well in mathematics. It would be an ecological fallacy to assume this correlation holds for any one particular student.
The assumption ignores real variation within each group. Some engineering students score lower than some liberal arts students, even though engineering students average a higher mathematics grade as a group.
The first person to describe ecological fallacies was the sociologist H. C. Selvin (1958). Selvin defined three different types of ecological fallacy:
-
The fallacy of composition: This occurs when it is assumed that what is true for the group must also be true for the individuals within that group.
-
The fallacy of division: This occurs when it is assumed that what is true for the individual must also be true for the groups to which they belong.
-
The fallacy of misplaced concreteness: This occurs when data from one level of analysis (for example, groups) is treated as if it were data from another level of analysis (for example, individuals).
Durkheim’s own data illustrate the fallacy of composition. If a country’s suicide rate rises alongside its proportion of Protestants, concluding that Protestants themselves are more suicide-prone than Catholics would be a fallacy of composition. The pattern more likely reflects Protestant communities’ social integration (Selvin, 1958).
The fallacy of division runs the other way. Meeting one highly numerate engineering student and inferring that engineering students in general must be exceptionally numerate is a fallacy of division. A single case says nothing reliable about the group it came from (Selvin, 1958).
Both errors are easy to make.
Since Selvin first described ecological fallacies, they have been widely discussed in the literature on statistics and research methods, with implications for fields as varied as epidemiology, criminology, and economics.
Examples
Robinson (1950): Literacy and Immigration
The ecological fallacy gets its name from a classic study. Sociologist William S. Robinson (1950) compared two kinds of correlation calculated from the same 1930 US census data.
Aim: Robinson wanted to test whether an ecological correlation, calculated across groups such as US states, could substitute for an individual correlation calculated across single people. Earlier researchers routinely used state-level data to draw conclusions about the individuals living in those states without checking whether the two levels actually agreed.
Method: Robinson used real data. Using the 1930 census, he computed both individual-level and state-level correlations between two pairs of variables: being foreign-born and being illiterate, and being Black and being illiterate.
Findings: The two levels pointed in different directions. At the individual level, foreign-born residents were slightly more likely to be illiterate (r ≈ +0.12). At the state level, the same relationship reversed to strongly negative (r ≈ –0.53): states with more foreign-born residents had lower illiteracy overall.
Conclusion: Robinson concluded the two levels measure different things. The same reversal turned up for race: a modest +0.20 individual correlation became a larger +0.77 at the state level (Robinson, 2009). Robinson himself called such inferences “cannot be justified” (Robinson, 2009).
His finding matters beyond census statistics. It shows that a real, measured relationship at the group level can point in the opposite direction from the same relationship measured on individuals.
Crime
A study of crime rates may find, for example, that neighborhoods with a higher proportion of African-American residents had higher crime rates.
However, when the data is analyzed at the individual level, it may show that African-American individuals were no more likely to commit crimes than white individuals.
The pattern is not about individuals at all. It reflects neighbourhood-level features: concentrated poverty, policing intensity, and housing instability.
Sampson, Raudenbush and Earls (1997) studied this directly in a large multilevel study of Chicago neighbourhoods. A neighbourhood’s collective efficacy, not the individual traits of its residents, best predicted its rate of violent crime.
This is not a harmless statistical slip. Correlations of this kind have historically been used to justify discriminatory policing and housing practices.
Breast Cancer and Fat Consumption
The link between breast cancer and dietary fat shows the same error in epidemiology. Several early studies found that countries with higher average fat intake also had higher breast-cancer rates.
Aim: Holmes et al. (1999) tested whether a woman’s own fat intake predicted her own breast-cancer risk, using individual-level data rather than national averages.
Method: The method was straightforward. The researchers pooled data from several large prospective cohort studies. They tracked individual women’s dietary fat intake over time and recorded who later developed breast cancer.
Findings: No meaningful link emerged. Women with high-fat diets were not meaningfully more likely to develop breast cancer than women with low-fat diets, once other risk factors were accounted for (Holmes et al., 1999).
Conclusion: The population-level link was real. It was not, however, a reliable guide to any individual woman’s risk. Countries with higher average fat consumption may well have higher average breast-cancer rates. It does not follow that one woman’s own fat intake drives her own risk (Holmes et al., 1999).
Real-World Applications Across Domains
The fallacy is not confined to census statistics. It shapes research and policy across several fields.
Political Science and Voting Behaviour
National-level data reliably show incumbent governments losing votes when the economy performs badly. Yet individual-level surveys often find no matching link between a voter’s own finances and their vote choice.
Kramer (1983) argued this gap is mostly a statistical artefact. It is not proof that voters ignore their own circumstances. A person’s economic change mixes two things: a government-caused part, and a purely personal part such as an illness or a house move.
Averaging across many voters cancels out the personal part. The government-caused part survives, producing a strong national pattern from a genuinely weak individual one.
A related problem arises when researchers try to reconstruct how individuals voted from precinct-level results alone, since ballots are secret. King (1997) developed an influential statistical method for estimating such individual-level bounds from aggregate returns.
Geography and the Modifiable Areal Unit Problem
Census tracts, postcodes and electoral wards are used constantly in geography and spatial epidemiology. Their boundaries are often administrative accidents rather than anything meaningful about what is being studied.
Openshaw (1984) formalised this as the modifiable areal unit problem. The finding is stark. Redrawing zone boundaries, or changing their number and size, can shift or even reverse a correlation computed across those zones, without the underlying individual data changing at all.
This is a close relative of the ecological fallacy. A single finding can carry both risks at once. It says nothing directly about individuals, and a different, equally defensible choice of boundaries might have produced a different aggregate result in the first place.
The lesson for any study built on postcodes, counties, or wards is the same. Check whether a result survives when the boundaries are redrawn.
Why They Happen
Ecological vs Individual Correlation
The term “ecological fallacy” comes from a distinction between two kinds of correlation. An individual correlation describes single, indivisible people.
Take a simple case. The correlation between eye colour and illiteracy for persons in the United States is an individual correlation. In an individual correlation, the variables are properties of people, such as height, income or race, not statistical constants like rates or means (Robinson, 2009).
An ecological correlation, by contrast, describes a group. The correlation between the percentage of a state’s population that is Black and the percentage that is illiterate, computed across US states, is an ecological correlation. The thing described is a state, not a person.
The ecological fallacy occurs when conclusions that can be drawn from ecological correlations are mistaken for conclusions about individual correlations (Robinson, 2009).
Three related reasons explain why the fallacy happens:
- Group Averaging: group data smooth over real differences between the individuals inside the group.
- Assumption of Homogeneity: people wrongly assume every group member shares the group’s average traits.
- Stereotyping: oversimplified beliefs about a group get applied to each person in it.
Group Averaging
Data from groups can mislead when applied to individuals, because group data average away the differences within the group (Sedgwick, 2015). A group average is a single number standing in for a whole distribution of different people.
The effect is not exotic. It happens to ordinary averages all the time.
Take the student math example again. A minority of engineering students may excel at maths, and a minority of liberal-arts students may struggle badly, with most students in between regardless of major.
The mean ends up higher for engineers as a group. Yet the median score for liberal-arts students could be just as high, or higher. A handful of very high or very low scorers can pull a mean away from where most of the group’s members actually sit.
The Assumption of Homogeneity
This bias is closely related.
A second reason is that people assume group members share the group’s characteristics. This is called the assumption of homogeneity, and it wrongly implies that every member of a group is alike.
The error is easy to make. A group average feels as though it must belong to each member individually, when it is really just a summary of many different scores.
Suppose someone meets a student from the class with the highest average writing scores in their school district. Assuming she is a literary genius is a fallacy.
Coming from the top-scoring class does not make her the top scorer in it. She could be the lowest scorer in a class that otherwise consists of exceptional writers (Sedgwick, 2015).
Stereotyping
Stereotyping works through the same logical error.
Ecological fallacies can also stem from stereotypes: oversimplified, often inaccurate beliefs about a group’s characteristics. Applying a group-level stereotype to any one member is exactly the same move as applying a group-level statistic to them. Both skip over the real variation within the group.
Someone might assume all engineers are intelligent but unsociable, based on a “nerdy” stereotype. Even if engineers as a group were less socially confident than average, individual engineers can be extremely gregarious (Sedgwick, 2015).
Avoiding Ecological Fallacies
Ecological fallacies can arise in any research that uses group data. Researchers who study social phenomena at the group level are especially exposed to the risk (Piantadosi, Byar, & Green, 1988).
Group data are sometimes the only data available. That is simply a fact of the field. Being aware of the risk and taking active steps to reduce it produces more accurate, honestly qualified conclusions. Piantadosi, Byar and Green (1988) recommend the following steps:
- Examine Individual Data: check group patterns against individual-level data wherever it exists.
- Use Variation Statistics: measure spread, not just averages.
- Question Your Assumptions: check what a test or measure is really capturing.
- Watch Your Own Bias: examine data without assuming what it “should” show.
Examine Individual Data
The first step is the simplest.
Even a modest individual-level sample can reveal whether a group-level pattern actually and reliably holds within groups, not just between them. Looking at data from individuals, not only groups, can uncover patterns a group-level summary hides.
Holmes et al.’s (1999) breast-cancer study, described above, is a direct example. The point generalises well beyond this one article’s own examples. Its own individual-level data revealed that a real population-level correlation did not hold for any individual woman.
That is not the end of the story.
This step is not always possible. Individual-level data are often expensive, slow, or ethically difficult to collect, and sometimes they simply do not exist for a given question. Checking both levels, whenever both are available, remains the single most direct way to catch a fallacy before it reaches print.
Use Variation Statistics
Tools such as the standard deviation and chi-squared tests describe the range and reliability of data among individuals, rather than collapsing it into one group-level number. A mean on its own hides how spread out a group really is.
The reason is simple.
Two groups can share an identical average while looking nothing alike underneath it. Reporting the spread alongside the average, not the average alone, is what lets a reader judge how much the group figure actually tells them about any one member.
Question Your Assumptions
Stereotypes about a culture’s abilities can leave a test-score gap unquestioned, an instance of confirmation bias. The disparity may instead reflect a test’s reliance on culturally specific knowledge rather than the ability it claims to measure.
A question asking test-takers to build a verbal metaphor using moccasins, for example, favours cultures where moccasins are familiar. It ends up measuring cultural familiarity, not abstract reasoning.
Researchers can become more aware of exactly which measures apply fairly across the different groups being compared. Assuming that a single test means the same thing for everyone who takes it is itself a risky assumption.
Treating the resulting score gap as though it reflected ability alone repeats the very fallacy this article warns against.
Watch Your Own Bias
Researchers can reduce the risk in part by staying alert to their own preconceptions when interpreting group-level results. Striving to examine data without assuming what it “should” show helps, even though it does not eliminate bias entirely.
Bias is easiest to catch in someone else’s reasoning. It is hardest to catch in one’s own, which is why good intentions alone are not a reliable safeguard.
Building a data-analysis team from people with varying backgrounds is a further, structural safeguard. Different training histories make different unwarranted assumptions visible, in a way no single researcher can manage alone.
This is not a call for paralysis. A team can still move forward on an aggregate finding while flagging, in its own write-up, that the group-level result may not travel to every individual case.
Critical Evaluation
Robinson’s demonstration is exceptionally clean. It identifies an exact, quantified mismatch between two named correlations in real data, not just a hypothetical risk. The debate over which level of analysis to trust echoes psychology’s own reductionism and holism debate.
Simpson’s Paradox and the Individualistic Fallacy
The ecological fallacy is often confused with Simpson’s paradox (Simpson, 1951). Both involve an association that changes, sometimes reversing sign, depending on how the data are grouped. The two problems are analytically distinct.
Simpson’s paradox concerns data pooled across a confounding subgroup within a single level of analysis. The ecological fallacy is about mistaking a group-level relationship for an individual-level one.
A second risk runs the other way. Diez-Roux (1998) named this the individualistic, or atomistic, fallacy: assuming individual-level data alone tells you everything relevant, while ignoring real group-level effects.
Subramanian, Jones, Kaddour and Krieger (2009) tested this directly. They reanalysed Robinson’s own 1930 census data using modern multilevel models.
A large, independent state-level effect remained, even after individual race and nativity were taken into account. Jim Crow-era school segregation law drove much of it.
Robinson’s own verdict goes further than his evidence supports. Discarding the group level entirely just swaps one error for another.
Contemporary Research
Recent work asks a sharper version of Robinson’s question. Does a correlation found across many people describe the processes happening inside any one of them? This extends psychology’s own nomothetic versus idiographic debate.
Molenaar’s Ergodicity Challenge
Molenaar (2004) argued that group-averaged data only justify individual-level conclusions when the underlying process is ergodic. An ergodic process is consistent across people and stable over time.
Not every psychological process meets that condition. Some clearly do not. On that basis, Molenaar called for idiographic, single-person methods to be restored as a first-class part of scientific psychology.
Fisher et al.’s Empirical Test
Fisher, Medaglia and Jeronimus (2018) supplied the large-scale test Molenaar’s argument had lacked. They analysed six datasets, each following dozens of participants across many repeated occasions.
They compared correlations calculated within each person against correlations calculated between people. Average correlations broadly agreed across the two levels.
But individual-level correlations varied two to four times more widely than the group-level figure would suggest. A person’s own correlation could sit far from, or even oppose, their group’s average. That is a large gap.
The researchers framed this as a research-ethics concern, not just a statistical one. They argued for wider use of idiographic designs and open-science practices. The stakes are not abstract.
Beltz, Wright, Sprague and Molenaar (2016) built a practical bridge between the two levels. Group-level and single-person symptom models, they showed, can diverge sharply for a given individual. The pattern repeats.
Piccirillo and Rodebaugh (2019) reviewed the wider idiographic literature and reached a consistent verdict. Within-person and between-person findings diverge often enough to justify wider use of idiographic designs in clinical psychology, though inconsistent methods across studies remain a real limitation.
Further Information
- Piantadosi, S., Byar, D. P., & Green, S. B. (1988). The ecological fallacy. American journal of epidemiology, 127(5), 893-904.
- Sedgwick, P. (2015). Understanding the ecological fallacy. Bmj, 351.
References
Beltz, A. M., Wright, A. G. C., Sprague, B. N., & Molenaar, P. C. M. (2016). Bridging the nomothetic and idiographic approaches to the analysis of clinical data. Assessment, 23 (4), 447-458.
Diez-Roux, A. V. (1998). Bringing context back into epidemiology: Variables and fallacies in multilevel analysis. American Journal of Public Health, 88 (2), 216-222.
Fisher, A. J., Medaglia, J. D., & Jeronimus, B. F. (2018). Lack of group-to-individual generalizability is a threat to human subjects research. Proceedings of the National Academy of Sciences, 115 (27), E6106-E6115.
Holmes, M. D., Hunter, D. J., Colditz, G. A., Stampfer, M. J., Hankinson, S. E., Speizer, F. E., … & Willett, W. C. (1999). Association of dietary intake of fat and fatty acids with risk of breast cancer. Jama, 281 (10), 914-920.
King, G. (1997). A solution to the ecological inference problem: Reconstructing individual behavior from aggregate data. Princeton University Press.
Kramer, G. H. (1983). The ecological fallacy revisited: Aggregate-versus individual-level findings on economics and elections, and sociotropic voting. American political science review, 77 (1), 92-111.
Molenaar, P. C. M. (2004). A manifesto on psychology as idiographic science: Bringing the person back into scientific psychology, this time forever. Measurement: Interdisciplinary Research & Perspective, 2 (4), 201-218.
Openshaw, S. (1984). The modifiable areal unit problem (Concepts and Techniques in Modern Geography 38). Geo Books.
Piantadosi, S., Byar, D. P., & Green, S. B. (1988). The ecological fallacy. American journal of epidemiology, 127( 5), 893-904.
Piccirillo, M. L., & Rodebaugh, T. L. (2019). Foundations of idiographic methods in psychology and applications for psychotherapy. Clinical Psychology Review, 71, 90-100.
Robinson, W. S. (2009). Ecological correlations and the behavior of individuals. International Journal of Epidemiology, 38 (2), 337-341.
Sampson, R. J., Raudenbush, S. W., & Earls, F. (1997). Neighborhoods and violent crime: A multilevel study of collective efficacy. Science, 277 (5328), 918-924.
Sedgwick, P. (2015). Understanding the ecological fallacy. Bmj, 351.
Selvin, H. C. (1958). Durkheim’s suicide and problems of empirical research. American Journal of Sociology, 63 (6), 607-619.
Simpson, E. H. (1951). The interpretation of interaction in contingency tables. Journal of the Royal Statistical Society: Series B, 13 (2), 238-241.
Subramanian, S. V., Jones, K., Kaddour, A., & Krieger, N. (2009). Revisiting Robinson: The perils of individualistic and ecologic fallacy. International Journal of Epidemiology, 38 (2), 342-360.