Sampling Bias: Types, Examples & How to Avoid It

Sampling bias occurs when a sample does not accurately represent the population being studied. This can happen when there are systematic errors in the sampling process, leading to over-representation or under-representation of certain groups within the sample.

Sampling bias means some members of a population were more likely to be selected than others, so the achieved sample does not represent the whole group.

This directly threatens a study’s external validity: findings from a biased sample can only be generalized to people who share its characteristics, not to the wider population. This bias is systematic, not random chance.

Sample Target Population

In medical research, this same pattern is called ascertainment bias, where one group of participants is over-represented in the sample. Excluding parts of the population this way misleads the results and discards valuable data that the study needed.

Sampling bias occurs during data collection, in the method itself, not in the resulting sample. It often happens without the researcher’s knowledge.

Key Takeaways

  • Definition: Sampling bias is a systematic difference between who is included in a sample and who is excluded, so some groups are over- or under-represented.
  • Not Random: Unlike ordinary sampling error, bias pushes results in one direction and does not shrink as the sample grows.
  • Common Causes: Sampling bias often comes from convenience-based selection, a sampling frame that does not match the target population, or non-response.
  • WEIRD Samples: Much of psychology’s evidence comes from Western, Educated, Industrialized, Rich and Democratic samples, which are frequent outliers on many measures (Henrich et al., 2010).
  • Size Isn’t Enough: A larger sample does not fix a biased sampling method; a huge but poorly sampled study can still be precisely wrong.
  • Fixes: Use random or stratified sampling, clearly define the target population and sampling frame, and follow up on non-responders.

Example of Sampling Bias

Imagine you want to study the prevalence of depression amongst undergraduate students at your university. You send out an email to the whole student body asking for volunteers.

This method leads to sampling bias, because only students who are comfortable discussing their depression will sign up. It is an example of voluntary response bias. Only people willing to talk about their experiences take part, so the sample is not representative.

Types of Sampling Bias

Undercoverage Bias

Undercoverage bias occurs when some population members are inadequately represented in the sample, so they have less chance of being selected than others.

For example, administering a survey online will exclude groups with limited internet access, such as the elderly and those in lower-income households.

Results are only biased if the excluded group differs in some meaningful way from those included, not simply because a group was missed.

Voluntary Response Bias / Self-Selection Bias

Self-selection bias is a type of bias that occurs when participants can choose whether or not to participate in the project.

Bias arises because people with specific characteristics might be more likely to agree to participate in a study than others, making the participants a non-representative sample.

For example, people with strong opinions or substantial knowledge about a specific topic may be more willing to spend time answering a survey than those without.

Volunteers also tend to differ from non-volunteers in predictable ways. Rosenthal and Rosnow’s (1975) analysis found volunteers are, on average, better educated, more sociable and more approval-seeking than people who do not come forward.

This shows up in health research too. In health research this pattern has its own name: the healthy-volunteer effect. People who volunteer for health studies tend to be healthier and more health-conscious than the population the study is meant to represent. This makes treatments look better than they are.

Survivorship Bias 

Survivorship bias refers to when researchers focus on individuals, groups, or observations that have passed some sort of selection process while ignoring those who did not.

In other words, only “surviving” subjects are selected. For example, in finance, failed companies tend to be excluded from performance studies because they no longer exist.

This causes the results to skew higher because only companies that were successful enough to survive are included.

Non-Response Bias

Non-response bias is a type of bias that arises when people who refuse to participate or drop out of a study systematically differ from those who take part.

For example, imagine a study on the prevalence of depression in a community. If people with depression are less likely to take part than those without it, the results will underestimate how common depression really is.

Recall Bias

Recall bias occurs when some members of your sample cannot remember important details accurately. As a result, they might provide incomplete or incorrect information that can distort your research findings.

This type of bias tends to affect retrospective surveys that rely on self-reported data. 

Exclusion Bias 

This bias results from intentionally excluding a particular group from the sample. Exclusion bias is closely related to non-response bias. 

Observer Bias

Observer bias refers to the tendency of observers not to see what is there, but instead to see what they expect or want to see.

This bias can result in an overestimation or underestimation of what is true and accurate, which compromises the validity of your research findings.

For example, researchers might unintentionally influence participants during interviews by focusing on specific statistics that tend to support the hypothesis instead of those that do not.

Causes of Sampling Bias

A common cause of sampling bias lies in the study’s design or data collection process, since researchers may favor or disfavor gathering data from certain people or conditions. Bias also creeps in when researchers rely on judgment or convenience rather than a properly randomized strategy.

This can happen in both probability and non-probability sampling.

Causes in Probability Sampling

In probability sampling, every member of the population has a known, equal chance of being selected, which reduces the risk of sampling bias but does not eliminate it.

Extracting random samples typically requires a sampling frame: a list of the population from which the sample is drawn. A sampling frame does not by itself prevent sampling bias.

A mismatched frame biases the sample. This can happen when a researcher fails to correctly determine the target population, or relies on outdated and incomplete information that excludes parts of it.

Even a properly selected frame can still produce a biased sample through non-response, if certain kinds of people are more likely to refuse to take part.

Mismatches between the sampling frame and the target population, together with non-response, are the main routes to a biased sample in probability sampling.

Causes in Non-Probability Sampling

In non-probability sampling, samples are selected using non-random criteria, such as convenience sampling, where participants are chosen based on accessibility or availability rather than a random process. These sampling techniques often produce biased samples, because some population members are more or less likely to be included than others.

Because the chance of being selected is unknown and often uneven, statistical generalisation to the wider population is far less secure than with probability sampling. This trade-off keeps them cheap, but riskier to generalise from.

How to Avoid Sampling Bias

  • Use random or stratified sampling → Stratified random sampling will help ensure you get a representative research sample and reduce the interference of irrelevant variables in your systematic investigation.
  • Avoid convenience sampling → Rather than collecting data from only easily accessible or available participants, you should gather data from the different subgroups that make up your population of interest. 
  • Clearly define a target population and a sampling frame → Matching the sampling frame to the target population as much as possible will reduce the risk of sampling bias. 
  • Follow up on non-responders → When people drop out or fail to respond to your survey, do not ignore them, but rather follow up to determine why they are unresponsive and see if you can garner a response.  Additionally, you should keep close tabs on your research participants, and follow up with them frequently to reduce attrition.
  • Oversampling → Oversampling can be used to avoid sampling bias in cases where members of the defined population are underrepresented.
  • Aim for a large, well-sampled study → A bigger sample captures more subgroups, but size alone cannot fix a biased sampling method; it must be paired with representative selection.  
  • Set up quotas for each identified demographic → If you think participant gender, age, ethnicity or some other demographic characteristic is a potential source of bias within your study, quotas will allow you to evenly sample people from different demographic groups within the study.

Critical Evaluation of Sampling Bias Research

Sampling bias is not just a technical flaw to correct case by case. It shapes what psychology as a whole can credibly claim about human behaviour, in ways that go well beyond any single study. It also explains why fixing one study’s sample does not fix the field’s evidence base.

The WEIRD Samples Problem

Henrich, Heine and Norenzayan (2010) found that most published psychology findings come from Western, Educated, Industrialized, Rich and Democratic (WEIRD) populations, with American undergraduates hugely over-represented. On measures ranging from visual perception to fairness and moral reasoning, these participants are frequent outliers rather than a neutral human default.

Arnett (2008) made a related point. Top psychology journals drew most of their samples from the United States, leaving what he called the ‘neglected 95%’ of the world.

Sears (1986) had earlier warned that social psychology’s heavy reliance on the college sophomore in the laboratory gives a distorted, age- and cohort-bound picture of human nature.

Convenience sampling is not a minor shortcut here. It is a systematic threat to how far psychology’s findings can generalise.

Why a Large Sample Doesn’t Guarantee an Unbiased One

Size is often mistaken for representativeness, but a large sample gathered by a biased method is simply a precisely wrong estimate. The clearest historical demonstration is the Literary Digest poll ahead of the 1936 US presidential election.

One case makes this vivid. The magazine mailed about ten million straw-vote ballots, drawn from its own subscriber list and from car-registration, telephone and club-membership lists. From the roughly 2.3 million ballots returned, it confidently predicted a landslide win for the Republican Alf Landon.

He lost by a landslide.

The sampling frame had silently excluded poorer, disproportionately Democratic voters who, in the Depression, owned neither a car nor a telephone. The low return rate added non-response bias on top of this frame error.

In the same election, George Gallup correctly called the result from a far smaller quota sample of around thirty thousand people.

The Replication Crisis

The Open Science Collaboration (2015) attempted to replicate 100 psychology studies and reproduced fewer than half of the original effects. The same narrow, convenience-based samples described above sit behind many of the effects that failed to replicate.

Unrepresentative sampling is therefore one reason findings fail to generalise across people and contexts. Sampling is not a peripheral detail. It is part of the field’s credibility problem.

Contemporary Research

Much data collection has moved to online labour markets such as Amazon Mechanical Turk and Prolific, which give faster, cheaper access to more varied samples than the campus subject pool.

Buhrmester, Kwang and Gosling (2011) tested this directly and found MTurk data was at least as reliable as data collected through traditional methods.

These are still convenience samples, with their own biases.

Paolacci and Chandler (2014) found crowdworkers are non-representative in age, income and attitudes, and often ‘non-naïve’: experienced workers have seen common experimental manipulations many times before, which can distort results.

The bias just moves online.

Peer et al. (2017) compared platforms directly and found meaningful differences in attention and data quality between MTurk and alternatives such as Prolific.

A related reform is pre-registration.

It means publicly stating a study’s design, hypotheses and analysis plan before data collection begins. Simmons, Nelson and Simonsohn (2011) showed that flexibility in when to stop collecting data inflates false-positive rates.

Pre-registration now typically fixes the sampling plan in advance, including a target sample size and stopping rules.

That shift is still underway.

Nosek et al. (2018) describe this as a shift toward transparent, confirmatory research rather than decisions made after seeing the data.

FAQs

What is the difference between sampling bias and sampling error?

Sampling error is the natural, random difference between a sample statistic and the true population value. It happens even with a perfectly unbiased sample and shrinks as the sample grows.

Sampling bias is different: it is a systematic distortion that does not wash out with a larger sample.

What is the difference between sampling bias and response bias?

Sampling bias occurs when some members of a population are systematically more likely to be selected in a sample than others. As a result, the sample misrepresents the group.

Response bias is a general term that refers to a wide range of conditions or factors that can lead participants to respond inaccurately or falsely to questions.

For example, there could be something about how the actual survey questionnaire is constructed that encourages a certain type of answer, leading to measurement error.

Which type of sampling is most at risk for sampling bias?

Non-probability sampling, specifically convenience sampling, is most at risk for sampling bias. With this type of sampling, some members of the population are more likely to be included than others.

Does sampling bias affect reliability?

Sampling bias mainly threatens external validity, not reliability. A study can still be reliable, meaning it gives consistent results, even when its sample is biased.

What a biased sample undermines is generalizability: findings may not extend beyond people like those actually sampled.

Why is it important to avoid sampling bias in research?

It is important to avoid sampling bias in research because otherwise, the population of interest will not be accurately represented. If the sample bias is not addressed then, your research loses its credibility.

Is probability sampling biased?

Probability sampling can significantly reduce sampling bias by giving every member of the population an equal chance of being included in the research.
This method can still result in a biased sample if your sampling frame does not match the population of interest.

Can sampling error be calculated?

Yes, sampling error is calculated by dividing the standard deviation of the population by the square root of the size of the sample.
The result is then multiplied by the confidence level.

Here’s the formula for calculating sampling error:

Sampling error = confidence level × [standard deviation of population / (square root of sample size)]

References

Arnett, J. J. (2008). The neglected 95%: Why American psychology needs to become less American. American Psychologist, 63(7), 602–614. https://doi.org/10.1037/0003-066X.63.7.602

Hamill, R., Wilson, T. D., & Nisbett, R. E. (1980). Insensitivity to sample bias: Generalizing from atypical cases. Journal of Personality and Social Psychology, 39(4), 578.

Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83. https://doi.org/10.1017/S0140525X0999152X

Nielsen, M., Haun, D., Kärtner, J., & Legare, C. H. (2017). The persistent sampling bias in developmental psychology: A call to action. Journal of Experimental Child Psychology, 162, 31–38.

Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716

Rosenthal, R., & Rosnow, R. L. (1975). The volunteer subject. Wiley.

Sears, D. O. (1986). College sophomores in the laboratory: Influences of a narrow data base on social psychology’s view of human nature. Journal of Personality and Social Psychology, 51(3), 515–530. https://doi.org/10.1037/0022-3514.51.3.515

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.

Julia Simkus

Psychology Researcher and Writer

BA (Hons) Psychology, Princeton University

Julia Simkus is a Princeton University graduate in Clinical Psychology (Magna Cum Laude) and holds a Master of Arts in Applied Psychology from New York University. During her studies she worked as a research assistant to Professor Nicole Avena at Princeton, co-authoring three published works on food addiction and substance use disorders in peer-reviewed journals and Oxford University Press. She wrote and edited over 70 articles for Simply Psychology between 2021 and 2024.