Sampling methods in psychology refer to strategies used to select a subset of individuals (a sample) from a larger population, to study and draw inferences about the entire population. Common methods include random sampling, stratified sampling, cluster sampling, and convenience sampling. Proper sampling ensures representative, generalizable, and valid research results.
Key Terms
- Sampling: the process of selecting a representative group from the population under study.
- Target population: the total group of individuals from which the sample might be drawn.
- Sampling frame: the actual, accessible list of population members from which the sample is drawn, such as a university’s enrolment register. It is rarely a perfect match for the target population.
- Sample: a subset of individuals selected from a larger population for study or investigation. Those included in the sample are termed “participants.”
- Generalizability: the ability to apply research findings from a sample to the broader target population, contingent on the sample being representative of that population.
- Biased sample: a sample that disproportionately represents certain segments of the population, leading to overrepresentation or underrepresentation of specific groups.
For instance, if the advert for volunteers is published in the New York Times, this limits how much the study’s findings can be generalized to the whole population, because NYT readers may not represent the entire population in certain respects (e.g., politically, socio-economically).
The Purpose of Sampling
Psychologists study large groups of people who share a characteristic relevant to the research question. This group is called the target population.
Target Population and Representativeness
In some types of research, the target population might be as broad as all humans. In other types, it might be a much smaller group, such as teenagers, preschool children, or people who misuse drugs. Most research narrows the group considerably.
Studying every person in a target population is impossible. So psychologists select a sample, a sub-group likely to represent the group they are interested in.
This matters for generalization: applying findings from the sample to the target population. A more representative sample earns more confidence in that step.
Representativeness means the sample mirrors the population on the traits that matter. That is what lets a small group stand in for a much larger one. A poorly chosen sample breaks this logic immediately, however large it is.
Sampling Bias
A sample can suffer from sampling bias. This is a systematic difference between those included in the sample and those excluded, so some groups are over- or under-represented.
Common sources include under-coverage, where the sampling frame omits part of the population, and non-response, where people selected decline or drop out. Both can skew results even when the original selection method was sound.
Bias like this is common.
Many psychology studies have a biased sample because they used an opportunity sample of university students as their participants (e.g., Asch). This limits how far the findings can be trusted.
Every study still has to decide how to select its sample.
The choice matters. The method chosen depends on factors such as time and cost. There is no single right answer.

Random Sampling
Random sampling is a type of probability sampling where everyone in the entire target population has an equal chance of being selected.
This is similar to a national lottery. If everyone who enters holds one ticket, then everyone has an equal chance of winning.
Random samples require naming or numbering every member of the target population. A raffle method, such as a random number generator, is then used to choose the sample.
Random sampling is regarded as the ideal method because, in principle, it eliminates systematic sampling bias.
Neyman (1934) gave the formal statistical case for this distinction. Only random selection lets a researcher quantify how far a sample estimate is likely to depart from the true population value. Purposive, judgement-based selection offers no such guarantee.
- Strengths: It is the least biased method, and the sample should represent the target population.
- Weaknesses: A complete list of the population is needed, which is time-consuming and costly to obtain. Even a perfectly random draw can be reintroduced to bias if people who refuse to take part differ from those who agree.
Stratified Sampling
During stratified sampling, the researcher identifies the different types of people that make up the target population and works out the proportions needed for the sample to be representative.
A list is made of each variable (e.g., IQ, gender, etc.) that might have an effect on the research. For example, if we are interested in the money spent on books by undergraduates, then the main subject studied may be an important variable.
Students studying English Literature, for example, may spend more money on books than engineering students. Using a large percentage of either group would therefore skew the results.
The researcher determines the relative percentage of each group at the university, then draws the sample so it contains every group in that same proportion:
- Engineering: 10%
- Social Sciences: 15%
- English: 20%
- Sciences: 25%
- Languages: 10%
- Law: 5%
- Medicine: 15%
- Strengths: The sample should be highly representative of the target population, so results can be generalized with confidence.
- Weaknesses: Gathering such a sample is extremely time-consuming, which is why this method is rarely used in psychology.
Opportunity Sampling
Opportunity sampling is a method in which participants are chosen based on their ease of availability and proximity to the researcher, rather than using random or systematic criteria. It’s a type of convenience sampling.
An opportunity sample is obtained by asking members of the population of interest if they would participate in your research. An example would be selecting a sample of students from those coming out of the library.
- Strengths: It is a quick and easy way of choosing participants.
- Weaknesses: It may not provide a representative sample and could be biased.
Systematic Sampling
Systematic sampling is a method where every nth individual is selected from a list or sequence to form a sample, ensuring even and regular intervals between chosen subjects.
Participants are systematically selected (i.e., orderly/logical) from the target population, like every nth participant on a list of names.
To take a systematic sample, list all the population members and decide on the desired sample size. Divide the population size by the sample size. This gives the sampling interval (n).
If you take every nth name, you will get a systematic sample of the correct size. If, for example, you wanted to sample 150 children from a school of 1,500, you would take every 10th name.
- Strengths: It is quicker than drawing individual random numbers and still spreads the sample evenly across the whole list.
- Weaknesses: If the list has a hidden pattern that lines up with the sampling interval, the sample can end up skewed. For example, every 10th house on a street might be a corner house.
Other Sampling Techniques
Beyond random, stratified, opportunity, and systematic sampling, several other techniques are widely used in psychological research.
Cluster Sampling
Cluster sampling divides the population into naturally occurring groups, called clusters, such as schools, hospitals, or geographic regions, and randomly selects whole clusters to study.
Rather than contacting every registered voter in a country, a researcher might randomly select one or two regions. All voters in those regions are then interviewed by telephone.
To study teaching practices across a city’s 200 primary schools, a researcher could randomly select 20 schools as the clusters. Every teacher in those schools is then observed, rather than sampling individual teachers from all 200 schools.
Multistage sampling extends this logic for very large populations. It combines several steps in sequence: randomly selecting regions, then schools, then classes, then individual pupils within each class.
- Strengths: Cluster sampling is far cheaper and more practical for large, spread-out populations. Data collection concentrates in a few locations, and only a list of clusters is needed, not every individual.
- Weaknesses: A selected cluster can differ systematically from the rest of the population. Cluster sampling therefore usually carries more sampling error than a simple random sample, because people within one cluster tend to resemble each other.
Volunteer Sampling
In volunteer, or self-selected, sampling, participants actively choose to take part, usually by responding to an advertisement or notice.
Because people who want to take part differ from those who do not, self-selection limits how far the findings generalize. The experiences of people who would never volunteer are simply missing from the data.
Hazan and Shaver’s (1987) “Love Quiz” study recruited a self-selecting sample of newspaper readers who chose to send in their answers. This is a recognized limitation of that attachment study, since readers who respond to a love questionnaire may not represent people in general.
- Strengths: Volunteer sampling is easy to organise and ethically straightforward, since participants have chosen to take part. It is useful for reaching people willing to discuss a specific topic.
- Weaknesses: Rosenthal and Rosnow (1975) found volunteers tend to be better educated, of higher social class, more sociable, and more approval-seeking than non-volunteers. They may also hold more extreme opinions, or try to “look good” or justify themselves, which biases the results.
Quota Sampling
Quota sampling is the non-probability counterpart of stratified sampling. The population is divided into subgroups, and a fixed quota of participants is set for each.
Unlike stratified sampling, participants within each subgroup are selected non-randomly, usually by opportunity, until every quota is filled.
A street surveyor might be told to obtain 25 men and 25 women. Once 25 men have completed the survey, the surveyor stops approaching men and continues with women only, until that quota is also full.
- Strengths: Quota sampling prevents any one group from being over-represented. It is quicker and cheaper than stratified sampling, because no full list or random draw is needed.
- Weaknesses: Selection within each quota is left to the interviewer, so quota sampling can smuggle in interviewer bias and cannot support proper statistical inference.
Purposive Sampling
In purposive, or judgement, sampling, the researcher deliberately approaches individuals expected to offer the most detailed or most relevant information for the study.
It is common in qualitative research, where the aim is rich, information-dense data from people with direct experience of the phenomenon. The goal is depth, not a statistically representative sample.
To study the lived experience of early-onset Parkinson’s disease, a researcher might purposively recruit people diagnosed before age 50. They hold relevant experience that a random cross-section of the public would not.
- Strengths: Purposive sampling yields rich, targeted data efficiently, and it is well matched to qualitative inquiry and to studying specialist or expert groups.
- Weaknesses: The researcher’s own judgement, and potential prejudices, determines who counts as “appropriate,” which can bias the sample.
Snowball Sampling
Snowball sampling starts with a few participants who then recruit others they know, so the sample grows through chains of personal referral.
Goodman (1961) gave the method its formal statistical treatment, and Biernacki and Waldorf (1981) developed it as a rigorous technique for studying concealed populations.
To investigate illegal drug use, for example, a researcher might recruit one or two users. These contacts then vouch for the study to acquaintances, reaching a hidden population that no register lists.
- Strengths: Snowball sampling reaches hidden or hard-to-access populations, such as drug users or members of secretive groups, that other methods cannot. Trust carried along referral chains raises participation.
- Weaknesses: Members recruit others like themselves from within their own social networks, so the sample is almost inevitably biased and unrepresentative of the wider population.
Sample Size
The sample size is a critical factor in determining the reliability and validity of a study’s findings. While increasing the sample size can enhance the generalizability of results, it’s also essential to balance practical considerations, such as resource constraints and diminishing returns from ever-larger samples.
Reliability and Validity
-
Reliability is the consistency of findings across occasions, researchers, or instruments. A small sample is more vulnerable to random error and outliers, while a larger sample produces more reliable results.
-
Validity is the accuracy of research findings. A small, unrepresentative sample compromises external validity, so a larger sample that captures more variability generalizes better to the wider population.
-
External vs Ecological Validity: A sample’s size affects external validity, generalising to other people, distinct from ecological validity, generalising to other settings. A reliable measure on an unrepresentative sample still generalises poorly, since reliability is about consistency of measurement, not about who was actually measured.
Statistical Power
In quantitative research, the number of participants needed is usually decided by a power analysis. Statistical power is the probability that a study detects a real effect, and it rises with sample size, effect size, and the significance criterion used. The logic is straightforward.
Cohen (1992) published a widely used “power primer” giving benchmark small, medium and large effect sizes, along with the sample sizes needed to detect each with adequate power. Tools such as G*Power (Faul et al., 2007) let researchers calculate the required sample size before data collection begins.
Under-powered studies miss real effects.
When they do reach significance, they tend to exaggerate the size of the effect. A small effect is genuinely hard to detect. Finding one reliably calls for a correspondingly large sample.
Qualitative Saturation
In qualitative research, sample size is not fixed by power but by saturation: the point at which further interviews or cases yield no new themes. Data collection becomes redundant beyond that point.
In qualitative research the aim is not statistical generalisation beyond the group studied.
Guest, Bunce and Johnson (2006) tracked how new codes emerged across sixty interviews. They found the basic thematic structure of a fairly homogeneous group was often in place after about twelve interviews. Most new codes appeared within the first six.
This gives qualitative researchers a defensible, evidence-based rule of thumb, rather than an arbitrary target sample size. Depth and richness matter more than sample size, so small, purposively chosen samples are entirely acceptable once they reach saturation.
The 1936 Literary Digest Poll
A large sample gathered by a biased method is not a safeguard. It only makes the wrong number more precise.
The Literary Digest magazine mailed around ten million straw-vote ballots ahead of the 1936 US presidential election. The list was drawn from its own subscribers and from car-registration, telephone and club-membership records. On the roughly 2.3 million ballots returned, it confidently predicted a landslide for the Republican, Alf Landon.
The result was not close.
Landon then lost in one of the largest landslides in US history. The sampling frame had silently excluded poorer, disproportionately Democratic voters who, in the depths of the Depression, owned neither a car nor a telephone. The low return rate added non-response bias on top.
In the same election, George Gallup correctly called the result using a far smaller quota sample of around thirty thousand people. A well-targeted small sample beat a huge biased one, because representativeness, not sheer size, is what licenses generalisation.
Practical Considerations
-
Resource Constraints: Larger samples demand more time, money, and resources. Data collection becomes more extensive, data analysis more complex, and logistics more challenging.
-
Diminishing Returns: Beyond a certain point, adding more participants yields only marginal benefit. Going from 50 to 500 participants can transform a study’s robustness, but going from 10,000 to 10,500 adds little for the extra cost.
Key Takeaways
- Sampling: Researchers study a sample because testing an entire target population is rarely possible.
- Sampling Bias: Occurs when the sample fails to represent the population, skewing results.
- Probability Methods: Random, stratified, systematic, and cluster sampling give everyone a known chance of selection and support statistical inference.
- Non-Probability Methods: Opportunity, volunteer, quota, purposive, and snowball sampling rely on convenience or judgement, so representativeness is less certain.
- Sample Size: Larger samples reduce random error, but size never fixes a biased sampling method.
- Generalizability: A representative sample lets findings generalize to the wider population; an unrepresentative one does not.
Critical Evaluation
Sampling is not just a technical step; it shapes how far a study’s conclusions can be trusted. Two influential critiques challenge how representative psychological research really is.
The WEIRD Samples Problem
Henrich, Heine and Norenzayan (2010) reviewed evidence across visual perception, fairness, cooperation and moral reasoning. They tested whether participants from Western, Educated, Industrialised, Rich and Democratic (WEIRD) societies represent humanity as a whole.
They found WEIRD participants are frequent outliers rather than a neutral default, making them among the least representative populations for generalising about human psychology. The pattern held across every domain tested.
American undergraduates are especially over-represented, even though most published claims are framed as describing human psychology in general, not just the psychology of one narrow, unusual group.
Arnett (2008) found that top psychology journals draw the great majority of their samples from the United States, sidelining the rest of the world’s population. Sears (1986) made a related point decades earlier: social psychology’s heavy reliance on undergraduate students gives a distorted, age-bound picture of human nature.
The Replication Crisis
The Open Science Collaboration (2015) attempted to replicate 100 psychology studies, originally published in three leading journals in 2008. Fewer than half of the original effects reproduced.
The team judged success against several strict criteria, not just a single significance test. Estimates ranged from around a third to about half of the original effects, depending on which criterion was used.
Some effects held up cleanly across every criterion tested; others collapsed under all of them. The pattern was not uniform.
This matters directly for sampling: many of the original studies relied on narrow, convenience-based participant pools of exactly the kind the WEIRD critique describes. Unrepresentative sampling is therefore not a peripheral detail.
It is part of why findings fail to generalise across people and contexts.
The Gap Between Theory and Practice
Textbooks present probability sampling, especially simple random sampling, as the gold standard. In practice, most psychological research relies on opportunity or volunteer samples instead. The reason is practical, not principled.
No complete sampling frame exists for most human populations, and budgets are limited, so researchers turn to whoever is easiest to reach, often undergraduate students.
This produces a persistent gap between how sampling is taught and how it is practised. Much published research generalises further than its sampling method can actually support.
Even a technically random draw is not immune. Participants who refuse or drop out are rarely a random subset of those invited, so the realised sample ends up less random than the design intended.
Large samples do not fix this: a big sample gathered by a biased method is simply a precisely wrong estimate.
A well-designed study therefore needs a defensible sampling method, not just careful measurement afterwards.
Contemporary Research
Much modern data collection has moved to online labour markets such as Amazon Mechanical Turk and Prolific. These give faster, cheaper access to more varied samples than the campus subject pool.
Buhrmester, Kwang and Gosling (2011) compared MTurk data against traditional recruitment methods and found it was at least as reliable. They are not risk-free, though.
Paolacci and Chandler (2014) found crowdworkers are non-representative in age, income and attitudes. Many are also non-naive: experienced workers have seen common manipulations and attention checks many times, which can distort results.
Peer et al. (2017) compared platforms directly. They found meaningful differences in attention and honesty between MTurk and alternatives such as Prolific.
A key reform from the replication crisis is pre-registration: publicly specifying the design, hypotheses and analysis before data collection begins. Nosek et al. (2018) describe this as a way of fixing the sampling plan and analysis decisions in advance. This makes the process transparent, not decided after the fact.
References
Arnett, J. J. (2008). The neglected 95%: Why American psychology needs to become less American. American Psychologist, 63(7), 602–614. https://doi.org/10.1037/0003-066X.63.7.602
Biernacki, P., & Waldorf, D. (1981). Snowball sampling: Problems and techniques of chain referral sampling. Sociological Methods & Research, 10(2), 141–163. https://doi.org/10.1177/004912418101000205
Buhrmester, M., Kwang, T., & Gosling, S. D. (2011). Amazon’s Mechanical Turk: A new source of inexpensive, yet high-quality, data? Perspectives on Psychological Science, 6(1), 3–5. https://doi.org/10.1177/1745691610393980
Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155–159. https://doi.org/10.1037/0033-2909.112.1.155
Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175–191. https://doi.org/10.3758/BF03193146
Goodman, L. A. (1961). Snowball sampling. The Annals of Mathematical Statistics, 32(1), 148–170. https://doi.org/10.1214/aoms/1177705148
Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), 59–82. https://doi.org/10.1177/1525822X05279903
Hazan, C., & Shaver, P. R. (1987). Romantic love conceptualized as an attachment process. Journal of Personality and Social Psychology, 52(3), 511–524. https://doi.org/10.1037/0022-3514.52.3.511
Henrich, J., Heine, S. J., & Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences, 33(2–3), 61–83. https://doi.org/10.1017/S0140525X0999152X
Neyman, J. (1934). On the two different aspects of the representative method: The method of stratified sampling and the method of purposive selection. Journal of the Royal Statistical Society, 97(4), 558–625. https://doi.org/10.2307/2342192
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606. https://doi.org/10.1073/pnas.1708274114
Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716. https://doi.org/10.1126/science.aac4716
Paolacci, G., & Chandler, J. (2014). Inside the Turk: Understanding Mechanical Turk as a participant pool. Current Directions in Psychological Science, 23(3), 184–188. https://doi.org/10.1177/0963721414531598
Peer, E., Brandimarte, L., Samat, S., & Acquisti, A. (2017). Beyond the Turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology, 70, 153–163. https://doi.org/10.1016/j.jesp.2017.01.006
Rosenthal, R., & Rosnow, R. L. (1975). The volunteer subject. Wiley.
Sears, D. O. (1986). College sophomores in the laboratory: Influences of a narrow data base on social psychology’s view of human nature. Journal of Personality and Social Psychology, 51(3), 515–530. https://doi.org/10.1037/0022-3514.51.3.515
