Cluster Sampling: Definition, Method and Examples

Cluster random sampling is a probability sampling method. Researchers divide a large population into smaller groups known as clusters, then randomly select among those clusters to form a sample.

This is typically used for large populations and sample sizes.

cluster sampling
A cluster sample divides a population into groups, or clusters, then randomly selects a subset of those clusters to study. This approach suits large, widely dispersed, or hard-to-reach populations, and works best when each cluster mirrors the population as a whole.

Cluster sampling reduces the number of participants a study needs when the full population is too large to test as a whole. Each cluster acts as a small-scale representation of that population.

This sampling method reduces the cost and time of a study by increasing efficiency. Researchers sometimes will use pre-existing groups such as schools, cities, or households as their clusters.

Key Terms

  • A sample is the group of participants selected from a target population to represent it in a study. Because testing an entire population is rarely feasible, researchers rely on this smaller group instead.
  • Representative describes how closely a sample mirrors a target population’s key characteristics, such as gender, ethnicity, and socioeconomic status. Psychologists choose their sampling method carefully to avoid sampling bias, the over-representation of one group in the sample, and build a representative sample.
  • Generalisability means the extent to which their findings can be applied to the larger population of which their sample was a part.

Cluster Sampling Techniques

Single-stage cluster sampling

    • A single-stage cluster is a type of cluster sampling where each unit of the chosen clusters is sampled. Researchers will first divide the total sample into a predetermined number of clusters based on how large they want each cluster to be.
    • Then, they randomly select and sample from the clusters and collect data from each individual unit in the selected clusters.

Double-stage cluster sampling

    • In two-stage cluster sampling, researchers will only collect data from a random subsample of individual units within each of the selected clusters to use as the sample.
    • This technique is less precise than single-stage sampling and should only be used when it is too challenging or expensive to test the entire cluster.

Multi-stage cluster sampling

    • This type of cluster sampling follows the same process as double-stage sampling but adds further rounds of random selection. A national survey might first randomly select regions, then randomly select schools within those regions, then classes, then pupils.
    • Each extra stage trades a modest increase in sampling error for a much more manageable and affordable sample to collect.

Applications of Cluster Sampling

Cluster sampling is used when the target population is too large or spread out to study individually. It suits any setting where no full list of individuals exists but natural groups do. Common applications include:

  • Geographic sampling: researchers group individuals within a community, neighborhood, or region into a single cluster, which is especially useful for widely dispersed populations.
  • Market research: cluster sampling is used when researchers cannot collect data from the whole target population directly.
  • Disaster and mortality assessment: researchers estimate mortality rates after wars, famines, or natural disasters, when a full census is impossible.

How to Cluster Sample?

  1. First, choose the target population you wish to study and determine your desired sample size.
  2. Divide the population into clusters. Each cluster should be diverse, mirroring the population’s overall characteristics, and roughly equal in size to the others.
  3. Randomly select clusters to preserve your results’ validity. The number selected depends on your target sample size.
  4. In single-stage sampling, collect data from every individual unit within the clusters you selected in Step 3.
  5. In double-stage or multi-stage sampling, randomly select units within each cluster to sample, then collect your data. This is usually easier than single-stage sampling because the sample is much smaller.

For example, to study teaching practices across a city’s 200 primary schools, a researcher might randomly select 20 of them as clusters. Every teacher within those 20 schools would then be observed, rather than sampling individual teachers from all 200 schools.

Cluster sampling method in statistics. Research on sample collecting data in scientific survey techniques.

Advantages of Cluster Sampling

Time and cost-efficient

Cluster sampling is cheaper and quicker than other sampling methods. For example, it reduces travel expenses for wide geographical populations.

High external validity

External validity is high only when each cluster mirrors the population’s mix of characteristics. If the clusters differ from one another, for example because a high-income region does not reflect the national spread of political views, the sample can be systematically skewed instead.

Practicality and ease

This type of sampling makes large populations manageable to study. No population-wide list of individuals is needed, only a list of the natural clusters, such as schools or postcodes, which is far easier to compile.

Limitations of Cluster Sampling

High sampling error

Sampling error is the natural gap between a sample’s result and the population’s true value.

Cluster sampling typically has a larger sampling error than a simple random sample of the same size. This happens because people within a cluster tend to be similar to one another, so the clusters do not fully mirror the population’s characteristics.

This error grows with more stages of clustering.

Complexity

Planning study designs for cluster sampling usually requires more attention because researchers need to determine how to divide up a larger population efficiently and properly.

Example Situations

Cluster sampling has supported real studies across many fields, including:

  • Assess immunization coverage (Henderson & Sundaresan, 1982).
  • Estimate density of waterfowl wintering (Smith, Conroy, & Brakhage, 1995).
  • Conduct a rapid assessment of health in communities affected by natural disasters (Malilay, Flanders, & Brogan, 1996).
  • Determine forest inventories (Roesch, 1993).
  • Assess the prevalence of irritable bowel syndrome in South China and its impact on health-related quality of life (Xiong, 2004).
  • Estimate the size of hidden and hard to access populations (Medina & Thompson, 2004).

Cluster Sampling vs. Stratified Sampling

Stratified sampling divides a population into smaller subpopulations, or strata, based on shared characteristics such as age, income, race, or education level. Researchers then randomly select members from each stratum, in proportion to its size in the population. This keeps the sample close to the population’s true mix.

Cluster sampling instead uses naturally occurring groups, such as city blocks or school districts, and randomly selects whole clusters rather than individual members.

Because people within a cluster tend to be similar to one another, cluster sampling usually carries a larger sampling error than stratified sampling of the same size.

Key Takeaways

  • Definition: Cluster sampling divides a population into naturally occurring groups, or clusters, and randomly selects whole clusters to study.
  • When Used: It suits large, widely dispersed, or hard-to-reach populations where no full list of individuals exists.
  • Stages: Single-stage sampling studies every member of the chosen clusters; double- and multi-stage sampling randomly subsamples within each cluster instead.
  • Main Strength: No population-wide list is needed, only a list of clusters, which makes it cheaper and more practical than other probability methods.
  • Main Limitation: Because people within a cluster tend to be alike, cluster sampling usually carries a larger sampling error than a simple random sample of the same size.
  • vs Stratified: Stratified sampling selects proportionally from every stratum; cluster sampling selects whole natural groups instead, which is faster but less precise.

References

Felix-Medina, M. H., & Thompson, S. K. (2004). Combining link-tracing sampling and cluster sampling to estimate the size of hidden populations. Journal of Official Statistics, 20(1), 19–38.

Henderson, R. H., & Sundaresan, T. (1982). Cluster sampling to assess immunization coverage: A review of experience with a simplified sampling method. Bulletin of the World Health Organization, 60(2), 253–260.

Malilay, J., Flanders, W. D., & Brogan, D. (1996). A modified cluster-sampling method for post-disaster rapid assessment of needs. Bulletin of the World Health Organization, 74(4), 399–405.

Roesch, F. A. (1993). Adaptive cluster sampling for forest inventories. Forest Science, 39(4), 655–669.

Smith, D. R., Conroy, M. J., & Brakhage, D. H. (1995). Efficiency of adaptive cluster sampling for estimating density of wintering waterfowl. Biometrics, 51(2), 777–788. https://doi.org/10.2307/2532964

Thompson, S. K. (1990). Adaptive cluster sampling. Journal of the American Statistical Association, 85(412), 1050–1059. https://doi.org/10.1080/01621459.1990.10474975

Xiong, L. S., Chen, M. H., Chen, H. X., Xu, A. G., Wang, W. A., & Hu, P. J. (2004). A population-based epidemiologic study of irritable bowel syndrome in South China: Stratified randomized study by cluster sampling. Alimentary Pharmacology & Therapeutics, 19(11), 1217–1224.

Further Information

For more on sampling methods and cluster sampling in practice, see:

Marketing researchers often use city blocks as clusters in cluster sampling. Using this fact, explain how a market researcher might use multistage cluster sampling to select a sample of consumers from all cities having a population of more than 10,000.

In multistage cluster sampling, the process begins by dividing the larger population into clusters, then randomly selecting and subdividing them for analysis.

For market researchers studying consumers across cities with a population of more than 10,000, the first stage could be selecting a random sample of such cities. This forms the first cluster.

The second stage might randomly select several city blocks within these chosen cities – forming the second cluster.

Finally, they could randomly select households or individuals from each selected city block for their study. This way, the sample becomes more manageable while still reflecting the characteristics of the larger population across different cities.

The idea is to progressively narrow the sample to maintain representativeness and allow for manageable data collection.

When is cluster sampling appropriate?

Cluster sampling is appropriate when:

1. The population is widespread geographically, and conducting simple random sampling is costly or impractical. Clusters can be geographically based to minimize travel costs.
2. Data collection involves face-to-face interviews or on-site inspections.
3. A list of individuals in the population is unavailable, but it’s possible to identify clusters representing the population.
4. The population is naturally divided into groups (clusters), and these clusters are internally heterogeneous, i.e., they reflect the diversity of the overall population.

It provides a balance between statistical accuracy and cost-effectiveness in such cases.

What is a cluster sample?

A cluster sample is a sampling method where the researcher divides the entire population into separate groups, or clusters.

Then, a random sample of these clusters is selected. All observations within the chosen clusters are included in the sample.

This method is typically used when the population is large, widely dispersed, and inaccessible. The clusters should ideally mirror the characteristics of the population as a whole.

Saul McLeod, PhD

BSc (Hons) Psychology, MRes, PhD, University of Manchester

Chartered Psychologist (CPsychol)

Saul McLeod, PhD, is a qualified psychology teacher with over 18 years of experience in further and higher education. He has been published in peer-reviewed journals, including the Journal of Clinical Psychology.


Olivia Guy-Evans, MSc

BSc (Hons) Psychology, MSc Psychology of Education

Associate Editor for Simply Psychology

Olivia Guy-Evans is a writer and associate editor for Simply Psychology, where she contributes accessible content on psychological topics. She is also an autistic PhD student at the University of Birmingham, researching autistic camouflaging in higher education.

Julia Simkus

Psychology Researcher and Writer

BA (Hons) Psychology, Princeton University

Julia Simkus is a Princeton University graduate in Clinical Psychology (Magna Cum Laude) and holds a Master of Arts in Applied Psychology from New York University. During her studies she worked as a research assistant to Professor Nicole Avena at Princeton, co-authoring three published works on food addiction and substance use disorders in peer-reviewed journals and Oxford University Press. She wrote and edited over 70 articles for Simply Psychology between 2021 and 2024.