Theoretical sampling is a data collection method used in grounded theory research. It involves collecting and analyzing data simultaneously, with the goal of developing a theory as it emerges.
Unlike traditional research, where sampling decisions are made upfront, theoretical sampling is iterative and guided by the emerging analysis. The sample evolves with the theory.
The researcher analyzes the first data, then seeks out the participants, documents, or observations most likely to enrich the developing categories and their relationships.
This process continues until theoretical saturation is reached, meaning that no new insights are being gained from the data.
Key Takeaways
- Sample Size: The goal is not a large number of participants but strategically chosen sources that offer the most relevant data for building the theory.
- Starting Point: Research begins with a small purposive sample. Later sampling decisions follow the categories that emerge from the analysis.
- Saturation: The researcher collects and analyzes data until new data no longer add new information or insights to the categories.
- Documentation: Researchers must record why they chose particular participants or data sources. This helps demonstrate the systematic and rigorous nature of grounded theory (GT) research.
- Origins: Barney Glaser and Anselm Strauss introduced theoretical sampling in 1967, in The Discovery of Grounded Theory.
Steps of Theoretical Sampling
Theoretical sampling is an iterative process, closely tied to the analysis of the data.
The researcher analyzes each new batch of data. That analysis can yield theoretical insights, which prompt further sampling decisions.
This cycle continues until the researcher reaches theoretical saturation, indicating that no new information is emerging from the data.
Theoretical sampling is employed once the researcher has begun developing tentative categories through initial coding and analysis.
Testing comes next. As the researcher identifies potential relationships between categories, theoretical sampling helps to gather data that can either confirm or refute these hypotheses.
1. Initial Purposive Sampling:
- Begin the research with a small, purposefully selected sample of participants or data sources. This initial sample is chosen based on the researcher’s preliminary understanding of the phenomenon under investigation and the research questions.
- The goal of this initial sampling is not to be representative but rather to gain a broad understanding of the phenomenon and identify potential areas for further exploration.
- The size of the initial sample should be manageable, allowing for in-depth analysis of the data.
2. Data Collection and Initial Coding:
- Collect data from the initial sample, using methods such as interviews, observations, or document analysis.
- Begin coding the data immediately, breaking it down into meaningful segments and assigning codes to these segments.
- This initial coding is open and exploratory, focusing on identifying key themes and concepts.
- Line-by-line coding is recommended for interview data (Charmaz, 2006). Use gerunds (“-ing” words such as managing disclosure) to capture action.
- In-vivo codes use the participants’ own words as labels, which keeps the analysis close to their meanings.
- Code generously at this stage. It is easier to discard a redundant code than to notice a pattern that was never coded (Holton, 2007).
3. Memo Writing and Reflection:
- Simultaneously with coding, write memos to document your analytical thoughts, reflections, questions, and emerging insights.
- Memos serve as a record of your thinking process, allowing you to track the development of your ideas and identify potential areas for further investigation.
- Memos are informal, dated notes to yourself, often messy. They are where a code first grows into a theoretical idea.
- A memo can also pose the next sampling question directly, such as which type of participant to recruit next.
4. Identifying Emerging Themes and Concepts:
- Through the process of coding and memo writing, identify patterns in the data and begin to develop tentative categories or concepts.
- These provisional categories may change as you collect and analyze more data.
- As codes accumulate, cluster them into fewer, more abstract categories.
- Describe each category by its properties (its characteristics) and its dimensions (the range along which each property varies).
- For example, one property of managing disclosure is how much a person tells others, ranging from total concealment to open disclosure.
5. Theoretical Sampling to Refine Categories:
- Based on the emerging categories, make strategic decisions about which participants or data sources to sample next.
- Theoretical sampling is directed by the developing theory and aims to gather data that will help to:
- Saturate Existing Categories: Collect data to further explore the properties and dimensions of existing categories, ensuring that you have captured their full range and variation.
- Develop Relationships Between Categories: Sample data to examine how different categories relate to each other, identifying potential connections, overlaps, or contradictions.
- Explore Emerging Questions: Pursue data that will help answer questions or address gaps that have arisen during analysis.
6. Iterative Process of Data Collection, Analysis, and Sampling:
- Continue the cycle of data collection, coding, memo writing, and theoretical sampling iteratively.
- Each round of data collection and analysis will inform subsequent sampling decisions, leading to a progressively refined and focused theoretical understanding.
7. Achieving Theoretical Saturation:
- Continue theoretical sampling until you reach theoretical saturation, the point at which collecting additional data does not yield new insights or properties within the categories.
- Theoretical saturation indicates that you have sufficiently explored the phenomenon and that your categories are well-developed and conceptually dense.
8. Documentation of Sampling Decisions:
- Throughout the process, meticulously document all sampling decisions, including the rationale for selecting specific participants or data sources.
- This documentation serves as an audit trail, enhancing transparency and allowing you to demonstrate the rigor and systematic nature of your sampling approach.
Practices Related to Theoretical Sampling
Origins in Grounded Theory
Barney Glaser and Anselm Strauss introduced theoretical sampling in The Discovery of Grounded Theory (1967). They built the method from their own fieldwork on how hospital staff and dying patients managed awareness of dying.
The principle is simple. Emerging concepts dictate who or what to study next, so the sample is never fixed in advance.
It served the wider aim of grounded theory: generating theory directly from data, rather than gathering confirmatory data for grand, untested theories. Theoretical sampling was one of several interlocking procedures behind this, alongside constant comparison, coding, memo-writing, and saturation.
Theoretical Sensitivity
Theoretical sensitivity is the researcher’s ability to recognize and extract meaningful patterns and themes from the data that contribute to the developing theory.
Theoretical sensitivity helps the researcher decide where to sample next. It tells them which participants or data sources are most likely to give rich information about underdeveloped parts of the theory.
Together, theoretical sensitivity and theoretical sampling keep the emerging theory constantly refined and grounded in the data. The result is a richer, more insightful understanding of the phenomenon under study.
Constant Comparison
Constant comparison involves continuously comparing data with data, data with codes, codes with codes, and so on.
Constant comparison exposes gaps and uncertainties in the emerging categories. Theoretical sampling then provides a roadmap for gathering specific data to address them.
For example, constant comparison may reveal a missing property of a category. The researcher might then use theoretical sampling to recruit participants likely to possess that property, or to analyze documents that focus on it.
Imagine a study of adjusting to chronic illness.
Comparing interview incidents of managing disclosure at work reveals properties such as who is told, how much is told, and when. Each property varies along a dimension, from full disclosure to total concealment.
Suppose one participant discloses only once symptoms become visible. The analyst’s memo then asks whether visibility changes the whole disclosure strategy. The next sampling decision follows directly: recruit people with an invisible condition and people with a visible one, and compare.
It’s crucial to understand that both constant comparison and theoretical sampling are not one-time activities in GT. They are iterative processes that occur throughout the research.
Theoretical Saturation
Theoretical saturation occurs when collecting additional data about a theoretical category reveals no new properties, dimensions, or insights about the emerging theory.
In essence, saturation means “theoretical completeness.” The categories are richly developed and the relationships between them are well established.
The goal of GT is not to collect a massive amount of data. It is to gather strategically the data most relevant for building and refining the theory.
Saturation is reached category by category. When a category is saturated, further data collection on it is unlikely to contribute any new understanding.
Therefore, the researcher can confidently cease theoretical sampling in those areas.
Four Meanings of Saturation
Researchers do not all mean the same thing by saturation. Data saturation, where new data only repeat what is already known, is not the same as theoretical saturation, where the categories themselves are fully developed.
Saunders et al. (2018) set out to untangle the term.
- Aim: To clarify how saturation is understood across qualitative traditions, and how to apply it consistently with a study’s methodology.
- Method: A conceptual analysis, not an empirical study. The authors reviewed and synthesised how saturation is defined and applied in the qualitative literature, then built a typology.
- Results: Saturation is used to mean several distinct things, and the authors set out four models. Which one applies depends on the analytic approach.
- Conclusion: Researchers should state which model of saturation they use and justify it against their design. For grounded theory, the target is fully developed categories, not repeated content.
The four models differ in what they treat as the point to stop:
- Theoretical saturation: the grounded theory sense, where the categories and emerging theory are fully developed.
- Inductive thematic saturation: no new codes or themes emerge as coding continues.
- A priori thematic saturation: how far codes or themes decided in advance are evidenced in the data.
- Data saturation: new data merely repeat what is already present.
This matters for sampling because a researcher cannot simply stop after a fixed number of interviews. Guest et al. (2006) studied sixty interviews with women in two West African countries. They found that themes reached saturation within the first twelve interviews.
That is a data-saturation benchmark. Grounded theory judges saturation by how fully the categories are developed, not by how many interviews have been done.
Ethical Issues
The flexible nature of theoretical sampling makes it hard to specify the exact population and sampling criteria in advance. Traditional research ethics applications commonly require this.
As the research progresses and new theoretical insights emerge, researchers might need to modify their sampling strategy, requiring amendments to their initial ethics applications.
This can cause delays and require justification. It can also pose challenges when dealing with ethics committees unfamiliar with GT methodology.
Researchers need to ensure that participants are continually informed about the study’s direction and have opportunities to reconsider their participation.
As theoretical sampling leads to collecting data from new sources, researchers must adapt their data management and security protocols to ensure the confidentiality of all participants. This involves:
- De-identifying data promptly and effectively to protect participant anonymity.
- Securely storing data, both physical copies and electronic files, to prevent unauthorized access.
- Developing clear procedures for data sharing and disposal that comply with ethical guidelines and data protection regulations.
Documentation of Decisions
Meticulously documenting theoretical sampling choices creates an audit trail that enhances the transparency of the research process.
Every sampling choice needs a recorded reason. The documentation should outline why each participant or data source was selected at each stage of the research. It should show that these decisions were grounded in the emerging theory, not driven by arbitrary choices or researcher bias.
This transparency allows others to understand how the theory evolved and strengthens the study’s credibility.
Clear documentation matters most when seeking ethical approval. It also helps when responding to queries from ethics committees.
Researchers need to provide a clear explanation of the principles and procedures of theoretical sampling, demonstrating its systematic and rigorous nature.
Examples help too. Showing how sampling decisions might evolve with emerging data can help ethics committees understand the rationale behind this flexible approach.
Analytical memos serve as a primary tool for documenting the researcher’s thinking process, including reflections on theoretical sampling decisions.
A dedicated sampling log can be used to systematically record details of each sampling decision, including:
- Date of the decision.
- Rationale for the decision, explicitly linking it to the emerging theory.
- Specific characteristics of the chosen participant or data source.
- Reflections on the potential contribution of this data to the theory development.
Example
To illustrate how theoretical sampling works in practice, let’s consider a hypothetical grounded theory study examining the experiences of nurses transitioning from hospital settings to community care roles.
Initial Purposive Sampling:
- The researcher might begin by purposefully sampling a small group of nurses who have recently made this transition.
- The initial sample could include nurses with varying levels of experience, different specialties, and from diverse geographical locations to ensure maximum variation in the early data.
Initial Data Collection and Analysis:
- The researcher could conduct in-depth interviews with these nurses, exploring their motivations for the transition, challenges they faced, coping mechanisms they employed, and their perceptions of their new roles.
- Initial coding and analysis of these interviews might reveal several emerging themes:
- Role Adjustment: Nurses describe difficulties adjusting to the increased autonomy, different skill sets, and changing patient dynamics in community care.
- Professional Identity: Nurses express feelings of uncertainty or a sense of loss related to their professional identity as they navigate their new role.
- Support Systems: Nurses highlight the importance of support from colleagues, supervisors, and family during the transition.
Theoretical Sampling Based on Emerging Themes:
- Based on these emerging themes, the researcher could then theoretically sample additional participants to refine and expand the developing categories.
- To further explore “Role Adjustment,” the researcher could seek out nurses who have experienced particularly challenging transitions, perhaps those working in specialized areas of community care or in remote locations with limited resources.
- To understand the nuances of “Professional Identity,” the researcher could interview nurses who have transitioned back to hospital settings after working in community care, examining the factors that influenced their decisions.
- To gain a deeper understanding of “Support Systems,” the researcher could interview supervisors and family members of nurses who have made the transition, exploring their perspectives on the challenges and support needs.
Iterative Process and Saturation:
- This process of data collection, coding, analysis, and theoretical sampling would continue iteratively, with each round informing the next.
- As the researcher collects more data and refines the categories, they would write memos documenting their analytical thoughts, the rationale behind their sampling decisions, and any adjustments made to the research process.
- The researcher would continue this cycle until they reach theoretical saturation, the point at which no new properties are emerging within the categories, and the relationships between categories are well-established.
Theoretical Sampling in Applied Research
The nurse example is hypothetical. Real health-services research uses the same logic. Foley and Timonen (2015) wrote a methodological guide for researchers in this field.
They illustrated grounded theory procedures, including theoretical sampling, with their own study of how people with amyotrophic lateral sclerosis (ALS) experienced health care.
Their argument is that these flexible procedures suit hard-to-reach population groups and topics not studied before. Sampling follows the data, not a fixed plan. A theory of how people engage with health care services can then build up from the data instead of being imposed on it.
Sources
Birks, M., & Mills, J. (2015). Grounded theory: A practical guide (2nd ed.). Sage.
Charmaz, K. (2006). Constructing grounded theory: A practical guide through qualitative analysis. Sage.
Chenitz, W. C., & Swanson, J. M. (1986). From practice to grounded theory: Qualitative research in nursing. Addison-Wesley.
Corbin, J., & Strauss, A. (1990). Grounded theory research: Procedures, canons, and evaluative criteria. Qualitative Sociology, 13(1), 3–21.
Foley, G., & Timonen, V. (2015). Using grounded theory method to capture and analyze health care experiences. Health Services Research, 50(4), 1195–1210. https://doi.org/10.1111/1475-6773.12275
Glaser, B. G. (2005). The grounded theory perspective III: Theoretical coding. Sociology Press.
Glaser, B. G., & Strauss, A. L. (1967). The discovery of grounded theory: Strategies for qualitative research. Aldine.
Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), 59–82. https://doi.org/10.1177/1525822X05279903
Holton, J. A. (2007). The coding process and its challenges. In A. Bryant & K. Charmaz (Eds.), The SAGE handbook of grounded theory (pp. 265–289). Sage. https://doi.org/10.4135/9781848607941.n13
Saunders, B., Sim, J., Kingstone, T., Baker, S., Waterfield, J., Bartlam, B., Burroughs, H., & Jinks, C. (2018). Saturation in qualitative research: Exploring its conceptualization and operationalization. Quality & Quantity, 52(4), 1893–1907. https://doi.org/10.1007/s11135-017-0574-8