Trustworthiness in qualitative research is akin to the concepts of validity and reliability in quantitative research.
Trustworthiness is the degree of confidence researchers have in the accuracy and truthfulness of their findings. A single test statistic cannot certify it. Researchers build it into the research process, then show readers how the data were gathered, interpreted and checked.
Egon Guba (1981) proposed a model for evaluating the trustworthiness of naturalistic inquiries, which is often applied to qualitative research. The model consists of four key criteria:
- Credibility (parallel to internal validity): Ensuring the findings accurately represent the reality of the participants, achieved through methods like prolonged engagement, peer debriefing (review by an outside colleague), and member checking (participants verifying interpretations).
- Transferability (parallel to external validity): Providing sufficient contextual information to allow readers to judge whether findings can be applied to other settings.
- Dependability (parallel to reliability): Demonstrating the research process is logical, traceable, and clearly documented.
- Confirmability (parallel to objectivity): Ensuring findings emerge from the data rather than researcher bias, through methods like audit trails and reflexivity.
While these criteria are presented as distinct, they are interconnected and mutually supportive.
They highlight the importance of rigor, transparency, and reflexivity in conducting and reporting qualitative studies.
Meeting these criteria makes qualitative findings more credible. It also builds a more trustworthy and meaningful body of knowledge.
Policymakers and practitioners rely on trustworthy research to make informed decisions.
Trustworthy qualitative studies provide valuable insights into human experiences and perspectives, which can be used to develop effective policies and improve service delivery.
Researchers have carried the framework into applied fields. Krefting (1991) translated Guba’s four criteria into concrete strategies for occupational therapy researchers.
Malterud (2001), writing in The Lancet, proposed relevance, validity and reflexivity as core standards for qualitative health research. Malterud also warned against treating checklists as a substitute for judgment.
1. Credibility (Truth Value)
Credibility, akin to internal validity in quantitative research, focuses on establishing confidence in the ‘truth’ of the findings.
It asks whether the findings are accurate and plausible. Do they reflect participants’ perspectives and experiences in a meaningful, believable way within the context where the study took place?
Strategies for Enhancing Credibility
- Prolonged engagement: Spending sufficient time in the field to build rapport with participants and gain a deep understanding of their perspectives.
- Persistent observation: Concentrating on the aspects of the setting that prolonged engagement has shown matter most, and examining them in depth (Lincoln & Guba, 1985).
- Triangulation: Using multiple data sources, methods, or researchers to corroborate findings and enhance their credibility.
- Member checking: Allowing participants to review and verify the accuracy of the data, interpretations, and conclusions.
Four Types of Triangulation
Triangulation strengthens an interpretation by showing that it holds up from more than one direction. Denzin (1978) distinguished four types:
- Data triangulation: Drawing on different data sources, such as comparable groups at different sites or the same group at different times.
- Researcher triangulation: Using more than one researcher to collect or interpret the same data, so a second analyst can confirm the themes.
- Theoretical triangulation: Approaching the same observations through more than one theoretical framework, which forces the researcher to justify the lens chosen.
- Methodological triangulation: Combining different methods on one topic, such as a questionnaire followed by interviews.
Combining methods rarely brings a study closer to one objective truth. Instead, it produces related but genuinely different sets of meaning, which the researcher must still reconcile through interpretive judgment.
Other Credibility Techniques
Researchers also use further checks. Each guards against a different way an account could drift from what participants meant.
- Establishing rapport: Building enough trust that responses reflect what participants actually think, not what they believe is expected.
- Negative case analysis: Deliberately searching for cases that do not fit an emerging interpretation, then revising it or stating its limits (Lincoln & Guba, 1985). A write-up can look credible simply by quietly setting awkward exceptions aside.
- Iterative questioning: Returning to a topic later in the interview, rephrased, to see whether the same picture holds up. It helps most on sensitive topics, where a guarded first answer may soften with rapport.
- Bracketing: Deliberately setting aside your own preconceptions before and during analysis (Tufford & Newman, 2010). Reflexivity differs: it accepts that assumptions cannot be fully suspended and makes their influence visible.
2. Transferability (Applicability)
Transferability is similar to external validity in quantitative research, and examines the extent to which the findings of a particular inquiry can be applied to other contexts or settings.
Generalizability is often not the goal in qualitative research. Transferability instead asks whether findings “fit into contexts outside the study situation that are determined by the degree of similarity or goodness of fit between the two contexts”.
Transferability is primarily the responsibility of the reader who wants to apply the findings to a new situation.
Guba (1981) called this fittingness: the degree of similarity, or goodness of fit, between the studied context and the reader’s own. The original study cannot establish that fit alone.
Strategies for Enhancing Transferability
- Thick Description: Providing rich, thick descriptions of the research context and participants.
- Sampling Detail: Clearly outlining the sampling procedures to show how participants were selected.
- Sample Relevance: Discussing the characteristics of the sample and their relevance to other groups.
Thick description supplies what readers need to judge fit. The term comes from the philosopher Ryle (1968/2009). A thin account describes a rapid eyelid contraction; a thick account explains it as a wink, a private signal to a co-conspirator. Geertz (1973) later developed the idea for anthropology.
3. Dependability (Consistency)
Dependability is comparable to reliability in quantitative research, and addresses the stability and consistency of the research findings over time.
It seeks to ensure that if the study were replicated with similar participants in a comparable context, consistent findings would emerge.
Strategies for Enhancing Dependability
- Audit Trail: Maintaining a comprehensive audit trail of all research decisions and modifications.
- Systematic Documentation: Employing systematic documentation methods, including dense descriptions of research procedures and the use of consistent coding techniques.
- Replication Checks: Engaging in peer debriefing and using stepwise replication (independent sub-teams analyzing the same data) or code-recode procedures (recoding after an interval) to assess stability.
Stepwise replication splits a research team into sub-teams that analyze the same data independently (Lincoln & Guba, 1985). Because the teams compare notes at agreed points, a disagreement can be traced to the stage where it arose.
The code-recode procedure works for a single researcher.
They recode a section of data after a deliberate interval, at least two weeks according to Lincoln and Guba (1985), and compare the two passes. Persistent disagreement means the coding scheme needs revising.
An audit trail is a record of the raw data, coding decisions and reasoning, kept from the start of a study. An outside auditor can review it. A dependability audit checks that the methods were applied consistently with accepted qualitative practice (Lincoln & Guba, 1985).
Halpern (1983) listed six categories an audit trail should retain:
- Raw Data: Field notes, transcripts and recordings.
- Data Reduction: Codes, memos and working summaries.
- Data Reconstruction: The themes, interpretations and conclusions built from the codes.
- Process Notes: Methodological decisions and the difficulties met along the way.
- Researcher Intentions: The original proposal and a reflexive journal.
- Instrument Development: Pilot forms and interview guides.
4. Confirmability (Neutrality)
This criterion replaces the concept of objectivity in quantitative research. It emphasizes the degree to which the findings are a product of the data, not of the researcher’s biases or interpretations.
Confirmability is usually demonstrated through an audit trail, which lets others trace how the conclusions were reached. An outside auditor can then run a confirmability audit, checking that the findings trace back to the data (Lincoln & Guba, 1985).
Sources of Bias in Qualitative Research
Qualitative researchers cannot eliminate bias, because the researcher is part of the instrument that produces the data. The discipline’s response is to make bias visible and account for it. Bias can start with participants or with the researcher.
Participant bias distorts what people tell the researcher:
- Acquiescence bias: A tendency to agree regardless of the question, countered by open-ended, neutral questions.
- Social desirability bias: Answering to appear likeable or acceptable, reduced by non-judgmental question phrasing.
- Dominant respondent bias: One group-interview participant crowding out the others, managed by the moderator actively sharing out turns.
- Sensitivity bias: Distorted answers to sensitive questions, addressed by building rapport, stressing confidentiality and raising sensitive topics gradually.
Researcher bias distorts what the researcher notices or reports:
- Confirmation bias: Noticing supporting information and discounting the rest. An independent observer can repeat the analysis, and reflexivity is the primary safeguard.
- Leading-question bias: Wording that nudges a respondent toward a particular answer, addressed with open-ended, neutral follow-ups in the participant’s own language.
- Question-order bias: Earlier answers shaping later ones, addressed by moving from general to specific questions.
- Sampling bias: A sample unsuited to the research aims, such as “professional participants” used to taking part in research.
- Biased reporting: Some findings omitted or under-represented in the write-up, countered by reflexivity and by reporting disconfirming data.
Reflexivity: Two Frameworks
Reflexivity and triangulation are the techniques most directly aimed at managing bias. Willig (2001) distinguished two kinds of reflexivity.
Epistemological reflexivity asks what the method itself can and cannot reveal. Participants who know they are observed may change their behavior, so conclusions need qualifying.
Personal reflexivity concerns the researcher. Their values, experiences and expectations may have shaped how they read the data. A researcher who has overcome a difficult experience might, for example, add a second interviewer as an independent check.
Walsh (2003) proposed four dimensions instead. They are personal, interpersonal, methodological and contextual. Interpersonal reflexivity looks at how relationships within a team, and between researcher and participants, shape what is said. Contextual reflexivity looks at how the wider cultural and historical moment shapes the questions asked.
Strategies for Enhancing Confirmability
- Peer debriefing: Consulting with colleagues or experts to verify interpretations and minimize researcher bias.
- Member checking: Seeking participant feedback to ensure the findings accurately reflect their experiences.
- Reflexive journaling: Documenting researcher reflections and biases to enhance transparency and mitigate subjectivity.
Critical Evaluation
Lincoln and Guba (1985) built their four criteria as a deliberate naturalistic parallel to the positivist criteria used in quantitative research. That gave qualitative researchers a vocabulary that reviewers trained in quantitative methods could recognize.
Critics argue it also risks importing positivist assumptions, such as a single true account to triangulate toward (Yardley, 2000). Several frameworks have tried to escape that tension.
Alternative Frameworks
- Authenticity Criteria: Lincoln and Guba (1986) added five criteria asking what the research did for and with its participants. They are fairness, plus ontological, educative, catalytic and tactical authenticity.
- Yardley (2000): Writing for health psychology, Yardley proposed four flexible principles: sensitivity to context, commitment and rigor, transparency and coherence, and impact and importance.
- Tracy’s Big Tent: Tracy (2010) proposed eight criteria broad enough to cover very different traditions: worthy topic, rich rigor, sincerity, credibility, resonance, significant contribution, ethics and meaningful coherence.
- Methodological Integrity: Levitt et al. (2017) replaced the fixed checklist with two questions: fidelity to the subject matter and utility in achieving the study’s research goals.
None of these has replaced Lincoln and Guba’s criteria, which remain the most widely taught starting point.
Limits of Common Techniques
Critics also question whether the techniques themselves deliver what they promise.
- Checklist Rigor: Barbour (2001) argued that triangulation, purposive sampling and respondent validation are often adopted as checklist items without a deeper rationale. Listing them signals compliance, not demonstrated credibility.
- Triangulation Illusion: Agreement across sources can look like independent corroboration when one shared bias, such as the researcher’s expectations, shapes every source (Barbour, 2001; Yardley, 2000).
- Member Checking: Birt et al. (2016) argued it is applied so inconsistently that it can be “a nod to validation” rather than a real test. Kullman and Chudyk (2025) respond with a five-step participatory model.
- Performative Reflexivity: Trundle et al. (2025) and Sibbald et al. (2025) argue that reflexivity statements have often become brief, formulaic paragraphs rather than evidence of sustained practice.
- Western Framing: Smith (2012) and Thambinathan and Kinsella (2021) argue that credibility standards are culturally situated Western constructs that can silence other ways of establishing rigor.
Contemporary Research
Recent work has moved in two directions: formalizing how credibility techniques are reported, and checking whether they work as claimed.
The APA’s reporting standards for qualitative research (Levitt et al., 2018) now ask authors to report exactly which techniques they used and how. Asserting rigor is no longer enough.
Building Trustworthiness Into Analysis
Nowell et al. (2017) offer a worked example of building trustworthiness into every phase of thematic analysis.
- Aim: To give practical guidance for conducting a thematic analysis that meets recognized trustworthiness criteria.
- Method: The authors laid out an auditable decision trail for each phase of thematic analysis. Their own mixed-methods case study of Strategic Clinical Networks in Alberta, Canada, served as the example.
- Results: Each of the four criteria could be turned into concrete, describable actions at specific points in the analysis. Credibility, for example, came from investigator triangulation during coding and member checking during interpretation.
- Conclusion: Trustworthiness has to be built throughout the analysis, not asserted afterwards. A documented audit trail makes the analysis easier for readers to trace and verify.
It is a practical guide, not an effectiveness test.
Reflexivity in Practice
Critics argue that published reflexivity statements are often brief and formulaic. De Smet et al. (2026) studied what researchers actually do while collecting and analyzing data.
- Aim: To examine reflexivity as it happens during research, not only as it is written up afterwards in a methods section.
- Method: In a researching-the-researcher design, seven junior researchers conducting an interview study were studied. Their in-the-moment reflexive questions and responses were analyzed with reflexive thematic analysis, alongside a post-research focus group.
- Results: The researchers practiced reflexivity as an ongoing, highly individual process of self-examination that extended beyond the interview into their own lives. Different researchers used the word to describe genuinely different things.
- Conclusion: Reflexivity cannot be captured in one retrospective statement. The authors recommend structured “reflexivity labs”, where a research team builds this capacity together.
This gives direct observational support to the critique of performative reflexivity above.
AI-Assisted Analysis
A newer strand applies the same criteria to a context Lincoln and Guba could not have anticipated: qualitative analysis done with artificial intelligence tools.
- Aim: To test whether AI-assisted qualitative analysis can meet the same rigor criteria expected of researcher-led analysis.
- Method: Lazarus et al. (2026) evaluated their own case study against Lincoln and Guba’s four criteria and reflexivity. It used a language-parsing tool with epistemic network analysis on medical students’ reflective diary entries about tolerating clinical uncertainty.
- Results: The approach identified three patterns of reflective depth across many entries efficiently. Meeting each criterion, however, required extra reporting, such as documenting the tool’s training data and configuration for dependability.
- Conclusion: AI assists analysis but does not reduce the researcher’s responsibility to build and report the criteria. It adds new things to report rather than replacing any.
The criteria still apply, but each asks researchers to report more.
Key Takeaways
- Four Criteria: Trustworthiness rests on credibility, transferability, dependability and confirmability, the qualitative parallels to internal validity, external validity, reliability and objectivity.
- Credibility: Researchers build it through prolonged engagement, triangulation, member checking and peer debriefing, not a single test.
- Transferability: Readers judge whether findings fit their own context, so thick description of the setting and sample is essential.
- Audit Trail: A retained record of decisions supports dependability and confirmability, and an outside auditor can review it.
- Beyond Checklists: Techniques help only when applied thoughtfully, so reporting a technique is not the same as practicing it well.
Reading List
- Anney, V. N. (2014). Ensuring the quality of the findings of qualitative research: Looking at trustworthiness criteria.
- Barbour, R. S. (2001). Checklists for improving rigour in qualitative research: A case of the tail wagging the dog?. BMJ, 322(7294), 1115–1117.
- Birt, L., Scott, S., Cavers, D., Campbell, C., & Walter, F. (2016). Member checking: A tool to enhance trustworthiness or merely a nod to validation? Qualitative Health Research, 26(13), 1802–1811.
- De Smet, M. M., Ekşi, E., & Truijens, F. L. (2026). Reflexivity in action: A qualitative study on how researchers interpret and practice reflexivity. Methods in Psychology, 15, Article 100273.
- Denzin, N. K. (1978). The research act: A theoretical introduction to sociological methods (2nd ed.). McGraw-Hill.
- Geertz, C. (1973). The interpretation of cultures: Selected essays. Basic Books.
- Guba, E. G. (1981). Criteria for assessing the trustworthiness of naturalistic inquiries. ECTJ, 29(2), 75–91.
- Halpern, E. S. (1983). Auditing naturalistic inquiries: The development and application of a model [Unpublished doctoral dissertation]. Indiana University.
- Krefting, L. (1991). Rigor in qualitative research: The assessment of trustworthiness. American Journal of Occupational Therapy, 45(3), 214–222.
- Kullman, S. M., & Chudyk, A. M. (2025). Participatory member checking: A novel approach for engaging participants in co-creating qualitative findings. International Journal of Qualitative Methods, 24.
- Lazarus, M. D., Zhao, L., Gibson, A., Martinez-Maldonado, R., & Stephens, G. C. (2026). Risky or rigorous? Developing trustworthiness criteria for AI-supported qualitative data analysis. Anatomical Sciences Education, 19(2), 330–337.
- Levitt, H. M., Bamberg, M., Creswell, J. W., Frost, D. M., Josselson, R., & Suárez-Orozco, C. (2018). Journal article reporting standards for qualitative primary, qualitative meta-analytic, and mixed methods research in psychology: The APA Publications and Communications Board task force report. American Psychologist, 73(1), 26–46.
- Levitt, H. M., Motulsky, S. L., Wertz, F. J., Morrow, S. L., & Ponterotto, J. G. (2017). Recommendations for designing and reviewing qualitative research in psychology: Promoting methodological integrity. Qualitative Psychology, 4(1), 2–22.
- Lincoln, Y. S., & Guba, E. G. (1982). Establishing dependability and confirmability in naturalistic inquiry through an audit.
- Lincoln, Y. S., & Guba, E. G. (1985). Naturalistic inquiry. Sage.
- Lincoln, Y. S., & Guba, E. G. (1986). But is it rigorous? Trustworthiness and authenticity in naturalistic evaluation. New Directions for Program Evaluation, 1986(30), 73–84.
- Malterud, K. (2001). Qualitative research: Standards, challenges, and guidelines. The Lancet, 358(9280), 483–488.
- Nowell, L. S., Norris, J. M., White, D. E., & Moules, N. J. (2017). Thematic analysis: Striving to meet the trustworthiness criteria. International Journal of Qualitative Methods, 16.
- Ryle, G. (2009). The thinking of thoughts: What is ‘Le Penseur’ doing? In Collected papers (Vol. 2, pp. 494–510). Routledge. (Original work published 1968)
- Schwandt, T. A., Lincoln, Y. S., & Guba, E. G. (2007). Judging interpretations: But is it rigorous? Trustworthiness and authenticity in naturalistic evaluation. New Directions for Evaluation, 2007(114), 11–25.
- Shenton, A. K. (2004). Strategies for ensuring trustworthiness in qualitative research projects. Education for Information, 22(2), 63-75.
- Sibbald, K. R., Phelan, S. K., Beagan, B. L., & Pride, T. M. (2025). Positioning positionality and reflecting on reflexivity: Moving from performance to practice. Qualitative Health Research.
- Smith, L. T. (2012). Decolonizing methodologies: Research and indigenous peoples (2nd ed.). Zed Books.
- Thambinathan, V., & Kinsella, E. A. (2021). Decolonizing methodologies in qualitative research: Creating spaces for transformative praxis. International Journal of Qualitative Methods, 20.
- Tracy, S. J. (2010). Qualitative quality: Eight “big-tent” criteria for excellent qualitative research. Qualitative Inquiry, 16(10), 837–851.
- Trundle, C., Araújo, N., Khan, S., & Phillips, T. (2025). Beyond the mirror: Challenging the common assumptions of reflexivity in qualitative research. International Journal of Qualitative Methods, 24.
- Tufford, L., & Newman, P. (2010). Bracketing in qualitative research. Qualitative Social Work, 11(1), 80–96.
- Walsh, R. (2003). The methods of reflexivity. The Humanistic Psychologist, 31(4), 51–66.
- Willig, C. (2001). Introducing qualitative research in psychology: Adventures in theory and method. Open University Press.
- Yardley, L. (2000). Dilemmas in qualitative health research. Psychology & Health, 15(2), 215–228.