When researchers evaluate a quantitative study, they reach for familiar tools: reliability coefficients, validity tests, statistical significance. But what happens when the research isn’t built on numbers at all? Qualitative research – the kind that explores lived experiences, social meanings, and complex human behaviour – operates under a fundamentally different logic. Applying quantitative yardsticks to it doesn’t just produce awkward results; it misrepresents what qualitative work is actually trying to do. This is exactly why researchers Yvonna Lincoln and Egon Guba proposed a distinct set of criteria for evaluating qualitative data – criteria centred on trustworthiness, and built around concepts like credibility, transferability, dependability, and confirmability.
Table of Contents
- Why qualitative research needs its own evaluation criteria
- The four pillars of trustworthiness
- Credibility
- Transferability
- Dependability
- Confirmability
- Credibility in depth: prolonged engagement and persistent observation
- Prolonged engagement
- Persistent observation
- Other key strategies for credibility
- Linking method-specific criteria to research design
- Why this matters beyond the academy
Why qualitative research needs its own evaluation criteria
The core tension is this: quantitative research seeks to measure an objective reality with precision and replicability. Qualitative research, by contrast, seeks to understand how people interpret and experience their world – and that world is shaped by context, culture, and subjectivity. Unlike quantitative research, which deals primarily with numerical data under a strictly objective paradigm, qualitative research handles non-numerical information whose interpretation is tied to human perspectives and experience. Applying the same evaluation criteria to both is not just inappropriate – it actively distorts the qualitative enterprise.
This is why, in their landmark 1985 work Naturalistic Inquiry, Lincoln and Guba proposed trustworthiness as the appropriate standard for assessing qualitative rigor. Trustworthiness refers to the assessment of the quality and worth of the complete study, helping determine how closely findings reflect the aims of the study according to the data provided by participants. The framework offers parallel – but not identical – alternatives to the quantitative concepts of internal validity, external validity, reliability, and objectivity.
The four pillars of trustworthiness
Egon Guba’s model for evaluating trustworthiness consists of four key criteria: credibility, transferability, dependability, and confirmability – and while they are presented as distinct, they are interconnected and mutually supportive. Each addresses a different dimension of research quality, and each requires specific strategies to achieve.
Credibility
Credibility is the qualitative parallel to internal validity – and it is arguably the most critical of the four. Credibility is concerned with the aspect of truth-value: do the findings accurately represent the reality of those being studied? It asks whether the researcher’s interpretations genuinely reflect the perspectives of participants, not just their own assumptions or biases.
According to Lincoln and Guba (1985), credibility is established through prolonged engagement, persistent observation, triangulation, peer debriefings, negative case analysis, referential adequacy, and member checks. Of these, prolonged engagement and persistent observation deserve particular attention, as they are the most distinctly qualitative in nature.
Transferability
Transferability is the qualitative counterpart to external validity or generalisability. Rather than claiming that findings apply universally, it asks whether they are applicable to similar contexts. Transferability addresses the applicability of findings to similar contexts or individuals – not to broader populations – and can be achieved through a “thick description” of findings from multiple data collection methods. This rich contextual description is what allows readers to judge for themselves whether findings resonate with their own situations.
Dependability
Dependability parallels reliability. It asks whether the research process is logical, traceable, and well-documented – so that another researcher could, in principle, follow the same path and reach comparable conclusions. The research process should be logical and transparent such that the process and procedures can be auditable and traced, ensuring coherence across methods and findings. An audit trail – a thorough record of research decisions – is the primary mechanism for achieving this.
Confirmability
Confirmability is concerned with objectivity: ensuring that the findings reflect the data and the voices of participants, rather than the researcher’s own preferences or preconceptions. Confirmability is achieved through methods like audit trails and reflexivity – ensuring findings emerge from the data rather than researcher bias. Reflexivity, in practice, means that researchers actively reflect on how their own position, experiences, and assumptions might be shaping their interpretations.
Credibility in depth: prolonged engagement and persistent observation
Among all the strategies for establishing trustworthiness, prolonged engagement and persistent observation are the most fundamental – and the most demanding. They are the bedrock practices that make credibility possible in field-based qualitative research.
Prolonged engagement
Prolonged engagement is the investment of sufficient time collecting data to have an in-depth understanding of the culture, language, or views of the people or group under study; to test for misinformation; and to ensure saturation of important categories. It is not simply about spending time – it is about spending enough time to move beyond surface impressions and detect distortions in the data.
Crucially, prolonged engagement is also what builds trust. Through prolonged engagement with participants, researchers can gradually transform from an “outsider” to more of an “insider” whom participants are comfortable to initiate conversation with – spurring them to provide richer data through numerous revelations. Without that trust, participants may give socially acceptable or diplomatic answers rather than honest ones, fundamentally compromising the quality of the data.
Prolonged engagement is particularly suited to longitudinal or ethnographic studies. It should only be utilised in studies that require the researcher to spend sufficient time – months or even years – in developing trust and a strong relationship with participants in that setting. It is not a blanket requirement for every qualitative study, but where the depth of social context matters, it is indispensable.
Persistent observation
While prolonged engagement provides scope – breadth of understanding across a context – persistent observation provides depth. As Lincoln and Guba (1985) famously put it: “If prolonged engagement provides scope, persistent observation provides depth.”
A key characteristic of persistent observation is that it enables the investigator to observe participants’ reactions, behaviour, facial expressions, and changes in the vocal spectrum – phenomena that cannot be comprehensively observed in telephone interviews or other forms of indirect communication. It brings the researcher into close, sustained contact with the specific aspects of a situation that are most relevant to the research question, filtering out what is irrelevant and focusing in on what is salient.
The relevance of persistent observation lies in the “in-depth pursuit” of those elements identified as salient through prolonged engagement. In other words, the two strategies work in tandem: prolonged engagement identifies what matters, and persistent observation examines it carefully.
In practice, persistent observation might look like a researcher constantly rereading interview transcripts, refining codes, and revising their conceptual categories across successive rounds of analysis – tracking how their understanding evolves and deepens over time. In one grounded theory study, researchers constantly read and reread data, analysed and theorised about it, and revised concepts accordingly – recoding and relabelling until the final theory provided the intended depth of insight. That iterative immersion is precisely what persistent observation looks like in analytic practice.
Other key strategies for credibility
Beyond prolonged engagement and persistent observation, several additional techniques support credibility. Used together, they create a multi-layered defence against the distortions that can arise in qualitative work.
Triangulation involves using multiple data sources, methods, or researchers to cross-check and corroborate findings. This may be done through data, investigator, or theoretical triangulation, strengthening the credibility of the qualitative data by ensuring appropriate multiple perspectives throughout data collection.
Member checking involves returning findings to participants so they can verify whether the researcher’s interpretations accurately reflect their experiences. Researchers tell respondents what they have heard so far – generating debate or confirmation – and may re-contact them later to check findings and ensure accurate understanding. This is often considered the single most powerful check on credibility available to qualitative researchers.
Peer debriefing requires the researcher to expose their work to a critical colleague who asks hard questions about the methodology and interpretations. Peer review and audit trails are commonly used standards of rigor to ensure trustworthiness and integrity within the data analysis process.
Negative case analysis involves actively searching for data that contradicts the emerging interpretation and revising that interpretation until it accounts for the exception. This disciplines the researcher against selectively confirming their initial hypotheses.
Linking method-specific criteria to research design
One of the important implications of Lincoln and Guba’s framework is that these criteria are not one-size-fits-all. Not all strategies might be suitable for every study – for example, a member check of written findings might not be possible for study participants with a low level of literacy. Researchers must think carefully about which trustworthiness strategies are appropriate given their specific method, population, and context.
This is what makes Lincoln and Guba’s framework genuinely method-appropriate rather than simply a qualitative substitute for quantitative norms. The nature and purpose of the quantitative and qualitative traditions are different enough that it is erroneous to apply the same criteria of worthiness or merit to both. Trustworthiness criteria must be selected and applied in ways that are sensitive to the research design – and built into that design from the start, not bolted on at the end.
A research question must be clear and focused and supported by a strong conceptual framework, both of which contribute to the selection of appropriate research methods that enhance trustworthiness and minimise researcher bias inherent in qualitative methodologies. Trustworthiness, in other words, is not something that is achieved after the study is complete – it is designed in from the beginning.
Why this matters beyond the academy
Some might wonder why these methodological debates matter outside university seminar rooms. The answer is straightforward: trustworthy qualitative research findings are important for informing policy decisions and improving the provision of services in various fields. When researchers study marginalised communities, patient experiences, educational practices, or social welfare systems, the quality of their methods directly shapes the quality of the interventions that follow.
Research that lacks credibility – because the researcher spent insufficient time in the field, failed to account for their own biases, or drew conclusions without checking them against participants’ perspectives – can produce findings that are misleading at best and actively harmful at worst. If a researcher or research team disseminating a qualitative investigation does not fully demonstrate that the work is trustworthy, it is up to the consumer to follow the age-old practice of “caveat emptor”. In high-stakes contexts, that is not good enough.
Method-appropriate evaluation criteria – centred on trustworthiness – are therefore not an academic nicety. They are the foundation of qualitative research that can be relied upon, built upon, and acted upon.
What do you think? If prolonged engagement and persistent observation are so central to credible qualitative research, how should researchers balance these demands against the practical constraints of time and access? And when findings from a qualitative study are used to inform public policy, who bears responsibility for evaluating whether the research meets adequate standards of trustworthiness – the researchers themselves, journal editors, or policymakers?
References
- https://pmc.ncbi.nlm.nih.gov/articles/PMC4535087/
- https://idcinternationaljournal.com/0ct-2019/Article_6_manuscript_IDC_August_October_2019_Dr_Annie_full.pdf
- https://www.simplypsychology.org/trustworthiness-of-qualitative-data.html
- https://pmc.ncbi.nlm.nih.gov/articles/PMC8816392/
- https://researchdesignreview.com/2019/07/31/critical-thinking-qualitative-design/
- https://resources.nu.edu/c.php?g=1013606&p=8394398
- https://www.emerald.com/qrj/article/23/4/372/359623/Application-of-Guba-and-Lincoln-s-parallel
- https://nursekey.com/trustworthiness-and-integrity-in-qualitative-research/
- https://nsuworks.nova.edu/cgi/viewcontent.cgi?article=5845&context=tqr
- https://www.quantilope.com/resources/glossary-trustworthiness-in-qualitative-research
- https://pmc.ncbi.nlm.nih.gov/articles/PMC7055404/
- https://journals.lww.com/dccnjournal/fulltext/2017/07000/rigor_or_reliability_and_validity_in_qualitative.6.aspx
- https://www.sciencedirect.com/science/article/pii/S2949916X24000045
- https://files.eric.ed.gov/fulltext/EJ1320570.pdf
Leave a Reply