Learn · Qualitative methods
What LLMs can do in qualitative research, and where they stop
An LLM can summarise, compare, and suggest language. It has no relationship with the people whose words it is reading, and it cannot carry responsibility for an interpretation. The question is which role the model should play, not whether it can produce themes.
After this guide you will be able to:
- Explain why an LLM cannot replace the interpretive work.
- Use one for the stages where it saves time.
- Name the five ways researchers over-trust these tools.
- Set a division of labour where the model suggests and you decide.
Why a model cannot replace the researcher
A large language model generates text by learning statistical relationships between words and predicting the most probable next one. That gives you fluency, not understanding, and there is no built-in mechanism to check whether an answer is true. Because it returns the most likely patterns from its training data, its default output is the conventional reading, and even careful prompting produces outputs that vary between runs and can include fabricated content.
Qualitative analysis is interpretive and contextual. It rests on embodied judgment and reflexivity, on a researcher who was in the room. A tool built on pattern-matching cannot stand in for that. Practitioners who have tested these models say the same thing: they produce passable surface-level analyses quickly, but rarely at the level of an experienced researcher. At best they speed up routine tasks.
Where LLMs help
Used deliberately, a model supports the early, labour-intensive stages. In published comparisons, models like GPT-4 and Claude pull coherent themes from large sets of narratives and catch many of the same core ideas a human team finds — a 2026 scoping review counted 75 such studies, most using models for coding and theme identification. They organise complex text fast and draft preliminary themes, which saves time.
Different models have different habits: some condense themes into broad buckets, others over-segment. The approach that works treats the model as a junior analyst. You review its themes, merge the overlapping ones, and hold the line on conceptual clarity. Role-specific prompting and a clear analysis plan produce richer output, and sometimes the model surfaces a reading that challenges your assumptions and sends you back to the data.
The ethics, in one line
A model can produce plausible errors and give different answers to the same prompt. Its training data may be undisclosed, and a provider may retain or reuse what you submit. Qualitative data often contain sensitive narratives, so check consent, ethics approval, institutional agreements, and the provider's data policy before uploading them. A local model can reduce exposure, but it does not replace documentation or human judgment. The full ethics guide covers this in depth.
The five ways researchers over-trust AI
Five epistemic risks are worth knowing by name, because naming one is how you catch yourself doing it.
- Category error. Mistaking the model's fluent language for genuine analytic insight.
- Unreliable outputs. The same prompt producing inconsistent, irreproducible results.
- Anthropomorphic fallacy. Attributing understanding or collaboration to a system that is repeating linguistic patterns.
- Causal misattribution. Blaming your prompts for a failure that is really a structural limit of the model.
- The oracle effect. Trusting an output as objective or neutral because it came from a machine.
The failure modes follow from the mechanism: a model has no reflexive capacity, over-weights frequent expressions, and can miss subtle or minority perspectives. Results shift with the prompt, the model, and the token limit. So check AI output against the raw data before it earns a place in the analysis.
A workable division of labour
The division below keeps two things straight: the role the model was given, and who stayed accountable.
| Research task | Suitable LLM role | Human responsibility |
|---|---|---|
| Research question | challenge breadth and ambiguity | decide the aim |
| Interview guide | flag leading or missing questions | judge relational and ethical fit |
| Transcription | draft the transcript | check against the audio |
| De-identification | flag possible identifiers | make the disclosure-risk call |
| Coding | suggest labels, or apply a fixed codebook | define, revise, and interpret codes |
| Theme development | propose alternatives and counterexamples | build the analytic account |
| Negative-case analysis | retrieve disconfirming passages | decide what they mean |
| Writing | structure and language editing | own every claim |
| Final interpretation | no autonomous role | researcher, with participants |
The rule underneath the table. A model may participate in the workflow, but the weight of each output has to be decided in advance: mechanical support, descriptive assistance, analytic challenge, or interpretive claim. Only the first three are safe to hand over. Interpretation, theory, and meaning stay human. The model can contribute. It cannot carry authorship, responsibility, or the relationship with the participant.
Keep the human in the loop, on the record.
You want the model's speed on coding and theme-spotting, but you also need to show a reviewer that you, not the model, made the analytic decisions. The Vahtian tools are built for that. MethodVahti checks that your method, terminology, and claims stay aligned, and flags combinations that do not belong together. QualiVahti Local runs the model as a suggestion engine you review, on your own machine, with every decision logged. The model suggests; you decide the meaning.
See MethodVahtiRelated guides
See also AI qualitative coding, human-reviewed, using LLMs on interview data ethically, and a short history of qualitative research. More guides are on the Learn page.