Blog
419 Qualitative Researchers Say Never. Others Say Adapt. Here Is a Third Answer.
- August 20, 2026
- Posted by: Dr. Kisito Futonge
- Category: Research

Using Large Language Models for Qualitative Data Analysis
In 2025, more than four hundred experienced qualitative researchers from over thirty countries signed an open letter rejecting the use of generative AI in reflexive qualitative research.
Not merely restricting it. Rejecting it—at every stage, including the first pass of coding.
Their argument deserves to be heard at full strength. A language model predicts text; it does not understand it. Qualitative analysis is an act of meaning-making, performed by someone, from somewhere, with a stake in the interpretation and an obligation to the people whose words are being interpreted. And the industry behind these tools depends on forms of extraction: of data, of energy, and of poorly paid human labour.
Refusal, the letter argues, is the position with integrity.
The response, when it came, landed punches of its own. Nobody serious needs to claim that a model “makes meaning.” Humans make meaning with tools, as they always have—with memos, whiteboards, software, colleagues, theories, and methods. Some of the intellectual traditions qualitative research itself values, including distributed cognition and sociomateriality, begin from the premise that thinking has never been a purely solitary activity. And while the harms surrounding generative AI are real, they do not automatically settle the question of individual abstention. One researcher refusing a tool may do less to change those systems than regulation, institutional standards, and collective pressure.
I have read both texts more times than I can count because I research this area. What strikes me most is that each side is weakest precisely where it attacks a position the other side does not actually hold.
The open letter demolishes the naive idea that a language model genuinely understands participants’ accounts. But the careful researchers I know who use these tools believe no such thing.
The response, meanwhile, defends the researcher’s control over how a tool is used in an individual study. But the letter’s deepest objections were never confined to individual practice. They concern the system in which the practice takes place.
Two ships, different oceans.
The debate has produced plenty of position-taking. It has produced far less procedure.
Here is the question I think matters more than Should you use it?
If a model quietly damaged your analysis, would you know?
Not in principle. Concretely.
What check would catch a fabricated quotation before it appeared in your findings chapter?
What figure would show that a model handles your descriptive codes reasonably well but performs disastrously on the interpretive ones?
What record could demonstrate—to an examiner, a reviewer, or simply to yourself—that your interpretation existed before the model offered you a smoother and more convenient one?
Most researchers on both sides of the debate cannot answer those questions, because the debate does not require them to.
Positions are cheap. Procedures have to be built.
That is why I created a course on using large language models in qualitative data analysis.
Across seven modules, learners build the evidence themselves using a purpose-built synthetic interview corpus, so no real participant data is ever placed at risk.
You try to make a model fabricate quotations—if the current generation will still oblige—and then build a process capable of catching the fabrication.
You code a transcript by hand first and date the memo, because that memo becomes the instrument on which every later comparison depends.
You put the model on a leash you designed: a fixed codebook, written inclusion and exclusion criteria, and outputs that can be checked rather than admired.
Then you measure performance per code, because an impressive headline agreement score can conceal failure on the one category that matters most.
You put the model’s themes on trial. Every quotation is searched back to the source. Every proposed theme is classified as grounded, plausible but generic, or unsupported.
And then comes the course’s most distinctive move.
Instead of treating a model’s conventionality only as a weakness, you use it as an instrument. Large language models tend to gravitate toward the readings a field has made most available. That very tendency can be used as a mirror: a way of surfacing an assumption in your own analysis that had become too familiar for you to see.
That protocol comes from a methodological proposal currently under peer review, and the course says so plainly. Promising and proven are different words, and a course about intellectual honesty should model the distinction.
By the end, learners produce two documents that the emerging norms of academic integrity increasingly demand.
The first is a disclosure statement generated from their own decision log.
The second is a one-page conditions note in which every delegated task is justified with evidence from their own exercises.
It ends with the sharpest question I know how to teach:
What evidence would show that your safeguards had failed?
Notice what the course does not do.
It does not tell you that the prohibitionists are wrong.
Several of their arguments survive every exercise intact. The bonus module gives their strongest case its full force, including the objections that no coding benchmark could ever answer because they were never really about coding: researcher formation, epistemic dependence, labour, environmental cost, institutional power, and the ethics of building scholarship on extractive infrastructures.
A learner who completes all seven modules and decides not to use these tools has used the course exactly as intended.
The final module even provides an affirmative non-use declaration, because in the current climate silence is no longer much of a statement.
What the course refuses is something narrower: any position, for or against, built primarily from vibes and a bibliography.
If you use these tools, you should be able to show your receipts.
If you refuse them, you should be able to state your grounds in a form that a sharp examiner cannot collapse with a single question. And “the machine cannot interpret” collapses faster than many people expect, because nobody needs the machine to be an interpreter for it to affect interpretation.
The course runs on two slogans.
Fluency is not analysis. A model’s polished, confident prose is a property of its training, not evidence about your transcripts.
Demand receipts. From the model, whose quotations and claims you verify. And from yourself, whose decisions and delegations you document.
The prohibitionists are right that something precious is at stake: the formation of researchers, the situatedness of interpretation, and the words of participants who trusted us with them.
The permissionists are right that these tools are here, that thinking has always been mediated by instruments, and that blanket refusal settles less than it sometimes appears to.
Both are also right that the other side, at its laziest, is maddening.
This course is for the space between those lazy positions: for researchers who want their eventual answer—yes, no, or only under specific conditions—to rest on evidence they generated themselves.
The seven modules are live now and free at BlendedIQ.ai.
Bring a browser, a spreadsheet application, access to any major AI chat tool, and, if you have them, your doubts.
The course was built to survive contact with skeptics.
It was built by one.
The course is free and self-paced, taking approximately 10–14 hours across seven modules plus a bonus philosophy seminar. A certificate is available on completion. No participant data is used or required at any point.