
When Your Themes Don't Answer Anything
A PhD student showed me the research questions from his progress report. There were five of them:
- How does AI enable customer experience in retail?
- How does AI improve operational efficiency in retail?
- How does operational efficiency influence customer experience?
- Which AI capabilities contribute most to customer satisfaction?
- How can retailers strategically leverage AI to improve both customer experience and operational performance?
He wasn't stuck for lack of effort. He had read widely, he had access to data, and he had an engaged supervisor. What he said was: I have an idea but I don't know how to formalise it and point out my contribution. And separately: I don't know what theoretical framework, method and dataset I need to finish it.
His supervisor's advice was to read more and collect more articles.
I've heard versions of this from bachelor's students through to PhD candidates, and I've come to think the reading advice is usually aimed at the wrong problem. More literature does not fix an unbounded question. It makes it worse, because every additional paper supplies more territory the question could plausibly cover, and the question has no principle for excluding any of it. You read for six months and end up with a broader problem than you started with, and a stronger feeling that you ought to be able to handle it.
Look at what those five questions actually do.
The question that can't be wrong
How does AI enable customer experience in retail?
Ask what a finding would have to look like for the answer to be no. There isn't one. The verb enable has already settled it; the question presupposes that AI enables customer experience and asks only for the mechanism. Whatever the data shows will be written up as a way AI enables customer experience.
The same is true of RQ2, where improve does the same work.
Compare RQ3: How does operational efficiency influence customer experience? That one is a real question. Efficiency might not influence experience at all. It might degrade it; faster checkout with fewer staff is a plausible route to a worse experience. Because it could come out false, it can also come out true, and either result is a finding.
Of the five, RQ3 is the only one that could survive contact with disconfirming data. It was also the one he was least attached to.
This is not an argument for hypothesis testing. Interpretive work can be disconfirmable too. The question just has to be about something specific enough that the data is able to push back.
The question that is several questions
RQ5, How can retailers strategically leverage AI to improve both customer experience and operational performance?, is not a fifth question. It's RQ1 through RQ4 restated in a single sentence, with strategically added.
The word both is the tell. Two outcomes, one question, and no account of what happens when they trade off, which is the interesting case, and the one the question is structured to avoid.
The test is mechanical. Count verb-object pairs. Each additional pair is another study, needing its own design and possibly its own data. The five here are really one causal chain (RQ1→RQ2→RQ3), one ranking problem (RQ4), and one summary (RQ5). The ranking problem alone requires a different method from everything around it — contribute most implies comparison across a defined set, which is not something the same material that answers a how question will give you.
That mismatch is why he couldn't name his method. It wasn't missing. His questions required three incompatible ones.
The term nobody defined
AI capabilities. Customer experience. Strategically leverage.
Not one of these was pinned down anywhere in the report. Until they are, no framework can be chosen, because a framework is precisely a commitment about what these things are and how they relate.
This is the pattern that survives contact with good software. Traceability will tell you which passage produced which theme. It cannot tell you that you were applying a slightly different sense of engagement in transcript 4 than in transcript 27, because that inconsistency never appears in the data. It lives in the judgments you made while reading, and you will not remember making them.
The fix is smaller than it looks. Write down what counts as an instance and what doesn't. Include two borderline cases and decide them before you meet them in the data. This isn't operationalising in a positivist sense — it's making your own criteria visible to yourself, which is what lets you apply them consistently and report them honestly.
Why it surfaces late
None of this is visible at proposal stage. A broad question reads as open and appropriately inductive, and no review panel objects to it. Fieldwork feels fine too, because nothing can be off-topic when the question never made anything off-topic.
The bill arrives when you have to say something specific. For that PhD student it arrived at what is my contribution. For a qualitative researcher it usually arrives at analysis — the first moment the question has to do real work, deciding which data matters and which doesn't. A question that never drew that line cannot draw it now, and no analytical software can draw it for you. It will faithfully organise everything you gave it, because you never supplied grounds to treat any of it as more relevant than the rest.
Then you sit looking at a clean set of themes for two weeks, unable to turn them into a finding.
If you're already in the data
Most people reading this have transcripts open, not a blank proposal. All three problems are repairable mid-project, and repairing them is faster than continuing to analyse around them.
Take an afternoon before you go further.
Write the disconfirming case. In one sentence: what would a participant have to say for your answer to be no? If nothing qualifies, narrow until something does. Check your verbs while you're there — enable, improve, enhance and support all quietly presuppose the finding.
Split the compound. Count verb-object pairs, pick one as primary, and say explicitly that the rest are secondary. In your notes now, in your limitations later.
Pin down one term. Take the central concept. Write what counts, what doesn't, and how you're deciding two cases at the edge.
You will usually find your existing data supports a sharper question than the one you set out with. The material is generally fine. What's missing is the frame that makes some of it evidence.
Then go back to the analysis. The themes will look different — not because they changed, but because you can finally see which ones are load-bearing.
A note on where I'm coming from
I build a platform for this stage — the structuring that happens before data collection, and the funding proposals that have to justify it. So I have an obvious interest in claiming it matters.
Two qualifications, then.
The first: this isn't fast. We have described parts of our own workflow in terms of minutes, and Susanne pushed back on that when we spoke. Her view was that a realistic timeline runs to a few hours once you account for scattered source material and the time it takes to read and judge the output. She's right, and we're changing how we describe it. Sharpening a research question is thinking work. Software can structure the thinking and hold the pieces in view. It cannot do the deciding, and anything claiming otherwise is selling you the part that was never the bottleneck.
The second: none of the three checks above needs software. They need a blank page and an uninterrupted afternoon. If you do them and never touch anything I've built, this post has done its job.
Where a platform earns its place is in the accounting — keeping the question, the sources and the reasoning in one traceable place, so that six months later you can reconstruct why you narrowed the way you did. That's the same discipline QInsights applies to analysis, applied one stage earlier. The two halves are continuous: a clear question makes analysis tractable, and analysis that stays linked to its evidence makes the question defensible when somebody asks.
Surya Yadav is co-founder of Aveksana, a platform for structuring research questions and preparing grant proposals — the stage before data collection begins.
