Painting of two young peasant women in conversation in a field, one resting her chin on her hand listening to the other, who stands holding a staff
Camille Pissarro, <em>Two Young Peasant Women</em> (1891–92). Public domain (CC0), via The Metropolitan Museum of Art.

A paper says not to use AI to write your interview questions. I half agree

UX ResearchUX MethodologyAI
View contents

I spent the last few weeks rewriting the discovery methods module of my introductory UX Research course, and I ran into an uncomfortable gap: I was teaching people what makes an interview question good, and not how to arrive at one. Every piece of material I knew, mine included, settled the construction of the guide in a single line. "Create a guide with topics, not closed questions." Right. And then what?

Looking around, I came across an article by Elly Phillips and Fiona Holland, from the University of Derby, published this year in Qualitative Research in Psychology. It's called Designing semi-structured interview guides for interpretative phenomenological analysis and reflexive thematic analysis: a practical guide, it's Open Access under a CC BY license, and it opens by naming exactly the gap I had: there's plenty of literature on the what of questions and almost none on how you get to them.

I come from a different context than they do. The article is written for psychology students doing a thesis with interpretative phenomenological analysis or reflexive thematic analysis, not for someone with forty-five minutes with a participant and a team waiting on decisions. Even so, the seven-step process they propose was more useful to me than almost anything I've read about interviewing in years. I've already adapted it and put it into the course.

But there's one section that won't leave me alone, and that's what I want to write about.

The warning

The authors include a frequently asked question: could I use AI to create the questions? Their answer is no. They argue three things: that models hallucinate, that their training doesn't necessarily include recent academic literature on the topic, and — the one I find strongest — that using generative AI is methodologically incongruent with approaches that are explicitly reflexive. They cite Jowsey and colleagues, who in 2026 came out against it from within the same field.

When I read that, my first reaction was defensive. A few months ago I published an experiment where I handed a Gemini GEM a qualitative analysis of twenty-one interviews that had taken me two months to do by hand. I didn't conclude that AI shouldn't be used. I concluded something more uncomfortable and more specific: that coding can be assisted, that situated interpretation as far as I've seen remains human work, and that what gets lost along the way has a name, agency, which I learned afterwards reading Ugaya-Mazza and his team.

So here I have an article that goes further than I do. And instead of dismissing it, I found myself thinking about why, in this specific part of the process, I think they're considerably more right than they are in general.

The seven steps aren't for producing questions

This is the part that changed my mind, and it comes from looking closely at the process rather than at its output.

Step 1 is breaking the research question into about three topics. Step 2 is generating a great many questions without filtering. And in both, what the authors highlight isn't what you produce, but what you find out while doing it:

  • If the topics you pulled out can be answered by a single question, your research question is too narrow.
  • If you end up with eight topics, it's too broad and won't fit in one interview.
  • If nothing comes to mind in step 2, the problem probably isn't your creativity either: it's the question.

It's a diagnostic. The steps aren't a machine for producing questions, they're a procedure for discovering that your research question was badly framed, at the only point in the project where fixing it is still cheap.

And that's the problem with delegating it. If you ask a model for twelve questions for your research question, it will give them to you. Good ones, even: open, not leading, well written. What it won't do is tell you that your research question didn't have twelve questions in it, because nobody asked it that and because the very format of the request assumes the question was sound. The deliverable looks identical. What disappears is the finding.

It's the same shape as the detachment I described in the Gemini experiment post, but one step earlier in the process and with an important difference: there what was lost was the memory of the route, how I arrived at those codes, which served me later when defending the findings. Here what's lost is information about the study's design, and it's lost before collecting a single piece of data. An interview with the wrong questions doesn't get fixed in the analysis.

There's another step that can't be delegated either, for simpler reasons: step 7 is testing the guide, first by reading it out loud yourself and then with a pilot interview. None of that happens in a chat.

Where I don't follow them

Where I don't go along is in the jump from there to methodological incongruence as a general principle. Phillips and Holland write for two explicitly reflexive approaches, in a context where the student's training is the point of the exercise: if a student delegates the design of their questions, they didn't learn to design questions, and that matters more than the guide they handed in. It's a solid pedagogical argument and it doesn't apply the same way to someone who has been running interviews for years and needs a guide by Thursday.

You can ask an AI to "write me a ten-question interview guide." I don't think the problem is that it writes them. The problem is who's leading.

You're the one leading. The one moderating, the one deciding what stays and what doesn't. If you hand that over, it doesn't matter how good the questions that come back look. You're the expert, and that part doesn't transfer.

It's the same mistake I made with Gemini, one step earlier in the process. There I assumed the transcripts were enough: I never gave it the client's context, or who would read the report, or what decisions would be made from the findings. The result was technically correct and strategically flat. The same thing happens here. With the research question alone, what comes back is ten well-written questions that don't know why they exist.

The context you have to give it

Leading is fairly concrete, and it starts before the first prompt. It's giving what we know and the model doesn't:

What the research is for and how the results will be used. Who will read them, what decision gets made with this. It comes first and it's what changes the output most.

Who these questions are for. Sometimes you already have an idea of the profile you need. If you don't, it has to be defined right there, before asking for anything. One, two, three profiles at most: that's what I recommend, because beyond that you can't cover it in the field or analyze it carefully afterwards.

What you want to know about each profile. And this from behavior, not from bias. Not "I want to confirm that older people struggle with payment," but "I want to understand how someone buying on the site for the first time works through payment."

There's something here I've seen around that strikes me as a mistake: writing exactly the same questions for every profile. They don't have to be. If you ask everyone the same questions to see how each one does, you're running a different study, one that compares people's performance against each other. It would be an interesting study, but it's a different one. What we usually want is to understand those people in that context: the use of a device, the purchase flow, why they decide to use or not use their headphones.

With all of that on the table, asking an AI for help seems reasonable to me.

The pass you don't skip

And afterwards it isn't "done, I have them."

What comes back gets read again through that same lens, question by question. Is it aligned with my research objective? Does it contribute or not? From there you edit the ones that need it, or you iterate, and then you can use them.

It's less flashy than "the AI wrote my guide" and it's more work than it looks. But it's the part that makes the guide yours, and it's also where the seven steps show up again: defining the objective, the profiles, and what you want to know about each one is, at bottom, doing steps 1 and 2 by hand.

The draft can be delegated. The diagnosis and the final decision are yours.

Which is more or less what Phillips and Holland are saying. Except they conclude that you'd therefore better not use it, and I conclude that you should therefore use it like this.

If you're just starting out

Before any of the above, there are two questions you can ask of each of your own questions from day one, without knowing any methodology.

Does this add information? Imagine the person has already answered. Does that answer help you understand your problem better? If you hesitate, it probably needs changing.

Am I fishing for a confirmation? The simplest example I can think of here is ice cream. "Do you like ice cream?" is a bad question, because the question itself already stakes out a position: you like it. "You don't like ice cream?" stakes out the opposite one and is just as bad. A better one would be "do you eat ice cream?", which asks about the behavior and not about the opinion I was already carrying. (Though who doesn't like it.)

They're the pocket version of nearly everything else, and they work the same whether you wrote the question or a model handed it back to you.

What I'm still turning over

I'm left with a doubt, and it's the kind that doesn't get resolved by writing a post.

Everything above works if you do the prior work. The risk is that this work is precisely the part you can skip without it showing: the guide that comes out of "write me ten questions" and the one that comes out of the whole process look fairly similar on paper. Neither carries a record of where it came from. And for someone just starting out, that's a problem the two questions above ease but don't solve: telling a considered guide from a plausible one takes practice, and practice isn't something a prompt gives you. Maybe Phillips and Holland are more right there than I gave them credit for.

The other thing I notice is that all of this has a name and I didn't learn it here. When I wrote about the Gemini experiment I ended up reading Ugaya-Mazza and his team, who study what happens to the sense of control, authorship, and decision of those of us doing qualitative research when we use computational tools. They call it agency, and one of their distinctions still strikes me as the most useful of all: transparency is not the same as agency. I can understand perfectly well how the model arrived at those ten questions and still not be the one who decided them. There I found it in the analysis. Here it applies just the same, before collecting a single piece of data.

In the meantime, the seven-step process adapted to UX is in module 4 of the introduction to UX Research course, free, with prompts and piloting included, which were the two things missing. And if you're earlier than that, still wrestling with the research question itself, I wrote a separate guide on how to formulate it.

AI can be a great help, but if you only play along with it, that's what you'll end up doing: following it, not directing it.

And the other thing, which holds with or without AI: you get plenty wrong at the beginning. So breathe, and make the best of what you have at hand.

I'd be glad to read corrections, especially from anyone who works with profiles and has reached a different conclusion about sharing questions across them.


Sources

  • Elly Phillips and Fiona Holland. Designing semi-structured interview guides for interpretative phenomenological analysis and reflexive thematic analysis: a practical guide. Qualitative Research in Psychology (2026). Open Access, CC BY 4.0. Link: https://doi.org/10.1080/14780887.2026.2719595
  • Luka Ugaya Mazza, Plinio Morita and James R. Wallace. The Shape of Agency: Designing for Personal Agency in Qualitative Data Analysis. University of Waterloo. Link: https://cs.uwaterloo.ca/~dvogel/gi2026/papers/1043a.pdf

Related Articles