Qualitative sample size

Justify how many interviews to plan and assess saturation during fieldwork, using the model by Malterud, Siersma and Guassora (2016) and the method by Guest, Namey and Chen (2020).

Is your study quantitative? Use the quantitative sample size calculator.

Before fieldwork: justify your sample

This tool does not calculate how many interviews you need. It helps you reason about your sample size with the information power model by Malterud, Siersma and Guassora (2016): the more information relevant to your question each participant holds, the fewer participants you need. The authors themselves warn that their model is not a checklist to calculate N (p. 1756), and that its five items trade off against each other. That is why they are not added up here.

How will you collect your data?
Study aimHow broad is the aim of your study?

A broad aim covers a more extensive phenomenon and calls for more participants (p. 1754). For example, understanding how people recover a password is narrower than understanding how they manage their digital security in general.

Sample specificityHow close are your participants to the experience you are studying?

A sample is dense when participants hold exactly the experience you are interested in, and also vary within it. If you recruit whoever is available, with no criteria, specificity drops and you need more people (p. 1755).

Established theoryDoes your study build on an existing theory or framework?

A theoretical framework helps explain relations within the data, and with it a small sample can still add something new. A study starting from scratch has to build its own foundation, and usually needs more participants (p. 1755).

Quality of dialogueHow do you expect the interview conversations to go?

It depends on the interviewer's skills, how articulate the participant is, and the chemistry between them. Malterud and colleagues acknowledge it is hard to predict in advance (p. 1755).

Analysis strategyHow will you analyze the data?

Comparing patterns across participants, as in a thematic analysis, calls for more participants than analyzing a few accounts in depth (p. 1756).

The tool does not suggest a number. If you already have one, it goes into the justification paragraph.

Empirical reference. In Hennink and Kaiser's (2022) systematic review, most of 16 tests with interview data reached saturation between 9 and 17 interviews (mean of 12 to 13). Nearly all came from health research, with homogeneous populations and narrowly defined aims. Take it as a starting point, not as your study's number: the authors found little evidence on how much each study characteristic matters.

Justification paragraph

Answer the five questions to build the paragraph.

During fieldwork: assess saturation

Malterud and colleagues recommend revisiting sample size during fieldwork. This part helps you do it with the method by Guest, Namey and Chen (2020): after each interview you record how many new codes came up, and the tool calculates how much new information the latest interviews add compared with the first ones.

Base size

The first interviews, which serve as the point of comparison. Guest and colleagues tested base sizes of 4, 5 and 6 and the outcome barely changed, so they suggest 4 as the default.

Run length

How many consecutive interviews each assessment looks at. Runs overlap: they move forward one interview at a time. Runs of 3 give a more conservative assessment.

New information threshold

The share of new information you accept as a sign of saturation. 0% is stricter: it takes more interviews before saturation counts as reached.

New codes per interview

A new code is one that did not come up in any earlier interview. Enter interviews in the order you conducted them, and count codes at a single level of your codebook: the method was tested on single-tier codebooks. If you work with several levels, you can run one assessment per level.

You need at least 6 interviews for the first assessment: 4 for the base and 2 for the run. You have 0.

What this calculation does not tell you

  • Meeting the threshold does not guarantee that saturation was reached. The authors compare the threshold to a p-value: a transparent convention, not a proof.
  • It does not tell you whether the interviews you did not conduct would have added something important. According to the authors, that cannot be known without conducting them; what the evidence does show is that new information tends to decline and that the most common themes come up early, as long as the interview guide and participant profile stay consistent.
  • The method was tested with inductive thematic analysis on narrowly defined questions. Its use with other epistemological perspectives is untested.
  • Counting new codes assumes a codebook-based analysis. Braun and Clarke (2021) consider saturation generally coherent with that kind of thematic analysis, but not with reflexive thematic analysis, where meaning is generated by interpreting the data rather than extracted from it. If that is your approach, this calculation does not fit it; they suggest reasoning about the sample with concepts such as information power, from the previous part.
  • For a more conservative assessment, use runs of 3 or the 0% threshold.

References

  • Braun, V., & Clarke, V. (2021). To saturate or not to saturate? Questioning data saturation as a useful concept for thematic analysis and sample-size rationales. Qualitative Research in Sport, Exercise and Health, 13(2), 201–216. doi.org
  • Guest, G., Namey, E., & Chen, M. (2020). A simple method to assess and report thematic saturation in qualitative research. PLOS ONE, 15(5), e0232076. doi.org
  • Hennink, M., & Kaiser, B. N. (2022). Sample sizes for saturation in qualitative research: A systematic review of empirical tests. Social Science & Medicine, 292, 114523. doi.org
  • Malterud, K., Siersma, V. D., & Guassora, A. D. (2016). Sample size in qualitative interview studies: Guided by information power. Qualitative Health Research, 26(13), 1753–1760. doi.org
Last updated: