Desk Research: Document and content analysis
TL;DR
Systematically analyse material that already exists and that nobody produced for your study: support tickets, reviews, forum threads, transcripts and internal documents.
Detailed description
Document and content analysis is secondary research on material that exists independently of the study: support tickets, store reviews, forum threads, sales transcripts or internal documents. Its defining property is that the data is non-reactive, because it was generated without anyone knowing it would be analysed: that removes the observer effect, but moves the problem to the origin of the corpus (Leavy, 2017, p. 257). It should not be confused with a literature review, which examines what others researched; here you examine traces of what people did or said for their own reasons.
The handbook by Schneijderberg, Wieczorek and Steinhardt defines it as the systematic, rule-based collection and analysis of texts that interprets their manifest and latent meaning by breaking them down into categories, patterns or topics (2026, p. 26). The manifest is what the text says; the latent has to be read between the lines, and that takes interpretation, not just counting.
Rose (2001, pp. 56-65) sets out the procedure in four steps. First, choose the material and the sample, which can be random, stratified, systematic or cluster-based; if you take one piece every so often, make sure the interval does not line up with a rhythm in the material itself: taking one day in seven always lands you on the same weekday, and the sample fills up with whatever happens that day. Second, define coding categories that are exhaustive, mutually exclusive and analytically useful, and pilot them until they no longer overlap. Third, code: for the result to be replicable, two people code independently and you measure how far they agree. Fourth, analyse: count frequencies, cross categories and group codes into themes. Rose applies it to images, but the method was born for written and spoken texts (p. 54).
Main objective
Extract patterns from a corpus that already exists, so you start from evidence rather than assumptions.
Use cases
When to use it
Early on, when material has already accumulated and budget or timeline do not allow a primary study. Before using it, pin down two things: the period the corpus covers and who it was produced about.
Effort level
Low to MediumRecommended number of users
Not directly applicable (document corpus)Advantages
Disadvantages
When to use
Metrics
Practical example
Review twelve months of support tickets and store reviews to rank problems by frequency before deciding what to test with users.