Desk Research: Document and content analysis

Discovery / ExplorationallBeginner

TL;DR

Systematically analyse material that already exists and that nobody produced for your study: support tickets, reviews, forum threads, transcripts and internal documents.

Detailed description

Document and content analysis is secondary research on material that exists independently of the study: support tickets, store reviews, forum threads, sales transcripts or internal documents. Its defining property is that the data is non-reactive, because it was generated without anyone knowing it would be analysed: that removes the observer effect, but moves the problem to the origin of the corpus (Leavy, 2017, p. 257). It should not be confused with a literature review, which examines what others researched; here you examine traces of what people did or said for their own reasons.

The handbook by Schneijderberg, Wieczorek and Steinhardt defines it as the systematic, rule-based collection and analysis of texts that interprets their manifest and latent meaning by breaking them down into categories, patterns or topics (2026, p. 26). The manifest is what the text says; the latent has to be read between the lines, and that takes interpretation, not just counting.

Rose (2001, pp. 56-65) sets out the procedure in four steps. First, choose the material and the sample, which can be random, stratified, systematic or cluster-based; if you take one piece every so often, make sure the interval does not line up with a rhythm in the material itself: taking one day in seven always lands you on the same weekday, and the sample fills up with whatever happens that day. Second, define coding categories that are exhaustive, mutually exclusive and analytically useful, and pilot them until they no longer overlap. Third, code: for the result to be replicable, two people code independently and you measure how far they agree. Fourth, analyse: count frequencies, cross categories and group codes into themes. Rose applies it to images, but the method was born for written and spoken texts (p. 54).

Main objective

Extract patterns from a corpus that already exists, so you start from evidence rather than assumptions.

Use cases

Customer supportE-commerce with reviewsCommunities and forumsProducts with prior research history

When to use it

Early on, when material has already accumulated and budget or timeline do not allow a primary study. Before using it, pin down two things: the period the corpus covers and who it was produced about.

Effort level

Low to Medium

Recommended number of users

Not directly applicable (document corpus)

Advantages

  • The data already exists: No recruiting, scheduling or moderating, so cost and timeline are a fraction of a primary study.
  • Non-reactive: The material was produced with no relation to your research, so nobody is telling you what they think you want to hear.
  • Volume: A corpus of tickets or reviews covers more cases than you could interview, and lets you count frequencies as well as read.
  • It always yields something: Even when it does not confirm what you were after, it narrows the ground; knowing what not to do is a usable result.

Disadvantages

  • Unknown currency: The material is dated, and a two-year-old complaint may describe a problem already fixed. Without pinning the period, the corpus mixes eras.
  • Someone else's population: The data was produced about other people and for another purpose, so the sample it carries belongs to whoever generated it, not to your question.
  • Who-writes bias: Only those who bothered to write leave a record, usually the very angry or the very happy; the silent user does not appear.
  • You cannot follow up: The corpus answers what it answers. If a pattern stays ambiguous, you need a primary method to close it.
  • Counting does not weigh cases: Frequency tells you where to look, but it does not tell a clear case from a doubtful one within the same code, and an infrequent theme can matter just as much. Read the rare cases before discarding them.

When to use

  • Research start
  • Tight budget or timeline
  • Prioritising reported problems
  • Preparing a primary study

Metrics

  • Volume by category
  • Frequency of recurring themes
  • Period covered by the corpus
  • Share of material discarded
  • Inter-coder agreement

Practical example

Review twelve months of support tickets and store reviews to rank problems by frequency before deciding what to test with users.

Related methodologies

Related Resources

Free tool by UXR — UX Research Consulting in Chile

Last updated: