Survey: SUS (System Usability Scale)

EvaluationquantitativeBeginner

TL;DR

The System Usability Scale (SUS) is a standardized survey instrument used to measure the perceived usability of a system. John Brooke published it in 1996 as a “quick and dirty” scale, meaning it is low-cost and easy to administer, but effective.

Detailed description

Despite its original description as a quick tool, the SUS has proven robust and reliable: in comparisons with other usability questionnaires it tends to deliver the most consistent results. Brooke built the scale on a precise idea of usability which, as he notes himself, is reflected in what was then the draft international standard ISO 9241-11: usability is not a property of the product but of its context of use, and the same system can score differently with a different kind of user or a different task. Of the three classes of measure that standard sets out (effectiveness, efficiency and satisfaction), the SUS covers the last.

The score is read against a benchmark, not in the abstract. The most widely cited reference average sits at around 68 (Sauro, 2011), calculated across a large number of usability studies. Bangor, Kortum and Miller (2008), for their part, measured an average of 70.1 across 2,324 surveys from their own corpus, and 69.7 when averaging by study: the two figures bracket the same order of magnitude, and which one you use matters less than what you compare against. Above that range the product's perceived usability is above average, and below it, below average. The comparison only holds if the reference study is comparable in tasks and in participant profile.

Calculating the score always follows the same procedure:

  1. For odd-numbered items (1, 3, 5, 7, 9), subtract 1 from the participant's response.
  2. For even-numbered items (2, 4, 6, 8, 10), subtract the participant's response from 5.
  3. Add the ten results: the total ranges from 0 to 40.
  4. Multiply that sum by 2.5 to bring it onto the 0-to-100 scale.

Main objective

The main goal of the SUS is to obtain a quantitative measure of the user's subjective satisfaction with the system or product.

Use cases

Redesigns with a previous version as a baselineComparison between versions of the same productBenchmark against the industry reference averagePrototypes and products already in productionNon-graphical interfaces: telephony and interactive voice response

When to use it

Immediately after the participant finishes the tasks, before any conversation about the session.

Effort level

Low

Recommended number of users

12-14 participants

Advantages

  • Fast to run: Ten questions and about two minutes to answer: it fits at the end of a session without making it longer.
  • One comparable number: It returns a single score from 0 to 100, comparable across products and across versions of the same product.
  • Efficient with small samples: It stays reasonably consistent with samples of around 12 to 14 participants (Tullis and Stetson, 2004).
  • Alternating statements: It alternates positive and negative statements, which makes it harder to answer on autopilot.
  • Free and already translated: It is free to use (with attribution) and has been translated and validated into several languages, so there is no instrument to build.

Disadvantages

  • Not diagnostic: It returns a global score: it tells you something is wrong, not what.
  • Not a percentage: A 68 does not mean 68% of tasks were completed, and that reading is the most common misunderstanding.
  • Measures reported perception: What a person says about the system does not always match the behavior that was observed.
  • Wide intervals with small samples: Differences of a few points between versions are not conclusive.
  • No substitute for qualitative analysis: The number does not explain its own cause.

When to use

  • At the close of a usability test, to put a number on the experience that was observed
  • To compare the perceived usability of two versions, or against a benchmark
  • To track the same screen over time and see whether an intervention moved the needle

Metrics

  • SUS score from 0 to 100
  • Number of participants in the measurement
  • Comparison against the industry reference average (around 68)
  • Acceptability range: not acceptable below 50, marginal between 50 and 70, acceptable above 70 (Bangor, Kortum and Miller, 2008)
  • Difference between versions, with its margin of error

Practical example

After completing five tasks in a payments portal, each participant answers the ten items and the group average is compared with that of the previous version.

Related methodologies

Related Resources

Free tool by UXR — UX Research Consulting in Chile

Last updated: