Q3.02.1SUS as a cross-product comparable scoredesignresearch

SUS supplies one overall score that can be compared across products

Aliases: SUS total score · cross-product usability score · 0–100 usability

What it is

The System Usability Scale (SUS) takes ten overall evaluations of the same system and, by a fixed recipe, composes a 0–100 score. Its job is a cross-product comparable index of perceived usability, not a description of which step failed inside the product. Odd items are positively keyed, even items reversed; each item is folded to a 0–4 contribution and the sum is multiplied by 2.5 so the result shares a hundred-point scale. Teams use it to compare versions, competitors, or channels, provided the instrument stays the same items and scoring.

Why it happens

A single total can be compared because the item set, polarity pattern, and linear transform are frozen, so answers on different products map onto one ruler. The ruler measures post-session overall impression—an amalgam of learning, efficiency, confidence, and willingness to use again—not a physical unit. Task difficulty, brand, and a recent frustration raise or lower the whole score, so “comparable” means relative position under similar tasks and similar samples, not that 72 equals the same usability in every setting. Reverse-keyed items dampen yea-saying so the total is less easily flattened by acquiescence; they do not turn the ruler into a diagnostic.

Studying it

Administer immediately after a representative task set, recording the tasks, the product, and the sample. When comparing products, covary task difficulty and familiarity, or lock the same task battery. Report the distribution and sample size with the mean; do not treat a small-sample point estimate as a product ranking. Scores are often read against published percentile tables for relative standing; those tables come from particular compilations and are not a norm for one’s own product. Missing items follow this instrument’s established rule, not some other scale’s imputation.

Where it stops holding

Switching the task from “find one control” to “open an account” can drop the score from the task alone, not from a worse interface. Different language versions, and object substitutions that replace “system” with an unrelated service, weaken cross-study comparability. Expert users and first-time users should not be pooled into one product score. SUS also does not replace completion or error rates: a high-scoring product can still fail a critical task. Filled from a home-page impression with no task, the score measures appeal, not post-use usability.

Applying it

  • To compare two products or versions, use the same SUS, the same task battery, and a like sample, then read the score gap.
  • Place the score beside task completion and errors; do not publish SUS alone as the usability conclusion.
  • Externally report the 0–100 total plus sample and task notes. If you map onto external percentiles, name the compilation; do not treat the percentile as your measurement.
  • Recheck with a small retest: on the same product and tasks the score should be roughly stable; if it jumps, inspect task and sample before claiming a product difference.

Related

  • Same group: Q3.02.2 The score does not locate specific problems · Q3.02.3 Use the standard items in the standard order
  • Adjacent: Q3.06 Task success rate · Q2.06 Usability testing
  • Search terms: System Usability Scale · perceived usability · cross-product comparison

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q3.02.1