Q3.02.2SUS score is non-diagnosticdesignresearch

The score does not locate specific problems

Aliases: non-diagnostic SUS · total cannot localize · overall score versus problem list

What it is

A SUS total is an overall impression after the session. 58 versus 81 can say which is higher; it cannot say whether search, forms, or permissions produced the low score. The ten items themselves are global evaluations (complexity, consistency, willingness to use frequently), not a problem inventory tied to screens or steps. Treating the total as a diagnosis rewrites “felt hard to use” as a located defect. Location still comes from task behavior, error types, and places participants can point to.

Why it happens

Each item asks about the whole use; respondents crush many moments into one number. A low score can be generated by a failed critical task, light friction throughout, one sentence that broke trust, or a prior brand attitude. The composite then sums ten global numbers, and path information disappears further. Even a single item such as “unnecessarily complex” has no coordinates: complexity may sit in navigation, in terminology, or in error recovery. The instrument’s goal is a stable overall score; that goal trades off against “point to the place”—comparability requires dropping location.

Studying it

Place SUS beside task-level behavior: completion, errors, requests for help, abandon points. Check whether the low-score group shares one failure location; if sources are scattered, the total only says “something is wrong” and cannot direct which screen to change. Item–behavior correlations are usually weak and unstable; an item should not be treated as a proxy for a defect class. Severity ratings or problem maps should carry diagnosis; SUS stays a session-level summary. An experiment that breaks one known step and shows the total dropping while no item points at that step demonstrates the non-diagnostic property.

Where it stops holding

With a tiny sample and one or two obvious blockages, a low score and those blockages may line up in the cases at hand; that alignment does not generalize to “SUS locates.” When a product has a single core task, the total is closer to an evaluation of that task and still has no step coordinates. If an open comment box is collected with SUS, diagnosis comes from the comments, not from the score. Splitting the ten items into homemade “usability / learnability” subscales exceeds the instrument’s design claim and still does not confer location.

Applying it

  • In the report template, SUS appears only in the “overall perception” column. The “what to change” column must cite a task failure, an error, or a location on video.
  • A low score triggers inspection of behavioral material, not a debate about which SUS item.
  • Do not add or remove features from a single item (for example, adding push notifications because “I would use this frequently” scored low).
  • When checking a redesign, a rising score plus disappearance of the key failures is what verifies that location happened. A rising score with the same failure points means diagnosis never occurred.

Related

  • Same group: Q3.02.1 SUS supplies one overall score that can be compared across products · Q3.02.3 Use the standard items in the standard order
  • Adjacent: Q2.06 Usability testing · Q3.08 Error rate and help requests
  • Search terms: non-diagnostic score · SUS limitation · usability problem localization

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q3.02.2