Q3.02.3Standardized SUS item wording and orderdesignresearch

Use the standard items in the standard order

Aliases: standard SUS order · do not rewrite SUS · fixed scoring rule

What it is

Cross-product comparability for SUS hangs on standard items, a fixed order, and odd-positive / even-negative scoring. Ten items appear in a set sequence, odd items positively keyed, even items reversed, with a five-point agreement grid, composing 0–100. Rewrite stems, shuffle order, drop items, switch to seven points, or drop reversals, and the number may still be called a usability score—it is no longer SUS, and it cannot share a ruler with history or with external percentiles. Replacing “system” with the product under evaluation (this app, this website) is an object substitution, not a rewrite of the construct or a reorder of the items.

Why it happens

Order supplies context: asking willingness to use frequently before complexity primes a different referent than the reverse. Reverse-keyed items sit in fixed slots; they share the load of acquiescence and make careless answering more visible as contradictions between polarities. Remove them or make every item positive, and acquiescence lifts the total while contradiction checks vanish. Drop any layer and the linear map—0–4 contributions times 2.5—breaks, which would need a new norm. Once a stem is rewritten as a feature (“is search easy”), the item leaves overall impression and becomes a feature rating; the total’s meaning moves with it. The protocol looks rigid; what it freezes is the ruler, not the product.

Studying it

Compare a rewritten version with the standard version in a split sample or successive administration, looking at total correlation, mean shift, and positive–negative consistency. High correlation is not permission to swap rulers: a constant offset can keep ranks stable while the absolute value can no longer be mixed with standard SUS. Record every departure (object word, order, number of points, missing-data rule). Publications and archives should state whether the standard ten or a variant was used; variants need a different name so the literature does not treat distinct instruments as one score.

Where it stops holding

Independently validated translations can count as the standard in that language, and should still be archived apart from the source version rather than stacked onto English norms. Disabled or low-literacy samples may need read-aloud or assistance; the assistance protocol must be fixed so moderator paraphrase does not become a private rewrite. A single “overall usability” item is not a shortened SUS. An in-house custom scale can fit the product better; that is a different ruler and should not inherit SUS percentiles.

Applying it

  • Lock the full standard ten items and their order in the item bank. Substitute only the object word with the product name; do not rewrite predicates.
  • Do not insert, delete, or reorder “to fit the business.” Extra questions go after the whole SUS block and are analyzed separately.
  • Encode odd-positive / even-negative scoring as a test: shuffled polarity or a wrong item count should error, not emit a 0–100.
  • Before a historical comparison, check an item-text hash or version mark. On mismatch, break the trend; do not splice a variant onto the old curve.

Related

  • Same group: Q3.02.1 SUS supplies one overall score that can be compared across products · Q3.02.2 The score does not locate specific problems
  • Adjacent: Q3.01 Questionnaire design · Q3.15 Questionnaires and rating scales
  • Search terms: standard SUS · item order · instrument integrity

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/Q3.02.3