Learnability, recognition, and comfort cannot be collapsed into one “ease of use” score
Aliases: composite usability · Likert collapse · single ease-of-use item · aggregated gesture score
What it is
Stuffing learnability, recognition, and comfort into one Likert item—“overall, was it easy to use?”—yields a one-dimensional score, not a summary of three-dimensional evidence. A high score may mean the move was easy to guess, even though it is often misread in the field. It may mean the demo always succeeded, even though the shoulder is already burning. One-dimensional scores also fake comparability across versions: the model changed, the score rose, people infer the gesture got better, while teaching and posture sat still. When a decision is needed, the score cannot point to a lever.
Why it happens
When people give a total, they grab whatever is currently salient. Just taught, salience is “I learned it.” Just failed by the recognizer, salience is “it does not listen.” Just dropped an aching arm, salience is fatigue. Question order, whether the demo succeeded, and whether the task was short, slide the weight toward one dimension. Adding or averaging three subscales assumes they compensate: a point lost on recognition can be bought back with a point of learnability. Physically they do not. Misreads do not shrink because the metaphor is cute; ache does not vanish because the model is accurate. Composite scores are also allergic to missing data: an item never tested for fatigue, with recognition filling the total, is systematically high.
Studying it
If a questionnaire is required, keep at least three separate items and report their correlations, not only a sum. Moderate-to-low correlations are evidence against merging. Predict that total from objective logs (first-try success, confusion, arm drops) to see which objective class kidnaps the score. Version comparisons must put three columns side by side; a t-test on the total is not a substitute. Generic instruments such as SUS can be a coarse product thermometer. They cannot rank items inside a gesture vocabulary.
Where it stops holding
A tiny vocabulary, one setting, one cohort of practiced users, can make the three items correlate highly, so a composite looks harmless. Grow the vocabulary or change the room and the correlation fans out. When management wants one dashboard number, the pressure returns; show three independent bars rather than one star. Accessibility evaluation must not composite: a gesture some people cannot form can still pass on the mean. Using “ease of use” as a formative weekly trend is less harmful than as a ship gate, but the trend still has to decompose.
Applying it
- Drop an “overall gesture ease” total from review packs. Use three columns, each colored on its own red–amber–green scale.
- If the organization demands one number, attach the three raw columns and “which column is this week's shortfall.” Never send the number alone.
- In release notes, say whether teaching, the model, or posture moved, and which column that maps to. Do not narrate with the total.
Related
- Same group: C4.05.1 Learnability, recognition, and comfort are separate evidence · C4.05.3 Learnable is not reliable: intuitive moves can still be misclassified · C4.05.4 Low effort is not comfort: small moves can still load posture
- Adjacent: C4.14 Gesture vocabulary size limits · C1.15 Pointing device metrics and throughput
- Search:
composite usability·unidimensional score·gesture evaluation