Subjective scales lose discriminating power at very low or very high load, showing ceiling and floor effects
Aliases: ceiling effect · floor effect · scale sensitivity
What it is
Subjective workload scales are sensitive to differences within a moderate load range, but when true load is low enough to feel essentially effortless, or high enough that users are clearly overwhelmed, ratings cluster at the low or high end of the scale and stop distinguishing finer degrees of difference. These two patterns are called the floor effect and the ceiling effect.
Why it happens
A scale's range and granularity are fixed, and people's own ability to discriminate load differences also declines at the extremes. At very low load, several tasks that all feel "effortless" no longer differ in subjective experience, so ratings pile up at the scale's lowest few points. At very high load, users often shift from analyzing specifically what's hard to a generalized sense of "I can't keep up," and tend to just pick the top of the scale, no longer distinguishing further increases in load. This is a joint product of the scale's own resolution and how people rate under extreme states — it isn't something a more finely graded scale alone can fully fix.
Studying it
A common way to check for ceiling or floor effects is to see whether ratings under a given condition pile up heavily at either end of the scale, and check whether the variance in that condition is noticeably smaller than under moderate-difficulty conditions. If a difficulty gradient was built into the design and the hardest and second-hardest conditions show almost no difference in subjective rating while objective performance differs clearly between them, that's a classic signal that the scale has lost discriminating power at the high end — an unchanged subjective rating alone cannot be used to conclude the two designs impose comparable load.
Where it stops holding
This problem is most acute when comparing multiple designs that are all "already quite hard" — the scale provides limited information in that regime. For comparisons within the moderate, everyday range of operation, the scale's discriminating power is generally reliable and doesn't need to be second-guessed.
Applying it
- When assessing a scenario known to be high-load (an interface juggling several simultaneous alerts, say), don't conclude two designs are "equally bad" just because their subjective ratings both sit near the top of the scale — supplement with objective performance metrics (error rate, reaction time) to tell them apart.
- For fine-grained optimization comparisons in an everyday, low-load scenario, the subjective scale has limited discriminating power — lean on behavioral indicators rather than differences in reported scores.
Related
- Same group: A9.07.1 Multidimensional scales split load into mental, physical, temporal and other components, scored separately then weighted · A9.07.2 A single overall scale is simpler to administer but cannot reveal where load comes from · A9.07.3 Retrospective ratings taken after a task are dominated by the peak and the ending, diverging from the true average
- Adjacent: A9.02 Measuring Load
- Search terms:
ceiling effect·floor effect·subjective rating scale sensitivity
Cards in the same group
- A9.07.1Multidimensional scales split load into mental, physical, temporal and other components, scored separately then weighted
- A9.07.2A single overall scale is simpler to administer but cannot reveal where load comes from
- A9.07.3Retrospective ratings taken after a task are dominated by the peak and the ending, diverging from the true average