A9.07.1Multidimensional subjective workload scaleresearchdesign

Multidimensional scales split load into mental, physical, temporal and other components, scored separately then weighted

Aliases: NASA-TLX · multidimensional workload scale · weighted workload index

What it is

A multidimensional subjective rating scale does not stop at asking how tiring an activity felt overall — it splits load into several qualitatively different components, commonly mental demand, physical demand, temporal demand, perceived performance, effort, and frustration, scored separately and then combined into an overall score. The NASA Task Load Index (NASA-TLX) is the most widely used instrument of this kind.

Why it happens

Splitting the score pays off because felt load is a composite of multiple sources: a single interaction can simultaneously involve "there's a lot to remember" (mental), "my hands can't keep up" (physical or temporal), and "I feel like I'm doing badly" (frustration) — sources that are qualitatively distinct. Asking for one overall number forces the respondent to silently weight these sources in their head before reporting a single figure, and that weighting process is opaque to the observer. Scoring each dimension separately first, then combining them with fixed or pairwise-derived individual weights, makes that implicit weighting explicit — the total score can then be traced back to which specific component is driving it.

Studying it

The typical procedure asks participants to score each dimension separately right after the task, usually on a 0–100 scale; some implementations add a preliminary pairwise-comparison round to derive individualized weights before computing the weighted total. The independent variable is usually task difficulty or interface condition; the dependent variables are the per-dimension scores and the weighted total. A methodological point worth watching: the dimensions are often not statistically independent — mental demand and frustration frequently rise together — and that covariation weakens the scale's power to diagnose a single source. A spike in one dimension alone should not be read as proof that load comes entirely from that source.

Where it stops holding

Splitting load into dimensions works best with participants who can introspect and distinguish different sources of pressure. When a task is complex enough that users themselves can't tell whether it's mental burden or time pressure, the boundaries between dimensions blur, and the split adds to the reporting burden rather than clarifying it. Multidimensional scales also take longer to complete than unidimensional ones, making them a poor fit for scenarios that need frequent in-task sampling.

Applying it

  • Use a multidimensional scale rather than a single overall rating when a usability test needs to diagnose where the load is coming from, not just whether it's high.
  • Once scores come in, focus first on whichever dimension stands out relative to the others before digging into the cause — a temporal-demand score clearly higher than the rest points toward checking the time limit or response speed, not simplifying the visual design.
  • Verification: score two versions of the same interface with the multidimensional scale and check whether the difference lands on the same dimension — this confirms the change actually addressed the source of load rather than shifting it elsewhere.

Related

  • Same group: A9.07.2 A single overall scale is simpler to administer but cannot reveal where load comes from · A9.07.3 Retrospective ratings taken after a task are dominated by the peak and the ending, diverging from the true average · A9.07.4 Subjective scales lose discriminating power at very low or very high load, showing ceiling and floor effects
  • Adjacent: A9.02 Measuring Load · A9.01 Three Types of Load
  • Search terms: NASA-TLX · multidimensional workload scale · subjective mental workload

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/A9.07.1