The ranking comes from magnitude-judgment experiments, not design intuition
Aliases: magnitude judgment · ratio estimation · graphical perception experiment · log absolute error
What it is
An experimental measurement of graphical judgment asks participants to estimate a ratio, difference, order, or value from controlled graphics and compares the response with truth. Cleveland and McGill proposed an ordering from theory and prior evidence, then used ratio judgments and transformed absolute error to test selected elementary perceptual-task comparisons. It was not a vote among designers, nor a complete fixed ranking produced directly by one experiment. Experiment turns intuition into a refutable hypothesis, but one study does not automatically create a universal design law.
Why it happens
“What percentage is the smaller of the larger?” avoids reading hidden source numbers, yet the response still includes estimation, numeric expression, and individual strategy. The classic log absolute error transforms discrepancies to reduce skew and compare conditions. Contemporary analysis can add hierarchical models, resampling, or robust summaries for participant and item variation. Two encodings with similar mean error may differ in directional bias, tail failures, or time, so the research question determines the outcome set.
Studying it
Specify the estimand, stimulus generator, sample, exclusions, repetitions, and analysis before collection, then randomize or balance channels and true-value ranges within participant. Retain raw judgments and report the error formula, bias, distribution or interval, time, and individual variation—not a ranking alone. Qualification and verifiable trials can protect crowdsourced quality, but exclusions should be prespecified rather than selected for a desired result. A replication needs comparable classic conditions plus separate tests on target devices, readers, and novel marks.
Where it stops holding
Ratio estimation is not arithmetic-free “pure perception”; percentage literacy, response controls, and strategy also matter. A small convenience sample can expose a large defect but cannot establish a population order, and many repeated trials do not repair coverage bias. Experimental error is not automatically real-world decision loss. Trend search, memory, and explanation need their own tasks. Design intuition remains useful for hypotheses and anomalous-result interpretation, not as a substitute for measurement.
Applying it
- Turn disagreement into a measurable task, such as estimating an A-to-B ratio or finding the maximum, and score against truth rather than asking which design “feels clearer.”
- Compare a new encoding with a credible baseline, randomize stimuli, and span true values, sizes, positions, and displays. Record missingness, device, and quality-control outcomes.
- Report error, systematic bias, time, tail failures, and participant variation together. Preserve uncertainty when conditions are close instead of forcing a unique rank.
- Bind conclusions to version, task, and audience. Cite those conditions in review and rerun after material product changes rather than freezing one mean into a component rule.
Related
- Same group: U1.01.1 Visual channels have a stable empirical accuracy ranking · U1.01.3 Quantitative and categorical data need two different channel rankings · U1.01.4 The most important variable should get the highest-ranked available channel · U1.01.5 The ranking is an accuracy ceiling; alignment and mark size erode it
- Adjacent: U1.02.1 Position along a common scale is the most accurate channel · U1.03.1 Length judgment requires a common starting point
- Search terms:
magnitude judgment·log absolute error·graphical perception experiment