Gaze must be interpreted with task performance
Aliases: eye tracking with outcomes · fixation efficiency · process tracing
What it is
The same long dwell can be interest and comprehension, or being stuck and rechecking without understanding. Those readings cannot be told apart from the gaze record alone. Joint interpretation puts gaze metrics in the same unit of analysis as that attempt’s success, time, errors, or choice: ask whether the task was done, then what the eyes were doing. Gaze is process evidence; performance is the test of whether the process helped.
Why it happens
Fixations are driven by goals, visual salience, and current uncertainty. Efficient succeeders may rarely regress because they already know where to go; failers may also rarely regress because they never found the patch. Confused people pile dwell on the wrong option. Without performance, “looked more” can be narrated as importance and “looked less” as neglect, and both stories close. Once aligned to the same attempt, three patterns separate: looked at the critical region and still failed (comprehension or mapping), never looked and failed (discoverability), looked and succeeded (process and outcome agree). Those three imply different redesigns.
Studying it
Give each attempt a row: performance plus gaze (whether a critical region was entered, time to first fixation, regressions, dwell). Stratify gaze by success/failure, or predict performance from gaze and inspect counterexamples. Define critical regions and the time window in which they “should” be looked at before seeing who won, so winners’ looking is not retrofitted as critical. Sample size must support the split; otherwise play successful and failed paths side by side as cases. Do not treat mean gaze across people as an explanation of mean performance across people.
Where it stops holding
In browsing, aesthetics, or ad-attention work with no right answer, performance may not be the target; joint interpretation then pairs gaze with choice, memory, or later recognition. If process tracing itself changes action—talking aloud, an uncomfortable headset—performance is contaminated and the joint analysis writes method reactivity into the conclusion. Experts’ short-gaze high-success pattern is efficiency; applying a novice “not enough looking” story reverses the meaning. Sessions with failed calibration should be dropped whole, not stripped of gaze while keeping performance.
Applying it
- Open the readout with a four-cell table of success/failure × entered-critical-region, not with a mean heatmap.
- For “looked and still wrong,” change copy, rules, or feedback; for “never looked and wrong,” change placement, timing, or contrast—do not apply the same patch to both.
- If gaze sits on price and conversion stays low, inspect what the price means and what follows; do not treat looking as conversion evidence.
- Accept a redesign only if performance improves and failed attempts’ gaze patterns move off “never found”; a change in where people look is not completion.