Each level requires its own measurement methods
Aliases: measurement mismatch · psychophysiology · experience sampling · delayed recall
What it is
The three levels form their judgments in different time windows and with different encodings, so each level's information is accessible only in a specific window — and each has its own measurement tradition: the visceral level is measured by instantaneous reactions, the behavioral level by task performance, the reflective level by memory and narrative. An instrument offset from its level produces measurement mismatch, and the mismatch is undetectable on its face — the numbers are real, the statistics significant, but they measured a different level.
Why it happens
The mismatch comes from encoding: how a level encodes experience determines what measurable trace it leaves. The visceral reaction exists only within a few hundred milliseconds of stimulus onset and is then covered by later processing; the only instruments that catch it are instantaneous ones — sub-second exposure judgments, millisecond-scale physiological signals, rapid preference tasks. A questionnaire filled out afterward captures only that reaction's residue. The behavioral sense of efficacy is distributed across every action in the task, so instantaneous measures under-sample it; performance data and the felt sense of control during operation catch it. The reflective verdict exists only afterward and long-term — asking "how do you like it" in the moment cannot reach it; delayed recall, diaries, and long-term attitude tracking can. Two canonical mismatches: a satisfaction questionnaire aimed at the visceral first impression — filled in after the fact, by which time the immediate reaction is complete and buried under the reflective summary, so it measures the summary, not the first impression; and instantaneous physiological response used to predict long-term attachment — skin conductance and pupil dilation reflect present arousal, with no stable channel to retention or advocacy weeks later, so betting on short-term arousal as a long-term predictor is a bet on a time-scale mismatch.
Studying it
Each level draws on a real measurement tradition. Visceral: liking ratings and two-alternative preference tasks after 50–500 ms exposure, affective priming, facial EMG (corrugator and zygomaticus activity separating valence), pupil dilation and skin conductance for arousal, first-fixation preference in eye tracking; exposure paradigms are cheap, physiological measures require lab equipment. Behavioral: completion rate, error rate, and task time as performance data, plus control and efficacy self-ratings filled in immediately after the task, and process-level think-aloud during operation. Reflective: delayed recall weeks later, experience sampling and usage diaries, longitudinal tracking of attitudes and advocacy, retrospective interviews. Methodological note: the measurement time window must match the level's time scale — an instrument offset by one window picks up the adjacent level's signal; correlations between the three levels' measures are typically modest, so no level's measure can stand in for another's; physiological measures give arousal and valence direction, not attribution to a specific design element, and must be paired with an exposure design to point at elements.
Where it stops holding
The level-to-instrument mapping is not one-to-one: some instruments span levels — eye tracking captures both first-fixation preference and search paths mid-task; the difference lies in the sampling window and the interpretation of the metric, not the device. Reflective measurement has the longest cycle and highest cost, and longitudinal tracking carries attrition bias — retained users are systematically over-sampled, and the missing reflective data belongs precisely to the users who left, so the bias points the same way as the conclusion. Lab instantaneous measures have limited ecological validity: high agreement of judgments under ultra-brief exposure does not mean equal response intensity in a real first visit. And questionnaires are not ruled out — state which level is being measured and place them in a matching window (immediately after the task for the behavioral level, weeks later for the reflective level); the same scale at different time points is a different level's instrument. The error is letting one time point's questionnaire pose as the conclusion for all levels.
Applying it
- Choose instruments by the decision's target level: visual and material decisions get exposure tests or rapid preference tasks; flow and interaction decisions get performance data plus immediately-post-task efficacy self-ratings; meaning and brand decisions get delayed interviews, diaries, and long-term metric tracking.
- Label every metric with which level it samples and in what time window; report per level, never merged into one total score.
- When a metric looks wrong, ask which level the instrument sampled before attributing anything to the design — a negative delayed-interview finding is not a brief for reworking the first screen, and a cold first impression is not a brief for reworking the flow.
- To validate: pair two instruments from different levels on the same decision (a first-impression exposure test plus a delayed interview weeks later) and check that each level's conclusion stays in its own lane rather than substituting for the other; any document using one level's data to argue for another level's change is on-site evidence that the mismatch has already happened.
Related
- Same group: P1.08.1 The three levels run on time scales orders of magnitude apart · P1.08.2 Visceral judgments are the most stable across cultures · P1.08.3 The reflective level rewrites the memory of lower-level experience
- Nearby: P1.01.4 The three levels can contradict each other · P1.03.1 Overall evaluation is dominated by the peak and the end
- Search terms:
measurement mismatch·facial EMG·experience sampling·delayed recall