A single evaluator finds only a small part of all problems, and individuals differ widely
Aliases: evaluator effect · heuristic evaluation · individual differences
What it is
The evaluator effect means heuristic-evaluation findings depend strongly on the person: with the same interface and checklist, different evaluators often find only partially overlapping problem sets. Any single-person report is a sample, not the complete problem inventory.
Why it happens
Detection depends on domain knowledge, platform experience, attention path, interpretation of the checklist, task script, and fatigue. Experts may see business semantics or platform conflicts more sharply yet skip areas they deem unimportant; novices may find surface gaps but miss permission and data consequences. The interface problem space is usually larger than one pass of attention.
Studying it
Have multiple evaluators inspect the same product independently, recording each person’s findings, time, entries used, and areas visited. After merging, calculate individual detection rate, pairwise overlap, unique findings, and coverage by severity. Comparing experience, checklists, and scripts explains sources of difference.
Where it stops holding
Difference is not error. Some unique findings are real; others are false positives, duplicates, or preferences, so verify rather than average. If the product is small, the task narrow, and evaluator backgrounds similar, overlap may be high. A single quick walkthrough is still useful early, but cannot claim completeness.
Applying it
- Label early single-person walkthroughs “single-rater, unverified,” then schedule independent review.
- Assemble complementary backgrounds covering at least two of domain, platform, accessibility, content, or engineering.
- Give evaluators the same task script and problem template, but require independent completion before merging.
- Evaluate the interface against the merged set, not against one person’s findings.
Related
- Same group: B3.19.2 Total findings show diminishing returns as evaluators increase, often flattening after three to five people · B3.19.3 Diminishing-return estimates assume independent evaluators with equal detection probability, conditions rarely met · B3.19.4 Low overlap means the problem space is large and more evaluators are needed, not that quality is poor · B3.19.5 Evaluate independently before aggregating; prior discussion erases independence and distorts the number effect
- Nearby: Q2 Usability Evaluation · Q4 Research Methods and Evaluation
- Search terms:
evaluator effect·heuristic evaluation·individual differences
Cards in the same group
- B3.19.2Total findings show diminishing returns as evaluators increase, often flattening after three to five people
- B3.19.3Diminishing-return estimates assume independent evaluators with equal detection probability, conditions rarely met
- B3.19.4Low overlap means the problem space is large and more evaluators are needed, not that quality is poor
- B3.19.5Evaluate independently before aggregating; prior discussion erases independence and distorts the number effect