Too many entries reduce the proportion actually used by evaluators
Aliases: heuristic checklist · evaluation fatigue · entry design · layered checklist
What it is
The longer a checklist gets, the more easily an evaluator drifts from "check each entry deliberately" to "hunt for problems by impression." A heuristic set should have a small number of stable core entries, paired with product-specific sub-checks; more entries is not automatically more rigorous. This card is about how large the checklist itself should be — a different stage of the same batch of problem records that the previous two cards address: those two govern how records get logged and aggregated during evaluation; this one governs what size the checklist should be set to before evaluation even starts, so the records themselves do not become distorted from the outset.
Why it happens
An evaluator's attention and working memory are a limited resource within a single review session. Dozens of abstract criteria laid out in front of them causes entries to overlap and conflict in ordering, and the evaluator quickly loses the ability to maintain "check each one carefully" as a sustained pace. The degradation typically follows one of two paths: either only the first few or most memorable entries stick, and the rest turn into mechanical checkmarks; or the evaluator abandons item-by-item comparison altogether, goes by first impression to find whatever "feels off," and only afterward searches the list for a roughly matching label. Both degradations systematically bias the outcome — interface areas reviewed later in the session get inspected less thoroughly regardless of their actual quality, and this is not a shortfall in the evaluator's ability; the checklist's length itself has exceeded the attentional budget a single review session can sustain. It is a checklist design problem, not an execution problem.
Where it stops holding
A complex product genuinely needs more check items than a generic checklist provides, but the fix is not cramming every item into one list and running it all at once — it is layering: a small, stable set of core principles forms the first layer, while domain-specific checks, platform-specific checks, accessibility checks, and concrete task scripts each form their own independent layer, with different review stages drawing on different subsets rather than requiring every evaluator to face the full set every time. Project-specific check items written for a concrete business scenario tend to catch real problems more easily than generic abstract entries, precisely because a concrete item lowers the translation cost of mapping an abstract principle onto the interface at hand — and that translation cost is exactly the cognitive resource that gets overwhelmed when the checklist runs too long.
Applying it
- Split the checklist into four categories — core entries, domain sub-checks, platform sub-checks, and task scripts — and select the relevant subset for the current interface type and review stage before starting, instead of defaulting to running the full set.
- Write each check item with the specific observable evidence to look for and how to judge it, and remove any item that merely restates another entry in different words, cutting down on overlap between entries.
- Track each evaluator's actual usage and finding rate per entry; an entry with persistently low usage that also overlaps heavily with others is a candidate to merge or demote from the core list to an optional supplement.
- How to check: cap the interface scope and duration of a single review session, and for larger products split the review into multiple rounds by module. After a review, compare the number of problems found in interface areas reviewed earlier versus later — a clear pattern of "fewer problems found the later it was reviewed" means the current checklist length has exceeded what one session's attentional budget can sustain, and it needs trimming or restructuring into layers.
Related
- Same group: B3.18.1 The same problem often fits several heuristics; classification disagreement does not invalidate the problem · B3.18.2 Heuristic entries are not mutually exclusive; counting problems per entry double-counts · B3.18.4 Problems fitting no entry should be kept; they show the current set is incomplete
- Nearby: Q2 Usability Evaluation · A1 Attention
- Search terms:
heuristic set·evaluation fatigue·checklist design·attentional budget