B3.17.2Golden Rulesdesign

The rules support generative design self-checks, not an evaluation scoring sheet

Aliases: generative self-check · heuristic evaluation · principle misuse · evidence gap

What it is

The golden rules are well suited to prompting the questions a designer should ask early on: is there a consistent model, who can use this, is feedback sufficient, how are errors prevented and recovered. Turning the eight rules directly into a 0–10 scorecard hides missing evidence; evaluation needs a task, users, data, and a specific question. This card pairs with "the eight rules are in tension with each other": that one explains why the rules cannot all be maximized at once; this one explains that even once you accept the tension, you still cannot quantify them into a comparable score table — quantification itself manufactures a false sense of precision that hides a lack of underlying evidence.

Why it happens

The golden rules are abstract directions and, by nature, carry none of the parameters an actual evaluation needs — a specific product's exposure volume, task criticality, role differences, and fix cost. When someone scores an interface against the eight rules one by one, what they are usually doing is conflating three entirely different things into one number: "I genuinely did not see this mechanism in the interface" (a missing-evidence finding), "I personally dislike this presentation" (a subjective preference), and "there is a legitimate reason this rule does not apply here" (a valid exception). Once these three are folded into a single score, that score can neither tell the team what to fix first nor explain itself to an outsider. A generative-stage self-check sidesteps this problem entirely, because it asks a different question: has the design set aside a mechanism to address a given rule, and has it written down how that will be verified later. It wants a yes/no and a plan, not a number precise to a decimal point — which is exactly why it works where a scorecard does not.

Where it stops holding

This does not mean the golden rules can never appear in an evaluation. They can generate specific check questions, such as "verify that every modal dialog can be exited" — once turned into a concrete, verifiable check item, this is a legitimate evaluation tool again. The problem is never "a rule appearing in an evaluation," it is "a rule being treated as a scoring dimension and scored directly without that translation." A review can also use a red/yellow/green marker to express "does evidence exist," instead of faking a precision that is not actually there. A genuine heuristic evaluation still depends on independent judgment from multiple evaluators plus observation of real tasks; a single person scoring against the eight rules alone, without either of those supports, produces a conclusion of limited value.

Applying it

  • During design review, translate each rule into a list of specific questions covering objects, paths, feedback, error, undo, control, memory, and closure — rather than scoring those dimensions directly.
  • Mark each check item only as "resolved," "undecided," "needs testing," or "known exception, risk accepted" — never a number that looks precise but has no grounding.
  • Prioritize evidence from real task observation, logs, and support data; the golden rules' role here is to help organize where to look, not to substitute for that first-hand evidence.
  • How to check: turn every "undecided" or "needs testing" item that surfaces during self-check into a concrete acceptance criterion or follow-up research question, and track whether it actually gets resolved in the next iteration, rather than letting the self-check list remain a one-off exercise.

Related

  • Same group: B3.17.1 The eight rules contain tensions; universal usability and expert accelerators pull against each other · B3.17.3 Closure requires explicit start and end markers for a task, a concern the other rules do not cover · B3.17.4 The rules are more abstract than specific criteria and must be translated into product-checkable clauses before use
  • Nearby: B3.11 Golden Rules · Q2 Usability Evaluation
  • Search terms: design review · heuristic evaluation · self-check · evidence gap

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/B3.17.2