B3.12.3Severity Ratingdesignresearch

Multiple evaluators rate independently before merging

Aliases: inter-rater agreement · independent rating · severity calibration

What it is

Severity should first be rated independently by multiple evaluators who have not seen one another’s scores; only then should differences be aggregated, evidence discussed, and a final level formed. Independent rating reduces the influence of personal experience, seniority, and persuasion order, and makes disagreement inspectable.

Why it happens

Rating depends on interpreting frequency, impact, persistence, and business constraints. A single rater may be anchored by the most dramatic clip, a recent event, or a familiar role; multiple sources cover different users, platforms, and professional paths. Discussion before rating anchors scores, while discussion afterward exposes evidence gaps. Merging is not simple averaging; record score distribution, reasons, and final assumptions.

Studying it

Give evaluators the same problem description, video or logs, task context, and level anchors. Ask for independent scores and written assumptions. Calculate score distribution or agreement, hold a divergence review, then revise the problem description or weights. A calibration set can periodically check whether evaluators assign similar levels to similar problems.

Where it stops holding

Too few raters, or raters with similar backgrounds, do not guarantee representativeness; add data or invite real task roles. Vague problem descriptions create false disagreement, so distinguish “different problems” from “different interpretations.” Urgent safety defects need not wait for full consensus; fix first and complete the record afterward.

Applying it

  • Have at least two evaluators with different backgrounds rate independently; record each score, rationale, and uncertainty.
  • When scores differ by more than one level, revisit evidence for frequency, impact, recovery, and audience.
  • Keep the final level, merging rationale, and rejected assumptions available for later review.
  • Build a calibration library so new evaluators rate samples before joining formal prioritization.

Related

  • Same group: B3.12.1 Severity synthesizes frequency, impact, and persistence · B3.12.2 Ratings support ordering rather than absolute judgment
  • Nearby: Q2 Usability Evaluation · Q4 Research Methods and Evaluation
  • Search terms: inter-rater reliability · severity calibration · heuristic evaluation

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/B3.12.3