V9.05.4Correlated errors defeat agreementdesignresearch

Agreement does not guarantee correctness; shared misconceptions err in unison

Aliases: correlated errors · illusion of consensus · shared bias

What it is

Redundancy rests on the inference "errors are independent, so agreement implies truth." That premise can fail: when contributors share a misconception — the same ambiguous instruction, the same cultural defaults, the same misleading interface — their errors are correlated, and they err in unison. Comparison then shows high agreement, and quality control delivers a false negative: five people all wrong the same way sail through voting, weighting, any aggregation. The boundary demotes agreement from a credibility metric to a "random error is controlled" metric — it detects noise, and is blind to systematic bias.

Why it happens

Correlated errors have three main sources. Shared input: every contributor reads the same instructions, so an ambiguity or wrong default there is copied into every judgment — agreement stays perfectly healthy while the conclusion shifts wholesale. Shared priors: a population's common cultural and experiential defaults (stereotyped categorization of certain images, shared readings of certain tones of text) push everyone the same direction on particular inputs; the more homogeneous the crowd, the more correlated the priors. Shared environment: presentation order, thumbnail quality, option layout — the same rendering bias acts on everyone alike. The common structure: random error decays with redundancy (more people, more stability) while systematic bias does not (more people just err more confidently) — redundancy has no immunity to the latter because all of its power is staked on independence. Seeded gold questions are the only routine instrument that pierces this: known correct answers bypass the crowd's consensus and expose collective drift directly.

Studying it

  • Paradigm: bias-injection experiments — plant a known-direction bias in the instructions or presentation, measure how the systematic shift of the answer distribution relates to redundancy level (expected: shift does not decay with more copies while variance shrinks — the distribution tightens around the wrong center); cross-population comparisons (judgment differences across cultural or background groups) estimate the shared-prior contribution.
  • Variables: bias source (instruction wording, presentation design, crowd composition) and redundancy level as independent variables; systematic bias of the aggregated result, within-group agreement, and gold-question hit rate as dependent variables.
  • Use in interface research: designing "bias sentinels" for crowdsourcing platforms — a gold subset continuously checked against aggregated output, raising an alarm when agreement is high but gold hit rate is low.
  • Methodological caveat: distinguishing correlated error from genuine task ambiguity requires gold or expert baselines — both look identical in the data (a stable deviation); cross-population comparisons must control for comprehension differences, or you measure a communication failure rather than a prior difference.

Where it stops holding

This boundary scopes redundancy-based quality control; it does not refute redundancy. For random error (fatigue, slips, random clicking) redundancy remains the first tool — the failure is only in cashing agreement in at full credibility value. Remedies have distinct homes: instruction-level correlation goes back to instruction revision and boundary examples; prior-correlated tasks call for crowd diversification (recruiting heterogeneous contributors) and cross-group contrast; environment correlation yields to presentation randomization. Expert judgment is not immune either — expert pools share training and paradigms, and their correlated errors can be tidier than the lay public's; professional consensus is not proof of independence.

Applying it

  • Run agreement and gold hit rate as paired indicators: agreement covers random error, gold covers systematic bias — watching only the former misses collective drift.
  • For high-stakes judgments, use stratified redundancy: separate contributor groups by background, require cross-group agreement, converting prior correlation into a between-group contrast.
  • Randomize the presentation layer: shuffle option order, item order, and zoom levels across contributors to dilute shared-environment error.
  • After any instruction revision, re-run the gold subset — revisions change systematic bias, not random error, and redundancy metrics will show no improvement.
  • Verification: periodically sample high-agreement items for expert review and track the share that is actually wrong; a rising share signals correlated errors accumulating — audit the instruction, crowd, and presentation layers in that order.

Related

  • Same group: V9.05.1 Having several people repeat the same task and comparing answers is the base quality control · V9.05.2 Redundancy scales cost linearly and must be tiered by task difficulty · V9.05.3 Seeded gold questions continuously estimate contributor accuracy · V9.05.5 Paying per completed item induces fast, low-quality work
  • Nearby: V9.04 Crowdsourcing Task Decomposition and Instructions · V9.07 Collaborative Filtering and the Limits of Collective Wisdom
  • Search terms: correlated error · systematic bias · wisdom of crowds

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/V9.05.4