Agreement does not guarantee correctness; shared misconceptions err in unison
Aliases: correlated errors · illusion of consensus · shared bias
What it is
Redundancy rests on the inference "errors are independent, so agreement implies truth." That premise can fail: when contributors share a misconception — the same ambiguous instruction, the same cultural defaults, the same misleading interface — their errors are correlated, and they err in unison. Comparison then shows high agreement, and quality control delivers a false negative: five people all wrong the same way sail through voting, weighting, any aggregation. The boundary demotes agreement from a credibility metric to a "random error is controlled" metric — it detects noise, and is blind to systematic bias.
Why it happens
Correlated errors have three main sources. Shared input: every contributor reads the same instructions, so an ambiguity or wrong default there is copied into every judgment — agreement stays perfectly healthy while the conclusion shifts wholesale. Shared priors: a population's common cultural and experiential defaults (stereotyped categorization of certain images, shared readings of certain tones of text) push everyone the same direction on particular inputs; the more homogeneous the crowd, the more correlated the priors. Shared environment: presentation order, thumbnail quality, option layout — the same rendering bias acts on everyone alike. The common structure: random error decays with redundancy (more people, more stability) while systematic bias does not (more people just err more confidently) — redundancy has no immunity to the latter because all of its power is staked on independence. Seeded gold questions are the only routine instrument that pierces this: known correct answers bypass the crowd's consensus and expose collective drift directly.
Studying it
- Paradigm: bias-injection experiments — plant a known-direction bias in the instructions or presentation, measure how the systematic shift of the answer distribution relates to redundancy level (expected: shift does not decay with more copies while variance shrinks — the distribution tightens around the wrong center); cross-population comparisons (judgment differences across cultural or background groups) estimate the shared-prior contribution.
- Variables: bias source (instruction wording, presentation design, crowd composition) and redundancy level as independent variables; systematic bias of the aggregated result, within-group agreement, and gold-question hit rate as dependent variables.
- Use in interface research: designing "bias sentinels" for crowdsourcing platforms — a gold subset continuously checked against aggregated output, raising an alarm when agreement is high but gold hit rate is low.
- Methodological caveat: distinguishing correlated error from genuine task ambiguity requires gold or expert baselines — both look identical in the data (a stable deviation); cross-population comparisons must control for comprehension differences, or you measure a communication failure rather than a prior difference.
Where it stops holding
This boundary scopes redundancy-based quality control; it does not refute redundancy. For random error (fatigue, slips, random clicking) redundancy remains the first tool — the failure is only in cashing agreement in at full credibility value. Remedies have distinct homes: instruction-level correlation goes back to instruction revision and boundary examples; prior-correlated tasks call for crowd diversification (recruiting heterogeneous contributors) and cross-group contrast; environment correlation yields to presentation randomization. Expert judgment is not immune either — expert pools share training and paradigms, and their correlated errors can be tidier than the lay public's; professional consensus is not proof of independence.
Applying it
- Run agreement and gold hit rate as paired indicators: agreement covers random error, gold covers systematic bias — watching only the former misses collective drift.
- For high-stakes judgments, use stratified redundancy: separate contributor groups by background, require cross-group agreement, converting prior correlation into a between-group contrast.
- Randomize the presentation layer: shuffle option order, item order, and zoom levels across contributors to dilute shared-environment error.
- After any instruction revision, re-run the gold subset — revisions change systematic bias, not random error, and redundancy metrics will show no improvement.
- Verification: periodically sample high-agreement items for expert review and track the share that is actually wrong; a rising share signals correlated errors accumulating — audit the instruction, crowd, and presentation layers in that order.
Related
- Same group: V9.05.1 Having several people repeat the same task and comparing answers is the base quality control · V9.05.2 Redundancy scales cost linearly and must be tiered by task difficulty · V9.05.3 Seeded gold questions continuously estimate contributor accuracy · V9.05.5 Paying per completed item induces fast, low-quality work
- Nearby: V9.04 Crowdsourcing Task Decomposition and Instructions · V9.07 Collaborative Filtering and the Limits of Collective Wisdom
- Search terms:
correlated error·systematic bias·wisdom of crowds
Cards in the same group
- V9.05.1Having several people repeat the same task and comparing answers is the base quality control
- V9.05.2Redundancy scales cost linearly and must be tiered by task difficulty
- V9.05.3Seeded questions with known answers continuously estimate contributor accuracy
- V9.05.5Paying per completed item induces fast, low-quality work