V9.05.3Seeded gold questions estimate worker accuracydesignresearch

Seeded questions with known answers continuously estimate contributor accuracy

Aliases: gold questions · trap questions · gold standard · worker reliability

What it is

Redundancy and aggregation infer quality sideways from answer distributions; gold questions (trap questions) measure it directly: items with known correct answers, disguised as ordinary tasks, seeded randomly among real items — each miss enters the contributor's error record. Every participant's accuracy is thereby estimated continuously, serving three uses: screening (below-threshold workers exit), weighting (historically accurate workers carry more weight in aggregation), and calibration (bounding the credibility of the whole batch). The anonymous, transient workforce stops being unexaminable — it is being examined all the time.

Why it happens

The ingenuity lies in indistinguishability: a gold item must match real tasks in appearance — same format, same difficulty distribution, random cadence — or participants will recognize the "testing zone" and treat it differently, and the measurement dies on the spot. With indistinguishability intact, performance on gold items samples real performance, and error-rate estimates converge with sample size. Gold items also rewrite the participant's game: knowing they exist, the expected cost of random answering goes from zero to dismissal and docked pay, so minimal diligence becomes the dominant strategy — deterrence and measurement in one body. Random placement adds temporal coverage: gold scattered through a session captures the fatigue curve (error rates rising late) and drift over time, which a one-time entrance exam cannot. And the information economics beat pure redundancy: redundancy buys truth with repeated workers, gold items calibrate each worker with a few known answers — the same budget supports much finer weighted aggregation.

Studying it

  • Paradigm: dose and placement experiments for gold items — manipulating density, disguise fidelity, and consequences (record-only / dismissal / pay-linked), measuring detection rate by participants, estimated versus true accuracy, and overall output quality; substantial re-analyses of platform data validate online gold against post-hoc human gold standards.
  • Variables: gold density, disguise homogeneity, cadence, and consequence design as independent variables; participant detection rate, deviation of estimated from true accuracy, overall output quality, and retention as dependent variables.
  • Use in interface research: gold-injection strategy in task dispatchers — adaptive density (tightening when error rates rise), templates homogenizing gold appearance with real items.
  • Methodological caveat: once remembered, gold leaks (veterans recognize repeats), so pools must roll and expand; density has a backlash point — participants who feel toyed with quit or protest, damaging validity and retention together; and gold difficulty must match the real task distribution, or easier gold overestimates true accuracy.

Where it stops holding

Gold calibrates judgment accuracy and does not apply to creative tasks (writing, design) — nothing standard to embed. It detects carelessness and incompetence, not "a systematically different reading than the designer's" (an insightful minority may be cleared out as error). Legal and platform-policy boundaries: docking pay for gold misses is restricted in some jurisdictions and platform rules, so consequences mostly stay at access and weighting, never retroactively touching completed piecework. In high-skill volunteer communities, embedding tests carries a relational cost — being tested is an affront.

Applying it

  • Maintain a rolling, expanding gold pool, strictly homogenized to the real tasks' format and difficulty distribution, retiring high-exposure items quarterly.
  • Make density adaptive: when error rates on baseline gold rise, automatically thicken seeding in the next batch.
  • Use estimates at three levels: below threshold exits, the middle band feeds aggregation weights, top scorers need less redundancy on their items.
  • Keep consequences at access and weighting — never retroactive to piecework pay — to stay clear of compliance risk.
  • Verification: audit dismissed and retained workers by sampled human review, comparing gold-based estimates against human judgment; low agreement means suspect the gold's homogeneity (recognized or mis-difficultied) before suspecting the workers.

Related

  • Same group: V9.05.1 Having several people repeat the same task and comparing answers is the base quality control · V9.05.2 Redundancy scales cost linearly and must be tiered by task difficulty · V9.05.4 Agreement does not guarantee correctness; shared misconceptions err in unison · V9.05.5 Paying per completed item induces fast, low-quality work
  • Nearby: V9.06 Contributor Motivation and Payment · V9.04 Crowdsourcing Task Decomposition and Instructions
  • Search terms: gold standard · trap question · worker reliability

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/V9.05.3