V9.04.3Boundary examples anchor judgmentdesignresearch

Boundary examples unify judgment better than abstract rules

Aliases: boundary cases · judgment anchoring · examples over rules

What it is

To make thousands of strangers judge by one standard, abstract rules will not do; boundary examples will. An abstract rule ("flag aggressive content as violating") handles clear cases identically for everyone — disagreement lives at the rule's edge: is sarcasm aggression, does quoting an insult count as spreading it. Boundary examples turn those edge cases into pre-decided instances ("this one counts, this one does not") placed inside the instructions; participants align on anchors at the edge, and the whole judgment distribution tightens. Rules give direction; examples nail the scale marks. Without examples, the scale marks are improvised from each person's defaults.

Why it happens

Human application of a category standard is example-driven, not rule-driven. Rules are encoded in language, and language's elasticity lets the same sentence unfold into different application boundaries in different heads; examples bypass language, supplying a non-negotiable "this one judges like so," and participants handle new inputs by analogy rather than interpretation — it sorts toward whichever example it resembles. Edge examples do the heavy lifting: clear examples (obvious extremes) merely confirm common sense and leave the contested band untouched, while an example placed exactly in the disputed zone pins the rule's elasticity to a concrete position. Examples also have a memory advantage — working memory stores instances, not clauses; what participants actually consult while judging is the example library. This is why adding ten boundary examples often beats three rewrites of the rule text: one edits a compressed representation of meaning, the other hands over expanded samples of the meaning itself.

Studying it

  • Paradigm: instruction-composition experiments — manipulate what the instructions contain (abstract rules only / rules plus clear examples / rules plus boundary examples / examples only), holding participants and inputs fixed, and compare the rise in inter-worker agreement and gold-standard hit rate; dose-response of example count and placement runs in the same setup.
  • Variables: presence, number, placement (clear zone versus contested zone), and positive-negative pairing of examples as independent variables; judgment agreement, deviation from expert baseline, and concentration of the distribution on edge inputs as dependent variables.
  • Use in interface research: making the example set a first-class object in task editors — pulling examples from historical contested judgments rather than leaving copywriters to invent them.
  • Methodological caveat: examples introduce their own anchoring bias — participants may mechanically sort new inputs toward the nearest example and ignore the rule, so the set must spread across regions of the rule space to disperse anchors. Contradictions between examples and rule text are fatal; cross-validation must precede publication.

Where it stops holding

Examples unify the perception of "where the edge is"; rules retain the explainability of "why the line sits there" — compliance settings where judgments must cite reasons and survive appeal need both, complementary rather than interchangeable. Example sets carry maintenance cost: standards drift with policy, and stale examples teach the old standard into the new regime. Highly professional judgment (medical imaging) is out of scope — there, uniformity comes from training and credentialing, not from a handful of examples on an instruction page.

Applying it

  • Fix the instruction structure as "one-sentence rule plus boundary example set": the rule gives direction, examples pin the edge; scale example count to the density of the contested band, pairing positive and negative instances.
  • Source examples from real contested judgments: after launch, extract inputs with the highest inter-worker disagreement, have experts decide them, and roll them into the set.
  • Attach one "why" line to each example — not a restatement of the rule, but the edge feature it sits on.
  • Run an example-rule consistency check before publishing: every example's verdict must follow from the rule text without contradiction.
  • Verification: track the entropy of judgment distributions over edge inputs as the example set grows; if entropy does not fall in the zone a new example targets, the example is not anchoring (its own verdict may be doubtful) — send it back for re-adjudication.

Related

  • Same group: V9.04.1 Crowdtasks must be decomposed to a granularity requiring no background knowledge · V9.04.2 Ambiguity in instructions converts directly into noise in results · V9.04.4 Unit duration determines mid-task abandonment · V9.04.5 The decomposition determines whether results can be reassembled
  • Nearby: V9.05 Quality Control and Redundancy in Crowdsourcing · V7.04 Content Governance
  • Search terms: boundary example · judgment anchoring · content moderation guidelines

Cards in the same group

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/handbook/V9.04.3