Underspecified Human Decision Experiments Considered Harmful

Explainable AI (XAI)AI-Assisted Decision-Making & AutomationData Scientists & AnalystsAI/ML Researchers & Engineers

Research Background and Issues

  • Identified Issues: There is no clear consensus on how to define decision-making problems in experiments or how to draw conclusions about human decision-making flaws. Many studies claiming to identify biases in human behavior lack the necessary experimental design foundation.
  • Significance: Human decision-making based on information may be influenced by AI-assisted tools or data visualization. In the absence of reliable and rigorous methods, research findings may become uninterpretable or mislead subsequent work.
  • Research Motivation and Related Work: Based on statistical decision theory and information economics, the authors aim to provide broadly applicable definitions and guidance for experimental design in human decision-making evaluation and assess whether current studies meet these standards.

Solution

  • Proposed Approach: The authors integrate statistical decision theory and information economics to define "decision problems" and "normative behavior" and recommend minimum standards that controlled experiments should meet to evaluate human decision quality.
  • Innovations: The framework allows for systematic identification of studies with insufficient experimental conditions and clarifies potential sources of performance loss (e.g., prioritization errors, reception errors), thus providing theoretical guidance for experimental design.
  • Implementation Steps:
    1. Define the basic components of a decision problem: state space, data generation model, signals, action space, and scoring rules.
    2. Derive normative behavior from experimental conditions, i.e., the optimal choices participants can make to maximize utility after receiving signals.
    3. Compare experimental design with the normative framework to identify potential issues of insufficient information.
    4. Recommend best practices for experimental design, such as providing clearly explained scoring rules, conveying sufficient information about the data generation model, and using standardized scoring rules.

Research Outcomes

  • Specific Findings: The authors analyzed 46 AI-assisted decision-making studies, of which 39 could be applied to the framework, but only 10 met the conditions for a well-defined decision problem. Most of the remaining studies failed to provide participants with enough information to make optimal responses.
  • Advantages Over Existing Solutions: The proposed framework enables researchers to interpret experimental results more rigorously while reducing noise caused by experimental errors. It allows researchers to more clearly identify potential sources of performance loss.
  • Experimental or Evaluation Results: The evaluation revealed that many experimental designs failed to use standardized scoring rules for incentives and evaluation, making results difficult to compare accurately. Additionally, over 70% of studies did not provide participants with sufficient information to make optimal responses.
  • Limitations and Future Directions:
    • Limitations: The framework requires participants to fully understand the scoring rules and data generation model, but participants may misunderstand due to the complexity of the information. Furthermore, learning effects caused by different experimental conditions have not been fully controlled.
    • Future Directions: Explore more effective ways to isolate specific sources of loss (e.g., cognitive processing errors or optimization failures) and apply this framework to more complex decision-making scenarios. Additionally, further research could focus on improving participants' understanding of scoring rules and experimental information.

Conclusion

This paper establishes new standards for experimental design and result interpretation by providing rigorous theoretical definitions and frameworks. The authors argue that research findings may be misleading when experimental designs fail to adequately define decision problems. By more clearly conveying experimental information, defining scoring rules, and setting evaluation standards, the interpretability and credibility of research in this field can be significantly improved.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189091/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714063
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Explainable AI (XAI), AI-Assisted Decision-Making & Automation
work
Professions
Data Scientists & Analysts, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
10 related papers