Underspecified Human Decision Experiments Considered Harmful
Authors
Research Background and Issues
- Identified Issues: There is no clear consensus on how to define decision-making problems in experiments or how to draw conclusions about human decision-making flaws. Many studies claiming to identify biases in human behavior lack the necessary experimental design foundation.
- Significance: Human decision-making based on information may be influenced by AI-assisted tools or data visualization. In the absence of reliable and rigorous methods, research findings may become uninterpretable or mislead subsequent work.
- Research Motivation and Related Work: Based on statistical decision theory and information economics, the authors aim to provide broadly applicable definitions and guidance for experimental design in human decision-making evaluation and assess whether current studies meet these standards.
Solution
- Proposed Approach: The authors integrate statistical decision theory and information economics to define "decision problems" and "normative behavior" and recommend minimum standards that controlled experiments should meet to evaluate human decision quality.
- Innovations: The framework allows for systematic identification of studies with insufficient experimental conditions and clarifies potential sources of performance loss (e.g., prioritization errors, reception errors), thus providing theoretical guidance for experimental design.
- Implementation Steps:
- Define the basic components of a decision problem: state space, data generation model, signals, action space, and scoring rules.
- Derive normative behavior from experimental conditions, i.e., the optimal choices participants can make to maximize utility after receiving signals.
- Compare experimental design with the normative framework to identify potential issues of insufficient information.
- Recommend best practices for experimental design, such as providing clearly explained scoring rules, conveying sufficient information about the data generation model, and using standardized scoring rules.
Research Outcomes
- Specific Findings: The authors analyzed 46 AI-assisted decision-making studies, of which 39 could be applied to the framework, but only 10 met the conditions for a well-defined decision problem. Most of the remaining studies failed to provide participants with enough information to make optimal responses.
- Advantages Over Existing Solutions: The proposed framework enables researchers to interpret experimental results more rigorously while reducing noise caused by experimental errors. It allows researchers to more clearly identify potential sources of performance loss.
- Experimental or Evaluation Results: The evaluation revealed that many experimental designs failed to use standardized scoring rules for incentives and evaluation, making results difficult to compare accurately. Additionally, over 70% of studies did not provide participants with sufficient information to make optimal responses.
- Limitations and Future Directions:
- Limitations: The framework requires participants to fully understand the scoring rules and data generation model, but participants may misunderstand due to the complexity of the information. Furthermore, learning effects caused by different experimental conditions have not been fully controlled.
- Future Directions: Explore more effective ways to isolate specific sources of loss (e.g., cognitive processing errors or optimization failures) and apply this framework to more complex decision-making scenarios. Additionally, further research could focus on improving participants' understanding of scoring rules and experimental information.
Conclusion
This paper establishes new standards for experimental design and result interpretation by providing rigorous theoretical definitions and frameworks. The authors argue that research findings may be misleading when experimental designs fail to adequately define decision problems. By more clearly conveying experimental information, defining scoring rules, and setting evaluation standards, the interpretability and credibility of research in this field can be significantly improved.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can decision problems be defined in experiments to ensure reliability of research conclusions?Category: Research Method Practice, Coding Workflows, and Design Research ReflectionSimilar questionsarrow_forward
- How can adequacy of experimental design in existing research be evaluated and potential problems identified?Category: Research Method Practice, Coding Workflows, and Design Research ReflectionSimilar questionsarrow_forward
- How does lack of information affect participants' decision quality in AI-assisted decision research?Category: Research Method Practice, Coding Workflows, and Design Research ReflectionSimilar questionsarrow_forward
Practical Problems
1- Users make suboptimal choices in AI-assisted decision-making due to insufficient information.Category: AI Decision Support and Reliance BehaviorSimilar questionsarrow_forward
- 100%
The Role of Initial Acceptance Attitudes Toward AI Decisions in Algorithmic Recourse
CHI '25· Explainable AI (XAI) +1
- 100%
The Amplifying Effect of Explainability in AI-assisted Decision-making in Groups
CHI '25· Explainable AI (XAI) +1
- 100%
Guided Reflection in AI-Assisted Decision-Making: Effects on AI Overreliance and Decision Accuracy
CHI '26· AI-Assisted Decision-Making & Automation +1
- 100%
Understanding the Effects of AI-Assisted Critical Thinking on Human-AI Decision Making
CHI '26· AI-Assisted Decision-Making & Automation +1
- 80%
Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
CHI '23· Explainable AI (XAI) +1
- 80%
Towards Estimating Missing Emotion Self-reports Leveraging User Similarity: A Multi-task Learning Approach
CHI '24· Explainable AI (XAI) +1
- 67%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 67%
Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine Learning
CHI '20· Explainable AI (XAI) +2
- 67%
No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
CHI '20· Explainable AI (XAI) +2
- 67%
User Ex Machina : Simulation as a Design Probe in Human-in-the-Loop Text Analytics
CHI '21· Explainable AI (XAI) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)