From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered Analysis
Authors
Research Background and Issues
-
What problems or challenges did the authors identify?
When using AI-assisted decision-making systems, people often fail to appropriately leverage AI model recommendations or explanations. This can lead to over-reliance (accepting incorrect AI decisions) or under-reliance (ignoring correct AI decisions), thereby reducing decision accuracy. Moreover, when AI models lack explainability, these issues become more pronounced, making it difficult for users to critically analyze AI suggestions, further deteriorating decision quality. -
Why is this issue important?
AI-assisted decision-making has been widely applied in critical domains such as criminal justice and financial investment. If human decision-makers fail to appropriately utilize AI recommendations, it could lead to severe consequences, such as fairness issues or resource wastage. Therefore, studying how to help decision-makers effectively use AI models in the absence of explanations is crucial. -
Research Motivation and Related Work
Traditional remedies include providing users with model decision explanations through explainable AI techniques, but these explanations are often difficult for users to understand. Recent advancements in large language models (LLMs) have demonstrated exceptional conversational and analytical capabilities, enabling further analysis of AI recommendations in natural language. This presents new opportunities to improve AI-assisted decision-making, but their effectiveness and the design of appropriate interaction frameworks still require validation.
Solution
-
What methods or solutions did the authors propose?
The authors proposed an algorithmic framework to improve decision-makers' appropriate reliance on AI recommendations and decision accuracy through dynamically selecting and presenting LLM-generated analyses:- Using LLMs to generate natural language analyses of task features' contributions to AI model decisions.
- Designing experiments to compare the effects of different presentation methods (sequential and concurrent) of LLM analyses on human decision-making.
- Developing a human behavior model to characterize the impact of LLM analyses on decision-making and using this model to dynamically select the optimal information to present to users.
-
What are the innovative aspects of this solution?
Unlike traditional AI explanation methods, this framework leverages LLMs to generate analyses in natural language that are directly understandable to users. Additionally, the algorithm dynamically selects the "most useful" LLM analyses to present to decision-makers, enhancing information utilization efficiency and reducing cognitive load. Furthermore, the framework optimizes the intelligent interaction process between humans and LLMs, significantly improving decision quality. -
What are the implementation steps and key technologies used?
- Constructing a task-specific random forest AI model and generating task instances.
- Using GPT-4 to generate feature contribution analyses for each task.
- Developing a three-part human behavior model:
- Initial state mapping: Encoding the decision state of humans before encountering LLM analyses.
- State update: Representing the dynamic changes in human states after receiving LLM analyses.
- Final decision mapping: Translating the final latent state from human interaction into specific decisions.
- Proposing an expected utility-based algorithm to dynamically select analyses that maximize decision effectiveness.
- Validating the framework's effectiveness through empirical experiments.
Research Outcomes
-
What specific outcomes were achieved?
- Experiments showed that merely presenting LLM analyses sequentially or concurrently does not significantly improve AI-assisted decision-making performance.
- After adapting LLM analyses using the proposed algorithmic framework, human decision-makers demonstrated significantly higher decision accuracy in two tasks (income prediction and recidivism prediction).
- The algorithm significantly reduced the number of interaction rounds between humans and LLMs (from approximately 4.87 rounds to about 2.99 rounds), improving interaction efficiency.
-
What advantages does it have compared to existing solutions?
- Dynamic selection of LLM analyses enhances the effectiveness of information presentation and reduces cognitive load.
- In addition to improving decision accuracy, it effectively reduces inappropriate reliance on AI recommendations.
- Experimental results show that the algorithm selects more concise and useful LLM analyses, offering advantages over mechanisms based on feature importance (e.g., LIME).
-
What are the experimental or evaluation results?
Experiments demonstrated that incorporating algorithm-optimized LLM analyses significantly improved decision-makers' accuracy compared to relying solely on AI recommendations or randomly presented analyses. Moreover, decision-makers' over-reliance on AI models was significantly reduced (by up to approximately 20%), indicating that the algorithm effectively optimized trust calibration. -
Limitations and Future Directions
Limitations:- The current study focuses only on decision-making tasks involving structured data, while complex scenarios such as tasks based on visual or textual data may require further adaptation.
- The human behavior model is based on group-average characteristics and does not fully account for individual differences.
- The study uses the base GPT-4 model as the source for analysis generation, and results have not been validated for domain-specific LLMs.
Future Directions:
- Explore applying the framework to unstructured data tasks and investigate how to automatically or manually extract features for LLM analyses.
- Consider incorporating personalized characteristics, such as user confidence and cognitive styles, to further optimize interaction processes.
- Validate the framework's applicability in expert scenarios or high-risk decision-making environments.
- Address potential ethical issues, such as preventing algorithmic selection from being used for harmful interventions or manipulating user behavior.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can dynamically selecting and presenting LLM-generated analyses improve decision-makers' appropriate reliance on AI recommendations and decision accuracy?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How do different presentation modes (sequential vs. simultaneous) of LLM analyses affect human decision-making?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How can computational models of human behavior optimize information presentation for more efficient and intelligent human-AI interaction?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- Users either over-rely on or ignore AI recommendations, degrading decision quality.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 83%
Simulating Cooperative Prosocial Behavior with Multi-Agent LLMs: Evidence and Mechanisms for AI Agents to Inform Policy Decisions
IUI '25· Human-LLM Collaboration +2
- 80%
Effects of Communication Directionality and AI Agent Differences in Human-AI Interaction
CHI '21· Human-LLM Collaboration +1
- 80%
AI Knowledge: Improving AI Delegation through Human Enablement
CHI '23· Human-LLM Collaboration +1
- 80%
Towards Human-AI Deliberation: Design and Evaluation of LLM-Empowered Deliberative AI for AI-Assisted Decision-Making
CHI '25· Human-LLM Collaboration +1
- 80%
A Survey on Interactive Reinforcement Learning: Design Principles and Open Challenges
DIS '20· Human-LLM Collaboration +1
- 67%
You Complete Me: Human-AI Teams and Complementary Expertise
CHI '22· Human-LLM Collaboration +1
- 67%
Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
CHI '23· Human-LLM Collaboration +1
- 67%
Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts
CHI '23· Human-LLM Collaboration +1
- 67%
Automatic Macro Mining from Interaction Traces at Scale
CHI '24· Human-LLM Collaboration +1
- 67%
A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness Evaluations
CHI '24· AI-Assisted Decision-Making & Automation +2
Based on Jaccard similarity of research subtopics & professions (≥60%)