Understanding the Effects of AI-Assisted Critical Thinking on Human-AI Decision Making
Honorable MentionPaper Title
Understanding the Effects of AI-Assisted Critical Thinking on Human-AI Decision Making
Publication Info
- Topic area: Human-AI decision making and critical thinking enhancement
- Keywords: AI-Assisted Critical Thinking, human-AI collaboration, decision-making, critical reflection, counterfactual analysis, Explainable AI, over-reliance, cognitive load, metacognition, user studies
Background and Problem
- Problem / challenge: Human-AI decision-making often underperforms due to humans’ superficial engagement with AI outputs, leading to over-reliance or under-reliance on AI. Existing approaches focus on reflecting on AI outputs rather than improving humans’ reasoning processes.
- Significance: Addressing flaws in human reasoning can improve decision quality, reduce inappropriate reliance on AI, and enhance human autonomy in high-stakes domains like medical diagnosis and financial auditing.
- Motivation and related work: Prior methods like Explainable AI (XAI), hypothesis-driven XAI, and reflective questioning focus on AI outputs but fail to model or critique humans’ reasoning. Recent frameworks like Human-AI Deliberation and ExtendAI elicit human reasoning but lack domain-specific, structured analysis. This paper addresses these gaps by focusing on improving human reasoning through AI-assisted critique and correction.
Solution
- Proposed approach: AI-Assisted Critical Thinking (AACT) framework
- Leverages a domain-specific AI model to analyze and critique human decision rationales via counterfactual analysis.
- Supports humans in identifying and correcting flaws in their reasoning.
- Novelty:
- Adapts the Recognition/Metacognition (R/M) model to human-AI decision making, focusing on humans’ reasoning rather than AI outputs.
- Introduces counterfactual perspective-taking to critique human arguments for incompleteness, unreliability, and conflicts.
- Implements a structured critique-and-correction workflow using conversational AI.
- Demonstrates heterogeneous impacts across user subgroups in a controlled study.
- Procedure and key techniques:
- Critique: AI identifies flaws in human arguments (e.g., missing features, unreliable evidence, conflicting alternatives) using counterfactual analysis.
- Correction: AI provides targeted self-reflection prompts, correction suggestions, and data-based triangulation to help users revise their arguments.
- Workflow: A structured process of targeted self-reflection, AI-based suggestions, and empirical data insights is delivered via a conversational interface.
Results
- Concrete findings:
- AACT significantly reduces over-reliance on AI (e.g., over-reliance ratio: 0.543 for AACT vs. 0.67 for Recommender).
- AACT increases cognitive load (mental demand: 3.887 for AACT vs. 3.33 for Recommender).
- AACT does not improve decision accuracy for the general population but benefits users with high AI familiarity (balanced accuracy: 0.819 for AACT vs. 0.695 for Human-only in this subgroup).
- Advantage over baselines:
- AACT reduces over-reliance on AI more effectively than Recommender and Analyzer treatments.
- Encourages deeper critical thinking, such as seeking counter-evidence.
- Experiments / evaluation:
- Conducted a randomized user study with 402 participants on Prolific using a house price prediction task.
- Compared AACT with Recommender, Analyzer, and Human-only baselines across metrics like decision accuracy, AI reliance, learning, and subjective perceptions.
- Subgroup analyses revealed heterogeneous effects based on task familiarity, AI familiarity, and education.
- Limitations and future work:
- Increased cognitive load and under-reliance on AI when correct.
- Limited generalizability due to low-stakes tasks, tabular data, and a well-performing AI model.
- Future work includes evaluating AACT in high-stakes domains, with different data modalities, and over longitudinal studies.
Summary
The paper introduces the AI-Assisted Critical Thinking (AACT) framework, which enhances human-AI decision making by critiquing and correcting human reasoning through counterfactual analysis. AACT reduces over-reliance on AI and promotes critical reflection but increases cognitive load and under-reliance on correct AI predictions. A user study demonstrates its effectiveness for certain subgroups, such as users with high AI familiarity. Future work will focus on improving scalability, adapting to diverse domains, and addressing ethical considerations of cognitive effort. AACT represents a step toward AI systems that act as thought partners in decision making.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
The Role of Initial Acceptance Attitudes Toward AI Decisions in Algorithmic Recourse
CHI '25· Explainable AI (XAI) +1
- 100%
The Amplifying Effect of Explainability in AI-assisted Decision-making in Groups
CHI '25· Explainable AI (XAI) +1
- 100%
Underspecified Human Decision Experiments Considered Harmful
CHI '25· Explainable AI (XAI) +1
- 100%
Guided Reflection in AI-Assisted Decision-Making: Effects on AI Overreliance and Decision Accuracy
CHI '26· AI-Assisted Decision-Making & Automation +1
- 80%
Who Should I Trust: AI or Myself? Leveraging Human and AI Correctness Likelihood to Promote Appropriate Trust in AI-Assisted Decision-Making
CHI '23· Explainable AI (XAI) +1
- 80%
Towards Estimating Missing Emotion Self-reports Leveraging User Similarity: A Multi-task Learning Approach
CHI '24· Explainable AI (XAI) +1
- 67%
Gamut: A Design Probe to Understand How Data Scientists Understand Machine Learning Models
CHI '19· Explainable AI (XAI) +2
- 67%
Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine Learning
CHI '20· Explainable AI (XAI) +2
- 67%
No Explainability without Accountability: An Empirical Study of Explanations and Feedback in Interactive ML
CHI '20· Explainable AI (XAI) +2
- 67%
User Ex Machina : Simulation as a Design Probe in Human-in-the-Loop Text Analytics
CHI '21· Explainable AI (XAI) +3
Based on Jaccard similarity of research subtopics & professions (≥60%)