Criticality: Scaffolding Decision-Making with Interactive Critical Thinking and Evidence-Based Reasoning Traces
Authors
Decision-making requires examining underlying assumptions and concepts, considering diverse perspectives, and weighing potential consequences with clear, accurate reasoning. Recent large language models (LLMs) show promise for assisting decision-makers by combining reasoning capabilities with the ability to retrieve relevant information from large documents. However, our formative study with five professional decision-makers revealed key limitations of using LLM in workflow: time-consuming alignment of user goals, lack of evidence-based grounding, overwhelmingly long outputs, and unsurfaced assumptions undermined user trust in the LLM output and the validity of the final decision. We introduce Criticality, a system that operationalizes the Paul-Elder Critical Thinking framework to structure reasoning into interactive Elements of Thought (e.g., purpose, assumptions, perspectives, implications), and evaluates and guides reasoning using Intellectual Standards (e.g., clarity, fairness, logic). It also retrieves evidence for each claim, classifies it as supporting, neutral, or contradictory, and explains the claim-evidence link. A within-subjects study (n=13) comparing Criticality to ChatGPT 5 Pro, a state-of-the-art reasoning model in conversational interface, found that Criticality improved user interaction of steering and repairing through the decision-making process, producing better decision rationales compared to the baseline.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering Systems
CHI '26· Human-LLM Collaboration +2
- 86%
Improving Human Verification of LLM Reasoning through Interactive Explanation Interfaces
IUI '26· Human-LLM Collaboration +2
- 75%
Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation
CHI '26· Human-LLM Collaboration +3
- 75%
Evalet: Evaluating Large Language Models through Functional Fragmentation
CHI '26· Human-LLM Collaboration +3
- 75%
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
CHI '26· Human-LLM Collaboration +3
- 75%
Key Considerations for Domain Expert Involvement in LLM Design and Evaluation: An Ethnographic Study
IUI '26· Human-LLM Collaboration +3
- 71%
AI of Oz: Enhancing Wizard of Oz Studies in HCI with AI Assistance for Human Moderation
CHI '26· Human-LLM Collaboration +2
- 71%
LLM-based In-situ Thought Exchanges for Critical Paper Reading
IUI '26· Human-LLM Collaboration +2
- 71%
DiscipLink: Unfolding Interdisciplinary Information Seeking Process via Human-AI Co-Exploration
UIST '24· Human-LLM Collaboration +2
- 67%
"Shall We Dig Deeper?": Designing and Evaluating Strategies for LLM Agents to Advance Knowledge Co-Construction in Asynchronous Online Discussions
CHI '26· Human-LLM Collaboration +3
Based on Jaccard similarity of research subtopics & professions (≥60%)