Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
Authors
Research Background and Problem
-
What problems or challenges have the authors identified?
- Although large language models (LLMs) have the potential to serve as daily assistants, their effectiveness in planning and sequential decision-making tasks remains unclear.
- The mechanisms for building and calibrating trust when users collaborate with LLM agents are not well understood.
- While LLMs perform well on simple tasks, their reliability and trustworthiness are questioned in high-risk tasks (e.g., financial transactions or credit card payments).
- Overtrust or distrust in AI can lead to improper reliance on the system or inefficiency.
-
Why is this problem important?
- LLM agents, as daily assistants, can enhance productivity and reduce human workload across various task scenarios, but their errors could result in high-risk consequences (e.g., financial loss).
- Research on human-AI collaboration is critical for designing more efficient AI systems and improving the user experience with AI.
-
Research Motivation and Related Work
- Existing studies show that LLMs exhibit high flexibility and interactivity in logical planning for task-solving, but their automated planning capabilities are limited, requiring user involvement to improve performance.
- Trust calibration and appropriate reliance are crucial for human-AI collaboration, but current experiments primarily focus on specific use cases, lacking comprehensive studies on LLM agents in everyday task scenarios.
Solution
-
What methods or solutions do the authors propose?
- Employ a "Plan-then-Execute" framework to separate the planning and execution processes of LLM agents. Users can participate in both the high-level planning and real-time execution phases to control plan quality and correct execution errors.
- Propose a 2×2 factorial experimental design to study the impact of the automation level of LLM agents (automated planning/user-involved planning, automated execution/user-involved execution) on user trust and task performance.
-
What is innovative about this solution?
- Integrates user involvement with LLM automation capabilities to explore the specific impact of human participation on trust calibration and task outcomes.
- Provides a user interaction interface, enabling users to flexibly edit and correct plans and execution results for more efficient task completion.
- For the first time, validates the effectiveness and limitations of LLM agents in comprehensive task scenarios (low-risk, high-risk, simple, complex).
-
What are the implementation steps and key technologies used?
- Planning Phase:
- Users evaluate and edit step-by-step plans generated by the LLM, including adding, deleting, or splitting steps.
- Execution Phase:
- Users review each execution action by the LLM, choosing to accept, provide feedback, or override the execution.
- LLM agent actions are simulated through predefined APIs.
- Experimental Design:
- 248 participants tested six tasks (e.g., financial transactions, credit card payments, travel planning, etc.).
- Data was collected on user trust, trust calibration (judgment of the accuracy of plans/execution results), task performance (plan quality and execution accuracy), and users' subjective experiences.
- Planning Phase:
Research Findings
-
What specific results were achieved?
- User involvement effectively improved task performance by correcting flawed plans and erroneous executions.
- Task performance was highly correlated with plan quality, but user involvement in plan editing sometimes reduced plan quality (e.g., when correct plans were mistakenly altered).
- Users exhibited uncalibrated trust in "superficially plausible but incorrect" plans generated by LLMs, leading to reliance errors.
- User involvement significantly increased cognitive load, reducing confidence in task performance.
-
What advantages does it have compared to existing solutions?
- Provides users with the ability to directly correct AI agents, especially when major errors are identified during execution.
- Balances trust, task performance, and user cognitive load, highlighting the need for differentiated human-AI collaboration designs for different task types.
-
What are the experimental or evaluation results?
- Trust Calibration: Users exhibited higher calibrated trust in high-quality plans compared to low-quality plans, but user involvement did not significantly improve this calibration.
- Task Performance:
- For significantly flawed plans (e.g., "plans with grammatical errors"), user involvement substantially improved overall plan quality.
- User involvement in execution improved final task accuracy, but automated execution outperformed user involvement in strictly evaluated action sequences.
- User Cognitive Load: Cognitive load was significantly higher in user-involved planning and execution compared to automated modes.
-
Limitations and Future Directions
- Limitations:
- The current experiments were conducted in controlled simulation environments, which do not fully capture the complexity of real-world tasks.
- Users are prone to high trust in "superficially plausible but incorrect" generated plans, leading to misuse.
- Future Directions:
- Develop dynamic and flexible collaborative workflows that allow users to make multiple corrections during planning and execution.
- Investigate how to provide transparency and feedback to reduce cognitive load while enhancing users' critical evaluation of LLM-generated content.
- Explore risk assessment mechanisms to dynamically adjust user involvement based on task complexity and risk.
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How do user involvement levels affect trust calibration and task performance for LLM (large language model) planning and execution across task contexts?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How can trust, task performance, and user cognitive load be balanced when users are involved in LLM planning and execution?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- How reliable are LLMs in high-risk tasks (e.g., financial trading), and can user involvement significantly reduce error rates?Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
Practical Problems
1- Users struggle to trust and correctly use AI assistants in high-risk tasks, potentially causing errors or losses.Category: LLM Trust and Over/Under-RelianceSimilar questionsarrow_forward
- 100%
Rethinking Interaction: From Instrumental Interaction to Human-Computer Partnerships
CHI '18· Human-LLM Collaboration +1
- 100%
To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models
CHI '25· Human-LLM Collaboration +1
- 100%
Enhancing AI-Assisted Group Decision Making through LLM-Powered Devil's Advocate
IUI '24· Human-LLM Collaboration +1
- 100%
C-PAK: Correcting and Completing Variable-length Prefix-based Abbreviated Keystrokes
UIST '23· Human-LLM Collaboration +1
- 67%
Automating Clinical Documentation with Digital Scribes: Understanding the Impact on Physicians
CHI '21· Human-LLM Collaboration +1
- 67%
Unified Conversational Models with System-Initiated Transitions between Chit-Chat and Task-Oriented Dialogues
CUI '23· Conversational Chatbots +2
- 67%
TiiS: A Review of User Interface Design for Interactive Machine Learning
IUI '19· Human-LLM Collaboration +2
- 67%
Induction of an active attitude by short speech reaction time toward interaction for decision-making with multiple agents
IUI '19· Agent Personality & Anthropomorphism +2
- 67%
Exploring the Effects of Machine Learning Literacy Interventions on Laypeople's Reliance on Machine Learning Models
IUI '22· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)