Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant

Human-LLM CollaborationAI-Assisted Decision-Making & Automation

Research Background and Problem

  • What problems or challenges have the authors identified?

    • Although large language models (LLMs) have the potential to serve as daily assistants, their effectiveness in planning and sequential decision-making tasks remains unclear.
    • The mechanisms for building and calibrating trust when users collaborate with LLM agents are not well understood.
    • While LLMs perform well on simple tasks, their reliability and trustworthiness are questioned in high-risk tasks (e.g., financial transactions or credit card payments).
    • Overtrust or distrust in AI can lead to improper reliance on the system or inefficiency.
  • Why is this problem important?

    • LLM agents, as daily assistants, can enhance productivity and reduce human workload across various task scenarios, but their errors could result in high-risk consequences (e.g., financial loss).
    • Research on human-AI collaboration is critical for designing more efficient AI systems and improving the user experience with AI.
  • Research Motivation and Related Work

    • Existing studies show that LLMs exhibit high flexibility and interactivity in logical planning for task-solving, but their automated planning capabilities are limited, requiring user involvement to improve performance.
    • Trust calibration and appropriate reliance are crucial for human-AI collaboration, but current experiments primarily focus on specific use cases, lacking comprehensive studies on LLM agents in everyday task scenarios.

Solution

  • What methods or solutions do the authors propose?

    • Employ a "Plan-then-Execute" framework to separate the planning and execution processes of LLM agents. Users can participate in both the high-level planning and real-time execution phases to control plan quality and correct execution errors.
    • Propose a 2×2 factorial experimental design to study the impact of the automation level of LLM agents (automated planning/user-involved planning, automated execution/user-involved execution) on user trust and task performance.
  • What is innovative about this solution?

    • Integrates user involvement with LLM automation capabilities to explore the specific impact of human participation on trust calibration and task outcomes.
    • Provides a user interaction interface, enabling users to flexibly edit and correct plans and execution results for more efficient task completion.
    • For the first time, validates the effectiveness and limitations of LLM agents in comprehensive task scenarios (low-risk, high-risk, simple, complex).
  • What are the implementation steps and key technologies used?

    1. Planning Phase:
      • Users evaluate and edit step-by-step plans generated by the LLM, including adding, deleting, or splitting steps.
    2. Execution Phase:
      • Users review each execution action by the LLM, choosing to accept, provide feedback, or override the execution.
      • LLM agent actions are simulated through predefined APIs.
    3. Experimental Design:
      • 248 participants tested six tasks (e.g., financial transactions, credit card payments, travel planning, etc.).
      • Data was collected on user trust, trust calibration (judgment of the accuracy of plans/execution results), task performance (plan quality and execution accuracy), and users' subjective experiences.

Research Findings

  • What specific results were achieved?

    • User involvement effectively improved task performance by correcting flawed plans and erroneous executions.
    • Task performance was highly correlated with plan quality, but user involvement in plan editing sometimes reduced plan quality (e.g., when correct plans were mistakenly altered).
    • Users exhibited uncalibrated trust in "superficially plausible but incorrect" plans generated by LLMs, leading to reliance errors.
    • User involvement significantly increased cognitive load, reducing confidence in task performance.
  • What advantages does it have compared to existing solutions?

    • Provides users with the ability to directly correct AI agents, especially when major errors are identified during execution.
    • Balances trust, task performance, and user cognitive load, highlighting the need for differentiated human-AI collaboration designs for different task types.
  • What are the experimental or evaluation results?

    • Trust Calibration: Users exhibited higher calibrated trust in high-quality plans compared to low-quality plans, but user involvement did not significantly improve this calibration.
    • Task Performance:
      • For significantly flawed plans (e.g., "plans with grammatical errors"), user involvement substantially improved overall plan quality.
      • User involvement in execution improved final task accuracy, but automated execution outperformed user involvement in strictly evaluated action sequences.
    • User Cognitive Load: Cognitive load was significantly higher in user-involved planning and execution compared to automated modes.
  • Limitations and Future Directions

    • Limitations:
      • The current experiments were conducted in controlled simulation environments, which do not fully capture the complexity of real-world tasks.
      • Users are prone to high trust in "superficially plausible but incorrect" generated plans, leading to misuse.
    • Future Directions:
      • Develop dynamic and flexible collaborative workflows that allow users to make multiple corrections during planning and execution.
      • Investigate how to provide transparency and feedback to reduce cognitive load while enhancing users' critical evaluation of LLM-generated content.
      • Explore risk assessment mechanisms to dynamically adjust user involvement based on task complexity and risk.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189597/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713218
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Human-LLM Collaboration, AI-Assisted Decision-Making & Automation
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
9 related papers