When Should Users Check? Modeling Confirmation Frequency in Multi-Step Agentic AI Tasks
Authors
Paper Title
When Should Users Check? Modeling Confirmation Frequency in Multi-Step Agentic AI Tasks
Publication Info
- Topic area: Optimizing user confirmation timing in multi-step AI tasks.
- Keywords: Agentic AI, user confirmation, error handling, decision-theoretic model, human-AI interaction, CDCR pattern, confirmation scheduling, task efficiency, user supervision, mixed-initiative systems.
Background and Problem
- Problem / challenge: Current AI agents typically rely on confirm-at-end strategies, which are prone to cascading errors and costly re-execution. Confirming every step avoids these issues but is inefficient. Balancing these extremes remains an unresolved challenge.
- Significance: Errors in long-horizon tasks can lead to significant time, monetary, and environmental costs. Efficiently scheduling user confirmations can enhance task reliability and user experience while reducing these costs.
- Motivation and related work: Prior HCI research has focused on providing user control and error recovery but has rarely addressed when interventions should occur in agent-led workflows. Reliability engineering offers models for inspection timing but focuses on economic costs, not user experience. This paper bridges these gaps by modeling confirmation timing in agentic AI tasks.
Solution
- Proposed approach: A decision-theoretic model for scheduling user confirmations during multi-step AI tasks to minimize task completion time while ensuring correctness.
- Novelty:
- Identification of the Confirmation–Diagnosis–Correction–Redo (CDCR) pattern from a formative study.
- Development of a mathematical model to optimize confirmation timing based on user interaction costs and agent accuracy.
- Empirical validation showing reduced task completion time and strong user preference for intermediate confirmation.
- Framing confirmation scheduling as a mixed-initiative design opportunity for user-supervised systems.
- Procedure and key techniques:
- Conducted a formative study with eight participants to identify user behavior patterns and preferences.
- Developed a stochastic decision process model incorporating state space, action space, and transition functions.
- Designed a dynamic programming algorithm to compute optimal confirmation checkpoints.
- Validated the model through a within-subjects study with 48 participants across three task domains (shopping, image editing, Overcooked).
Results
- Concrete findings:
- Intermediate confirmation reduced task completion time by 13.54% (35.84 seconds on average).
- Confirmation frequency varied by error location, with the largest savings (≈29%) for early errors.
- Intermediate confirmation reduced redo time by 22% and diagnosis time by 17%.
- Advantage over baselines:
- Strong user preference: 81% of participants preferred intermediate confirmation over confirm-at-end.
- Reduced cognitive burden and improved error detection compared to confirm-at-end strategies.
- Comparable or better performance than probability-based confirmation strategies in simulations.
- Experiments / evaluation:
- Within-subjects study with 48 participants across three task domains (shopping, image editing, Overcooked).
- Simulated environment with fixed execution times and error rates to isolate confirmation frequency effects.
- Quantitative metrics (completion time, confirmation, diagnosis, redo) and qualitative feedback (user preferences, perceptions).
- Limitations and future work:
- Current model focuses only on time cost; future work could incorporate real-world costs, safety, and trust calibration.
- Limited personalization of parameters; future systems could adapt dynamically to individual user behavior.
- Need for better alignment with user perceptions of task importance and difficulty.
Summary
This paper introduces a decision-theoretic model for optimizing user confirmation timing in multi-step AI tasks, addressing the trade-off between autonomy and control. A formative study identified the CDCR pattern, which informed the model's design. Validation with 48 participants showed a 13.54% reduction in task completion time and strong user preference (81%) for intermediate confirmation over confirm-at-end strategies. The model offers a scalable solution for enhancing user supervision in agentic AI systems, with potential extensions to incorporate broader cost factors, safety, and trust dynamics.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 86%
Agentic Audio Moderators vs Humans in Think-Aloud Usability Testing
CHI '26· Generative AI (Text, Image, Music, Video) +3
- 75%
Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
IUI '26· Explainable AI (XAI) +4
- 75%
Mapping the Design Space of User Experience for Computer Use Agents
IUI '26· Human-LLM Collaboration +4
- 71%
Researching AI Legibility through Design
CHI '20· Explainable AI (XAI) +2
- 71%
Re-examining Whether, Why, and How Human-AI Interaction Is Uniquely Difficult to Design
CHI '20· Generative AI (Text, Image, Music, Video) +2
- 71%
Effects of LLM-based Search on Decision Making: Speed, Accuracy, and Overreliance
CHI '25· Human-LLM Collaboration +2
- 71%
Do Integral Emotions Affect Trust? The Mediating Effect of Emotions on Trust in the Context of Human-Agent Interaction
DIS '21· Agent Personality & Anthropomorphism +2
- 71%
Teaching-Learning Interaction: A New Concept for Interaction Design to Support Reflective User Agency in Intelligent Systems
DIS '21· Human-LLM Collaboration +2
- 71%
How Users Perceive Mixed-Initiative AI: Attitudes Toward Assistance in Problem Solving
IUI '26· Human-LLM Collaboration +2
- 67%
Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
CHI '21· Explainable AI (XAI) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)