Do It For Me vs. Do It With Me: Investigating User Perceptions of Different Paradigms of Automation in Copilots for Feature-Rich Software
Authors
Research Background and Problem
-
Identified Issues or Challenges:
Users face significant challenges when using software assistants based on large language models (LLMs), such as "copilots." These include misunderstandings of user intent, outputs deviating from expectations, and limited capabilities for task learning and feature exploration. Fully automated assistants ("Do It For Me") lack transparency when completing simple tasks, preventing users from learning software functionalities. Additionally, for feature-rich software, users often prefer to "learn by doing" to familiarize themselves with new features, a need overlooked by fully automated systems. -
Significance:
It is crucial to help users complete complex tasks while enabling them to learn software functionalities and maintain control over tasks. This not only enhances user productivity but also fosters autonomy and engagement with the software. Designing assistive systems that achieve these goals is key to improving human-computer interaction and enhancing user experience. -
Research Motivation and Related Work:
The HCI and AI communities have been exploring ways to optimize LLM-based assistants, but existing systems often focus on technical improvements without adequately addressing user experience and interaction needs. Numerous studies have shown that step-by-step guidance, graphical interface visual references, and adaptive assistance strategies can significantly improve task completion rates, learning outcomes, and trust. The primary motivation of this study is to explore the optimal balance between the "fully automated" (Do It For Me) and "semi-automated guidance" (Do It With Me) design paradigms.
Solution
-
Proposed Method or Solution:
The authors developed two distinct AI assistant prototypes:- AutoCopilot (Fully Automated Assistant), which completes entire tasks based on user input.
- GuidedCopilot (Semi-Automated Assistant), which combines step-by-step guidance, visual support, and automation for repetitive tasks.
-
Innovations:
- GuidedCopilot balances user control and automation by handling only simple tasks via semi-automation and providing step-by-step visual guidance.
- Subsequent design explorations introduced enhanced features, such as task- and state-aware preview clips (GuidedCopilotVisual) and adaptive mixed-media guidance (GuidedCopilotADP), to improve personalized interaction and real-time adaptability.
-
Implementation Steps and Key Technologies:
- Designed two assistant systems:
- AutoCopilot processes tasks fully automatically based on user queries.
- GuidedCopilot triggers step-by-step automation and provides visual reference guidance, such as UI element illustrations.
- Developed structured components, including:
- Text and visual storage modules: Dynamically generate steps using NLP and retrieval-augmented generation (GraphRAG).
- Automation functionality modules: Implemented AppScript in Google Sheets and JavaScript in Figma.
- User interface design: Combined LLM-generated textual guidance with visual illustration outputs.
- Subsequent improvements:
- GuidedCopilotVisual integrated context-specific preview videos into task progress.
- GuidedCopilotADP dynamically adjusted guidance steps based on users' task completion progress.
- Designed two assistant systems:
Research Outcomes
-
Specific Results:
- User Studies and Experiments:
- Conducted comparative experiments with 20 users on the two assistants (tasks in Google Sheets and Figma). Results showed that GuidedCopilot significantly outperformed AutoCopilot in task completion rate (88.5% vs. 35.0%) and accuracy (82.0% vs. 12.0%).
- Follow-up design extension studies (10 participants) indicated that the introduction of GuidedCopilotVisual and GuidedCopilotADP features in Photoshop further enhanced task support and adaptability.
- User Feedback:
- The step-by-step guidance and visual references provided by GuidedCopilot enhanced users' sense of control and facilitated learning of software functionalities.
- AutoCopilot's full automation saved time for simple repetitive tasks but performed poorly for complex or creative tasks.
- User Studies and Experiments:
-
Advantages Over Existing Solutions:
- Semi-automation and step-by-step guidance enable the completion of complex tasks without sacrificing user control.
- The introduction of dynamic, adaptive user guidance addresses the lack of support for learning and control in existing AI assistants.
-
Experimental or Evaluation Results:
Experimental data demonstrated that GuidedCopilot has clear advantages in terms of user learning outcomes, productivity improvement, and support for complex tasks. -
Limitations and Future Directions:
- The study's application scope is primarily limited to Google Sheets and Figma; future work should validate its effectiveness in other feature-rich software.
- The lack of transparency and explainability in fully automated assistants requires further investigation, such as integrating explainable AI (XAI) to enhance understanding of decision-making processes.
- Explore how user needs for different levels of automation change with increased proficiency in long-term usage scenarios.
Conclusion
Through empirical studies and user testing, this research provides an in-depth comparison of the advantages and disadvantages of fully automated and semi-automated design paradigms in feature-rich software environments. The findings offer valuable insights for designing more efficient, controllable, and user-friendly AI assistants. The authors recommend that future assistants balance automation and user learning based on user skill levels and task complexity, while leveraging dynamic visual guidance and adaptive design to further enhance human-computer collaboration efficiency and user experience.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can LLM-driven assistants (e.g., Copilot) help users learn software features while automating task completion?Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
- What behaviors and preferences do users exhibit when using fully automated versus semi-automated assistants?Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
- How can semi-automated assistants improve user experience through dynamic visual guidance and adaptive design?Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
Practical Problems
1- Users struggle to learn software features while using AI assistants to complete tasks and lack a sense of control.Category: LLM Code Generation and Programming AssistantsSimilar questionsarrow_forward
- 100%
From Operation to Cognition: Automatic Modeling Cognitive Dependencies from User Demonstrations for GUI Task Automation
CHI '25· Human-LLM Collaboration +1
- 80%
Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation
CHI '24· Human-LLM Collaboration +1
- 80%
"If the Machine Is As Good As Me, Then What Use Am I?" – How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and Accomplishment
CHI '24· Human-LLM Collaboration +1
- 80%
An Exploratory Study on How AI Awareness Impacts Human-AI Design Collaboration
IUI '25· Human-LLM Collaboration +1
- 80%
"It would work for me too": How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools
IUI '25· Human-LLM Collaboration +1
- 80%
Type, Then Correct: Intelligent Text Correction Techniques for Mobile Text Entry Using Neural Networks
UIST '19· Human-LLM Collaboration +1
- 80%
Generative Trigger-Action Programming with Ply
UIST '25· Human-LLM Collaboration +1
- 75%
Tap&Say: Touch Location-Informed Large Language Model for Multimodal Text Correction on Smartphones
CHI '25· Human-LLM Collaboration
- 67%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 67%
Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
CHI '24· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)