GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
Authors
Minh Duc Vu
CSIRO's Data61Jieshan Chen
CSIRO's Data61Zhenchang Xing
CSIRO's Data61 & Australian National UniversityChunyang Chen
Technical University of MunichVirtual assistants have the potential to play an important role in helping users achieve different tasks. However, these systems face challenges in their real-world usability, characterized by inefficiency and struggles in grasping user intentions. Leveraging recent advances in Large Language Models (LLMs), we introduce GPTVoiceTasker, a virtual assistant poised to enhance user experiences and task efficiency on mobile devices. GPTVoiceTasker excels at intelligently deciphering user commands and executing relevant device interactions to streamline task completion. For unprecedented tasks, GPTVoiceTasker utilises the contextual information and on-screen content to continuously explore and execute the tasks. In addition, the system continually learns from historical user commands to automate subsequent task invocations, further enhancing execution efficiency. From our experiments, GPTVoiceTasker achieved 84.5% accuracy in parsing human commands into executable actions and 85.7% accuracy in automating multi-step tasks. In our user study, GPTVoiceTasker boosted task efficiency in real-world scenarios by 34.85%, accompanied by positive participant feedback. We made GPTVoiceTasker open-source, inviting further research into LLMs utilization for diverse tasks through prompt engineering and leveraging user usage data to improve efficiency.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
MUD: Towards a Large-Scale and Noise-Filtered UI Dataset for Modern Style UI Modeling
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 83%
"If the Machine Is As Good As Me, Then What Use Am I?" – How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and Accomplishment
CHI '24· Human-LLM Collaboration +1
- 83%
How the Role of Generative AI Shapes Perceptions of Value in Human-AI Collaborative Work
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 83%
GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
CHI '25· Generative AI (Text, Image, Music, Video) +2
- 83%
New Enactions of Expertise: Software Engineers’ Evaluation and Demonstration of Coding Expertise with AI Coding Assistants
CHI '26· Human-LLM Collaboration +2
- 83%
From Junior to Senior: Allocating Agency and Navigating Professional Growth in Agentic AI-Mediated Software Engineering
CHI '26· Human-LLM Collaboration +2
- 83%
"It would work for me too": How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools
IUI '25· Human-LLM Collaboration +1
- 83%
Developer Interaction Patterns with Proactive AI: A Five-Day Field Study
IUI '26· AI-Assisted Decision-Making & Automation +2
- 83%
Generative Trigger-Action Programming with Ply
UIST '25· Human-LLM Collaboration +1
- 71%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)