VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
Authors
Mobile task automation is an emerging field that leverages AI to streamline and optimize the execution of routine tasks on mobile devices, thereby enhancing efficiency and productivity. Traditional methods, such as Programming By Demonstration (PBD), are limited due to their dependence on predefined tasks and susceptibility to app updates. Recent advancements have utilized the view hierarchy to collect UI information and employed Large Language Models (LLM) to enhance task automation. However, view hierarchies have accessibility issues and face potential problems like missing object descriptions or misaligned structures. This paper introduces VisionTasker, a two-stage framework combining vision-based UI understanding and LLM task planning, for mobile task automation in a step-by-step manner. VisionTasker firstly converts a UI screenshot into natural language interpretations using a vision-based UI understanding approach, eliminating the need for view hierarchies. Secondly, it adopts a step-by-step task planning method, presenting one interface at a time to the LLM. The LLM then identifies relevant elements within the interface and determines the next action, enhancing accuracy and practicality. Extensive experiments show that VisionTasker outperforms previous methods, providing effective UI representations across four datasets. Additionally, in automating 147 real-world tasks on an Android smartphone, VisionTasker demonstrates advantages over humans in tasks where humans show unfamiliarity and shows significant improvements when integrated with the PBD mechanism. VisionTasker is open-source and available at https://github.com/AkimotoAyako/VisionTasker.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 83%
"If the Machine Is As Good As Me, Then What Use Am I?" – How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and Accomplishment
CHI '24· Human-LLM Collaboration +1
- 83%
"It would work for me too": How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools
IUI '25· Human-LLM Collaboration +1
- 83%
Generative Trigger-Action Programming with Ply
UIST '25· Human-LLM Collaboration +1
- 71%
Adapting User Interfaces with Model-based Reinforcement Learning
CHI '21· Human-LLM Collaboration +2
- 71%
ChainBuddy: An AI-assisted Agent System for Generating LLM Pipelines
CHI '25· Human-LLM Collaboration +1
- 71%
Assistance or Disruption? Exploring and Evaluating the Design and Trade-offs of Proactive AI Programming Support
CHI '25· Human-LLM Collaboration +2
- 71%
PointAloud: An Interaction Suite for AI-Supported Pointer-Centric Think-Aloud Computing
CHI '26· Human-LLM Collaboration +2
- 71%
Situated, Dynamic, and Subjective: Envisioning the Design of Theory-of-Mind-Enabled Everyday AI with Industry Practitioners
CHI '26· Brain-Computer Interface (BCI) & Neurofeedback +2
- 71%
State Your Intention to Steer Your Attention: An AI Assistant for Intentional Digital Living
CHI '26· Human-LLM Collaboration +2
- 71%
ChoiceMates: Supporting Unfamiliar Online Decision-Making with Multi-Agent Conversational Interactions
IUI '26· Human-LLM Collaboration +2
Based on Jaccard similarity of research subtopics & professions (≥60%)