LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
Authors
The emergent large language/multimodal models facilitate the evolution of mobile agents, especially in mobile UI task automation. However, existing evaluation approaches, which rely on human validation or established datasets to compare agent-predicted actions with predefined action sequences, are unscalable and unfaithful. To overcome these limitations, this paper presents LlamaTouch, a testbed for on-device mobile UI task execution and faithful, scalable task evaluation. By observing that the task execution process only transfers UI states, LlamaTouch employs a novel evaluation approach that only assesses whether an agent traverses all manually annotated, essential application/system states. LlamaTouch comprises three key techniques: (1) On-device task execution that enables mobile agents to interact with realistic mobile environments for task execution. (2) Fine-grained UI component annotation that merges pixel-level screenshots and textual screen hierarchies to explicitly identify and precisely annotate essential UI components with a rich set of designed annotation primitives. (3) A multi-level application state matching algorithm that utilizes exact and fuzzy matching to accurately detect critical information in each screen, even with unpredictable UI layout/content dynamics. LlamaTouch currently incorporates four mobile agents and 496 tasks, encompassing both tasks in the widely-used datasets and our self-constructed ones to cover more diverse mobile applications. Evaluation results demonstrate LlamaTouch’s high faithfulness of evaluation in real-world mobile environments and its better scalability than human validation. LlamaTouch also enables easy task annotation and integration of new mobile agents. Code and dataset are publicly available at https://github.com/LlamaTouch/LlamaTouch.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 100%
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
UbiComp '24· Human-LLM Collaboration
- 75%
AffectiveSpotlight: Facilitating the Communication of Affective Responses from Audience Members during Online Presentations
CHI '21· Social & Collaborative VR +1
- 75%
Evaluating Large Language Models in Generating Synthetic HCI Research Data: a Case Study
CHI '23· Human-LLM Collaboration +1
- 75%
DynEx: Dynamic Code Synthesis with Structured Design Exploration for Accelerated Exploratory Programming
CHI '25· Human-LLM Collaboration +1
- 75%
SummAct: Uncovering User Intentions Through Interactive Behaviour Summarisation
CHI '25· Human-LLM Collaboration +1
- 75%
Towards Rapid Interactive Machine Learning: Evaluating Tradeoffs of Classification without Representation
IUI '19· Human-LLM Collaboration +1
- 75%
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
IUI '25· Human-LLM Collaboration +1
- 75%
Assuage: Assembly Synthesis Using A Guided Exploration
UIST '21· Human-LLM Collaboration +1
- 75%
Shapir: Standardizing and Democratizing Access to Web APIs
UIST '21· Human-LLM Collaboration +1
- 75%
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
UIST '24· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)