DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
Authors
Paper Title
DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
Publication Info
- Topic area: Mobile information retrieval and automation using multi-LLM systems.
- Keywords: Mobile agents, multi-LLM collaboration, information seeking, task decomposition, UI automation, report synthesis, transparency, user intervention, privacy-aware systems.
Background and Problem
- Problem / challenge: Mobile information seeking is fragmented due to isolated app ecosystems, requiring users to manually switch contexts and re-enter data. Existing solutions, such as LLM-driven agents and API-based tools, lack transparency, cross-source integration, and effective user intervention mechanisms.
- Significance: Addressing these challenges can reduce cognitive load, improve workflow efficiency, and enhance user trust in mobile automation systems.
- Motivation and related work: Prior work includes LLM-driven web search agents and mobile task-execution systems, which struggle with login-gated content, dynamic in-app data, and transparency. Existing tools lack structured progress monitoring, verifiable evidence, and effective intervention mechanisms, leaving gaps in achieving reliable and user-controllable mobile information retrieval.
Solution
- Proposed approach: DroidRetriever, a multi-LLM-based system for transparent and steerable mobile information seeking.
- Novelty:
- A multi-LLM architecture with task decomposition, UI copilot, and report synthesis modules.
- A transparent dashboard for real-time task progress and exploration traces, enabling user intervention.
- Privacy-aware mechanisms that pause on high-risk screens and prompt user confirmation.
- Citation-linked reports for verifiable and structured information synthesis.
- Procedure and key techniques:
- Task Decomposition: Breaks tasks into app-specific sub-tasks, selects relevant apps, and determines search modes (focused, list-view, or multi-page).
- UI Copilot: Automates navigation using vision-based UI comprehension, action planning, and scrolling screenshots, with proactive pausing for privacy-sensitive actions.
- Report Synthesis: Processes screenshots to generate structured reports (e.g., tables, summaries) with citation-linked evidence for traceability.
- User Intervention: Provides a dashboard and widget for task monitoring, manual corrections, and takeover during navigation.
Results
- Concrete findings:
- Achieved 96% app-level decomposition accuracy and 82% page-level decomposition accuracy across 22 tasks.
- Demonstrated higher report coverage (p = .0096) and comparable accuracy (p = .77) to human participants in Study 1.
- Reduced user workload and improved perceived certainty and confidence in multi-app tasks compared to manual and baseline systems.
- Advantage over baselines:
- Outperformed LLM-driven search engines (Qwen, ChatGPT) and mobile agents (Claude, Mobile-Agent-v2) in coverage, accuracy, and user trust.
- Provided better transparency and control through progress visualization and citation-linked reports.
- Reduced token usage (19,459 tokens/task) compared to Claude (210,128 tokens/task).
- Experiments / evaluation:
- Study 1: Compared DroidRetriever’s report synthesis against human performance on 13 tasks, evaluating coverage, accuracy, and redundancy.
- Study 2: Assessed task decomposition, UI copilot, and overall system performance on 22 single-app and multi-app tasks, including comparisons with baseline tools.
- Study 3: Conducted qualitative interviews and case studies to evaluate user experience, trust, and intervention mechanisms.
- Limitations and future work:
- Struggles with dynamic content (e.g., video streams) and sudden screen changes.
- Relies on generalized LLMs, leading to latency and occasional inaccuracies.
- Future work includes enhancing support for dynamic interfaces and exploring multi-modal models for faster, end-to-end navigation.
Summary
DroidRetriever is a transparent and steerable mobile information retrieval system leveraging multi-LLM collaboration. It automates task decomposition, UI navigation, and report synthesis, providing real-time progress dashboards and citation-linked reports for user verification. Experiments demonstrated its superior performance in coverage, accuracy, and user trust compared to existing tools. While effective in addressing fragmented mobile information seeking, future work aims to improve responsiveness to dynamic content and optimize LLM-based navigation. The system is well-suited for everyday, cross-app tasks requiring actionable and verifiable information.
Research Questions / Practical Problems
Question signals indexed for this paper.
- 71%
LLM-box vs. Thinking-box: Designing for Deliberate User Engagement with Distorted Information in Conversational Search
CHI '26· Human-LLM Collaboration +2
- 71%
A New Taxonomy of Web Search: A User-Centered Framework for Search Intent in the AI Era
CHI '26· Exploratory Search & Information Seeking +2
- 67%
Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences
UIST '24· Human-LLM Collaboration
Based on Jaccard similarity of research subtopics & professions (≥60%)