GuideMe: A VLM-Based System Assisting Independent Smartphone Learning for Older Adults

Aging-Friendly Technology DesignHuman-LLM CollaborationAI-Assisted Decision-Making & AutomationBehavior Change & Reflection TechnologyElderly Care WorkersFamily CaregiversAssistive Technology Specialists

Paper Title

GuideMe: A VLM-Based System Assisting Independent Smartphone Learning for Older Adults

Publication Info

  • Topic area: Assistive technology for older adults' independent smartphone learning
  • Keywords: Vision-Language Models, older adults, independent learning, smartphone usability, cognitive load, in-situ highlighting, conversational agents, multimodal context, accessibility, assistive technology

Background and Problem

  • Problem / challenge: Older adults face significant challenges in learning smartphone applications due to cognitive decline, difficulty articulating technical problems, and high visual and cognitive loads when following instructions. Existing independent learning tools and in-person instruction methods are either ineffective or inaccessible.
  • Significance: Addressing these challenges is essential to prevent older adults from being marginalized in the digital world and to enhance their autonomy, self-esteem, and access to critical services.
  • Motivation and related work: Prior research has explored interactive tutorials, AR-based learning, and conversational agents, but these approaches are often application-specific, resource-intensive, or fail to address older adults' unique cognitive and visual challenges. The need for a robust, generalizable, and user-friendly learning support tool remains unmet.

Solution

  • Proposed approach: GuideMe, an on-device conversational instruction system leveraging Vision-Language Models (VLMs) to assist older adults in learning smartphone applications independently through multimodal context analysis, clarifying questions, and in-situ visual guidance.
  • Novelty:
    1. Integration of VLMs to analyze multimodal context and assist in intent clarification.
    2. Use of in-situ highlighting to reduce visual search and cognitive load.
    3. Simulation of effective in-person instruction patterns, including intent confirmation and step-by-step guidance.
  • Procedure and key techniques:
    1. Capture real-time UI context using Android’s Accessibility Service and Media Projection API.
    2. Use GPT-5 to analyze user queries and app context, generating clarifying questions for intent confirmation.
    3. Highlight target UI elements directly on the app interface using a semi-transparent overlay.
    4. Implement fail-safe mechanisms for error recovery, network latency, and unrecognized queries.

Results

  • Concrete findings:
    • GuideMe reduced UI search time (M = 1.26s) compared to an AI search engine (M = 4.23s) and was close to in-person instruction (M = 0.81s).
    • Incorrect clicks were significantly lower with GuideMe (M = 0.11) compared to the search engine (M = 2.94).
    • Task completion time was faster with GuideMe (M = 56.5s) than with the search engine (M = 82.9s).
  • Advantage over baselines:
    • Comparable performance to in-person instruction in usability and cognitive load reduction.
    • Superior to AI search engines in task efficiency, accuracy, and user experience.
  • Experiments / evaluation:
    • Formative study (N=16) to identify older adults' challenges in smartphone learning.
    • User study (N=18) comparing GuideMe with an AI-powered search engine and in-person instruction across six app tasks.
    • Metrics included task completion time, UI search time, incorrect clicks, and subjective ratings (NASA-TLX, UEQ-S).
  • Limitations and future work:
    • Limited to Android devices; privacy concerns with cloud-based VLMs.
    • Conducted in a controlled laboratory setting with Chinese participants; generalizability to other demographics and real-world scenarios needs exploration.
    • Future work includes developing on-device AI models, dynamic transparency mechanisms, and long-term "in-the-wild" deployments.

Summary

GuideMe is a VLM-based system designed to assist older adults in learning smartphone applications independently by simulating effective in-person instruction patterns. Through multimodal context analysis, clarifying questions, and in-situ highlighting, it reduces cognitive and visual loads while improving task efficiency and accuracy. A user study demonstrated its performance comparable to in-person instruction and significantly better than AI search engines. While promising, future work is needed to address privacy concerns, expand platform compatibility, and validate its effectiveness in diverse real-world contexts.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222985/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791448
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Aging-Friendly Technology Design, Human-LLM Collaboration, AI-Assisted Decision-Making & Automation, Behavior Change & Reflection Technology
work
Professions
Elderly Care Workers, Family Caregivers, Assistive Technology Specialists
article
Content Status
Full text indexed
hub
Related Papers
1 related papers