Unblind Text Inputs: Predicting Hint-text of Text Input in Mobile Apps via LLM
Honorable MentionAuthors
Title of the Paper
Predicting Text Input Hint Text in Mobile Applications Using Large Language Models (LLM): Design and Evaluation of HintDroid
Paper Information
- Research Area: Human-Computer Interaction, Assistive Technology, Applications of Large Language Models
- Keywords: Mobile Application Design, Accessibility, User Interface, Large Language Models, Text Input, Feedback Mechanism, Visually Impaired Users
Research Background and Issues
-
Problems and Challenges:
- Accessibility issues are prevalent in mobile applications, particularly the lack of hint text in text input components.
- An analysis of 4,501 Android applications revealed that over 76% of text input components lack hint text, creating barriers for visually impaired users relying on screen readers to understand input requirements.
- The absence of hint text is often due to developers' lack of awareness about accessibility or insufficient information descriptions.
- Existing research primarily focuses on generating accessibility labels for image components, with little attention to the automatic generation of text input hint text.
-
Significance:
- According to the World Health Organization, at least 2.2 billion people worldwide face vision impairment, necessitating more inclusive mobile application designs.
- The lack of hint text prevents screen readers from conveying the semantic information of input fields, significantly hindering the interaction and experience of visually impaired users.
- Generating appropriate hint text can enhance the overall accessibility of mobile applications, providing better functional support for visually impaired users.
-
Motivation and Related Work:
- Large Language Models (LLMs) have demonstrated outstanding performance in natural language processing and generation tasks, inspiring their application in accessibility-related text generation tasks.
- Very few studies have addressed semantic understanding and label generation for software interfaces, and no existing methods are tailored for generating text input hint text.
- HintDroid is proposed as a novel solution in this research domain to address these challenges.
Solution
-
Method or Solution:
- A system named HintDroid is proposed, which combines LLMs with graphical user interface (GUI) information to generate hint text for text inputs through contextual learning.
- A feedback mechanism is introduced to optimize the generated hint text by determining whether the generated input content meets the requirements and whether it triggers the next page, providing feedback to refine the results.
-
Innovations:
- A contextual learning approach based on LLMs transforms user interface information into comprehensible natural language descriptions.
- A feedback mechanism (including error information extraction) is designed to evaluate and improve the quality of the generated hint text.
- Hint text generation is integrated with input content generation to validate its relevance and improve outcomes.
-
Implementation Steps:
- GUI Entity Extraction: Extract and organize information about the application, current page, and input components, including component text, IDs, relative positions, etc.
- Hint Generation: Use contextual learning to construct prompts with examples to help the LLM understand the task.
- Feedback Optimization: Evaluate the quality of hint text by detecting whether the input content triggers the next page and refine the generated results based on error information.
- Automated Experiments and Evaluation: Generate hint text using LLMs and validate its effectiveness through natural language evaluation metrics such as BLEU and ROUGE.
Research Outcomes
-
Specific Results:
- HintDroid achieved leading performance in large-scale experiments, with BLEU@1 reaching 83%, ROUGE scoring 63%, and CIDEr scoring 62%.
- User studies showed that after using HintDroid, the rate of visually impaired users correctly filling in input content increased by 152%, while the coverage of exploring more page activities and states increased by 77% and 66%, respectively.
- The average time for generating hint text was 1.86 seconds per page, significantly improving development efficiency.
-
Advantages Compared to Existing Solutions:
- Significantly outperforms existing methods, including rule-based, deep learning, and current LLM-based approaches, in terms of hint text generation accuracy and coverage.
- Provides a feedback-based optimization mechanism, making it adaptable to diverse application scenarios.
- Combines hint text generation with content generation, ensuring the practicality of the hint text and improving user experience.
-
Experimental or Evaluation Results:
- Large-scale comparative experiments showed that HintDroid outperformed the best existing methods by over 82%.
- User experiment feedback indicated that using HintDroid reduced input completion time by 139%.
-
Limitations and Future Directions:
- Limitations in generation accuracy: Hint text generation quality is affected when GUI contextual information is insufficient.
- Currently limited to the Android platform; future work could expand to iOS, web, and other platforms.
- Plans to enhance the model's contextual understanding capabilities and explore real-time interaction implementation.
- Users expressed a desire for more personalized and dynamically adjustable hints, such as providing input format or interactive feedback support.
Conclusion
HintDroid leverages the powerful natural language understanding capabilities of large language models, combined with contextual learning and feedback mechanisms, to effectively bridge the information gap for visually impaired users in mobile applications by providing accessible text input hints. Future work includes improving scenario adaptability, achieving cross-platform expansion, and incorporating personalized design features.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can LLMs automatically generate hint text for mobile app text input components?Category: Screen Reader and Interface AccessibilitySimilar questionsarrow_forward
- How can hint text generation incorporate GUI context to improve accuracy and relevance?Category: Screen Reader and Interface AccessibilitySimilar questionsarrow_forward
- How effective is optimizing generated hint text with feedback mechanisms?Category: Screen Reader and Interface AccessibilitySimilar questionsarrow_forward
Practical Problems
1- Blind users struggle to understand input requirements in apps via screen readers.Category: Screen Reader and Interface AccessibilitySimilar questionsarrow_forward
- 67%
Generative and Malleable User Interfaces with Generative and Evolving Task-Driven Data Model
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers
CHI '25· Generative AI (Text, Image, Music, Video) +1
- 67%
GeneyMAP: Exploring the Potential of GenAI to Facilitate Mapping User Journeys for UX Design
CHI '25· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)