ARTiST: Automated Text Simplification for Task Guidance in Augmented Reality
Authors
AR Navigation & Context AwarenessHuman-LLM CollaborationAlgorithmic Transparency & AuditabilityUniversity Professors & ResearchersUI/UX Designers
Title of the Paper
ARTiST: Automated Text Simplification for Task Guidance in Augmented Reality
Paper Information
- Domain: Text simplification and task guidance in Augmented Reality (AR)
- Keywords: Augmented Reality, Text Simplification, Large Language Models, User Experience, Human-Computer Interaction
Research Background and Problem Statement
- Challenges: Using text-based task guidance on AR devices (e.g., Head-Mounted Displays, HMDs) presents the following issues:
- Limited display area on head-mounted devices, where lengthy text can obstruct the user's view and increase cognitive load.
- Existing text simplification methods are primarily designed for individuals with low reading proficiency and are not tailored for AR environments, lacking consideration of physical tasks and contextual scenarios.
- Current textual prompts often fail to incorporate spatial information or provide clear task details.
- Significance:
- Enhancing the readability and conciseness of textual task guidance can significantly improve users' efficiency during real-world tasks.
- Improving user experience and reducing cognitive load is critical to expanding the practical applications of AR in industries such as manufacturing, healthcare, and education.
- Research Motivation:
- To address the specific needs of AR, develop an automated text simplification system that adapts to the constraints of AR environments.
- Bridge the gap in traditional text simplification methods, which lack specificity for AR applications.
Proposed Solution
- Method/System: The study introduces a text simplification system named ARTiST, which leverages OpenAI's GPT-3 large language model and employs the following innovative techniques to optimize text-based task guidance in AR environments:
- Plan-of-Technique: Guides the language model to sequentially perform content simplification, syntactic simplification, lexical substitution, and spatial information augmentation in multiple steps.
- Chain-of-Thought Method: Enhances multi-step reasoning capabilities, reducing potential conflicts between different simplification methods.
- Error-Aware Calibration: Corrects potential biases in the model's output to ensure the simplified text retains the original semantics and adheres to design principles.
- Spatial Information-Based Fine-Grained Simplification: Collects spatial information about objects in the scene and integrates it into the textual guidance.
- Innovations:
- Customization for AR-specific requirements, including minimizing text occlusion within the display area and optimizing the reading experience.
- Achieves precise text generation through few-shot learning without requiring extensive labeled datasets.
- Combines semantic simplification with spatial information augmentation, achieving the first fully automated text optimization tailored for AR.
- Technology and Implementation Steps: The system relies on GPT-3, using a small number of examples to construct the training dataset, and integrates Detic for real-time spatial object localization. A backend service is built using Flask to push the simplified text results to HoloLens 2 AR devices.
Research Outcomes
- Key Results:
- Developed the ARTiST system—a GPT-3-based automated text simplification framework specifically designed for AR.
- Adhered to technical guidelines informed by user and expert feedback to support accurate and interactive AR text interfaces.
- Validated through user experiments that ARTiST significantly reduces cognitive load and operational errors.
- Experiment Summary:
- The experiments consisted of two parts: measuring task completion efficiency, user cognitive metrics (NASA TLX), and text recall accuracy.
- Tests with 16 participants demonstrated that ARTiST significantly reduced error rates and improved the accuracy of recalling task steps compared to unoptimized text and traditional simplification methods (e.g., T-5 model).
- Simplified text received higher subjective ratings in readability, task guidance, and trustworthiness.
- Advantages Compared to Existing Methods:
- Unlike traditional text simplification methods, ARTiST systematically considers AR display characteristics (e.g., field-of-view limitations).
- Guided by the current GPT-3 model, ARTiST balances dynamic information addition with spatial expression simplification, showcasing the flexibility of the system design.
- Limitations and Future Directions:
- Limitations:
- Performance in scenarios involving continuous movement or complex interactions (e.g., surgeries) remains unverified.
- The system has primarily been tested on head-mounted display devices and has not yet been extended to other AR devices or application scenarios.
- Technical issues such as API request delays may lead to display update latency.
- Future Directions:
- Explore the potential for combining text simplification with multimedia elements (e.g., images, audio) for complex tasks.
- Enhance research on personalized configuration capabilities, enabling the system to better adapt to different users' preferences and background knowledge.
- Improve support for dynamic task scenarios (e.g., mobile contexts), including enhancing real-time responsiveness.
- Limitations:
Further Open-Source Contributions
- The authors have provided the implementation code, which is publicly available at ARTiST Code Repository, allowing the academic community to improve and replicate the system.
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can task instruction text on AR devices be simplified to reduce users' cognitive burden and operational errors?Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
- How can text simplification systems for AR environments combine spatial information to optimize presentation?Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
- In AR scenarios, how can semantic preservation in text simplification be balanced with effective subsequent task guidance?Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
lightbulb
Practical Problems
1- Lengthy text on AR devices obstructs the field of view and increases users' cognitive burden.Category: AR Task Guidance, Tutorials, and Prompt DesignSimilar questionsarrow_forward
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642772
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
AR Navigation & Context Awareness, Human-LLM Collaboration, Algorithmic Transparency & Auditability
work
Professions
University Professors & Researchers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers