ARTiST: Automated Text Simplification for Task Guidance in Augmented Reality

AR Navigation & Context AwarenessHuman-LLM CollaborationAlgorithmic Transparency & AuditabilityUniversity Professors & ResearchersUI/UX Designers

Title of the Paper

ARTiST: Automated Text Simplification for Task Guidance in Augmented Reality

Paper Information

  • Domain: Text simplification and task guidance in Augmented Reality (AR)
  • Keywords: Augmented Reality, Text Simplification, Large Language Models, User Experience, Human-Computer Interaction

Research Background and Problem Statement

  • Challenges: Using text-based task guidance on AR devices (e.g., Head-Mounted Displays, HMDs) presents the following issues:
    • Limited display area on head-mounted devices, where lengthy text can obstruct the user's view and increase cognitive load.
    • Existing text simplification methods are primarily designed for individuals with low reading proficiency and are not tailored for AR environments, lacking consideration of physical tasks and contextual scenarios.
    • Current textual prompts often fail to incorporate spatial information or provide clear task details.
  • Significance:
    • Enhancing the readability and conciseness of textual task guidance can significantly improve users' efficiency during real-world tasks.
    • Improving user experience and reducing cognitive load is critical to expanding the practical applications of AR in industries such as manufacturing, healthcare, and education.
  • Research Motivation:
    • To address the specific needs of AR, develop an automated text simplification system that adapts to the constraints of AR environments.
    • Bridge the gap in traditional text simplification methods, which lack specificity for AR applications.

Proposed Solution

  • Method/System: The study introduces a text simplification system named ARTiST, which leverages OpenAI's GPT-3 large language model and employs the following innovative techniques to optimize text-based task guidance in AR environments:
    1. Plan-of-Technique: Guides the language model to sequentially perform content simplification, syntactic simplification, lexical substitution, and spatial information augmentation in multiple steps.
    2. Chain-of-Thought Method: Enhances multi-step reasoning capabilities, reducing potential conflicts between different simplification methods.
    3. Error-Aware Calibration: Corrects potential biases in the model's output to ensure the simplified text retains the original semantics and adheres to design principles.
    4. Spatial Information-Based Fine-Grained Simplification: Collects spatial information about objects in the scene and integrates it into the textual guidance.
  • Innovations:
    • Customization for AR-specific requirements, including minimizing text occlusion within the display area and optimizing the reading experience.
    • Achieves precise text generation through few-shot learning without requiring extensive labeled datasets.
    • Combines semantic simplification with spatial information augmentation, achieving the first fully automated text optimization tailored for AR.
  • Technology and Implementation Steps: The system relies on GPT-3, using a small number of examples to construct the training dataset, and integrates Detic for real-time spatial object localization. A backend service is built using Flask to push the simplified text results to HoloLens 2 AR devices.

Research Outcomes

  • Key Results:
    1. Developed the ARTiST system—a GPT-3-based automated text simplification framework specifically designed for AR.
    2. Adhered to technical guidelines informed by user and expert feedback to support accurate and interactive AR text interfaces.
    3. Validated through user experiments that ARTiST significantly reduces cognitive load and operational errors.
  • Experiment Summary:
    • The experiments consisted of two parts: measuring task completion efficiency, user cognitive metrics (NASA TLX), and text recall accuracy.
    • Tests with 16 participants demonstrated that ARTiST significantly reduced error rates and improved the accuracy of recalling task steps compared to unoptimized text and traditional simplification methods (e.g., T-5 model).
    • Simplified text received higher subjective ratings in readability, task guidance, and trustworthiness.
  • Advantages Compared to Existing Methods:
    • Unlike traditional text simplification methods, ARTiST systematically considers AR display characteristics (e.g., field-of-view limitations).
    • Guided by the current GPT-3 model, ARTiST balances dynamic information addition with spatial expression simplification, showcasing the flexibility of the system design.
  • Limitations and Future Directions:
    • Limitations:
      • Performance in scenarios involving continuous movement or complex interactions (e.g., surgeries) remains unverified.
      • The system has primarily been tested on head-mounted display devices and has not yet been extended to other AR devices or application scenarios.
      • Technical issues such as API request delays may lead to display update latency.
    • Future Directions:
      • Explore the potential for combining text simplification with multimedia elements (e.g., images, audio) for complex tasks.
      • Enhance research on personalized configuration capabilities, enabling the system to better adapt to different users' preferences and background knowledge.
      • Improve support for dynamic task scenarios (e.g., mobile contexts), including enhancing real-time responsiveness.

Further Open-Source Contributions

  • The authors have provided the implementation code, which is publicly available at ARTiST Code Repository, allowing the academic community to improve and replicate the system.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147808/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642772
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
AR Navigation & Context Awareness, Human-LLM Collaboration, Algorithmic Transparency & Auditability
work
Professions
University Professors & Researchers, UI/UX Designers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers