Title of the Paper

PANDALens: Towards AI-Assisted In-Context Writing on OHMD During Travels

Paper Information

  • Research Area: Application of human-computer interaction and augmented reality technologies in travelogue content generation
  • Keywords: HMD, smart glasses, artificial intelligence, large language models, multimodal information, human-AI collaborative writing, in-situ writing, travel blogs

Research Background and Problem

  • Problem or Challenge: Traditional in-situ writing tools are relatively passive, offering limited intervention in users' experiences and observation behaviors, which leads to lower writing quality and issues with memory decay. Additionally, users' writing interest and attention are often disrupted by unintelligent or poorly designed devices.
  • Importance: Writing effectively helps users record personal experiences and share unique feelings, especially during travels, which is crucial for fostering deeper reflection and enhancing travel experiences.
  • Research Motivation and Related Work: Existing studies have explored the importance of automatically capturing key moments and reducing subsequent editing workload. However, current systems lack sufficient human-AI collaboration and context-based interaction mechanisms during document generation.

Solution

  • Proposed Method: Introducing an active AI-assisted narrative recording system named PANDALens, which integrates optical transparent head-mounted display (OHMD) design and a multimodal information analysis framework. The system uses artificial intelligence to analyze user behavior and the environment in real time, supporting personalized recording and writing.
  • Innovations:
    1. Integration of multimodal information (visual, audio, spatial, and temporal) for interest recognition and context-based writing.
    2. Utilization of large language models (LLM) to generate relevant questions in real time, encouraging user expression and gradually creating detailed travel narratives.
    3. Design of a hybrid interaction model, combining human and automated interactions to optimize the writing experience.
    4. Provision of low-distraction notification design and content generation strategies to enhance users' focus on primary tasks (e.g., traveling).
  • Implementation Steps:
    1. Capturing Moments of Interest: Using OHMD multimodal information tools to detect user behavior, combined with voice input and user gestures to record situational moments.
    2. Generating Context-Relevant Questions: Using LLM to generate questions based on captured multimodal data, encouraging users to express themselves in depth.
    3. Draft Writing and Content Editing: Users can select content to include, and the LLM generates high-quality documents based on personalized writing styles.

Research Outcomes

  • Specific Results:
    1. PANDALens increases the number of moments recorded (by 19.2%) and comments made (by 72.6%) during travels through proactive suggestions and real-time questioning, enriching the content.
    2. Compared to smartphone-based writing tools (LiveSnippets), PANDALens significantly improves language quality, creativity, appeal, and alignment with writing style. Overall user satisfaction scores for writing increased from 60.88 to 82.19 (out of 100).
    3. Post-writing editing time was reduced by 70.7%.
  • Advantages:
    1. The hybrid interaction model reduces the burden of recording for users compared to tools requiring active user initiation, improving recording quality.
    2. Integration and utilization of multimodal information significantly enhance the expressiveness and personalization of writing content.
    3. Distraction management design ensures negligible impact on primary travel activities during the recording process.
  • Experimental or Evaluation Results: Tested in real museum scenarios and compared with LiveSnippets, PANDALens demonstrated significant advantages in writing quality, travel experience enhancement, and user satisfaction.
  • Limitations and Future Directions:
    1. Hardware limitations (e.g., current OHMD weight and transparency issues) may affect user experience.
    2. For more complex scenarios, such as social interactions or noisy environments, more advanced voice recognition and image analysis technologies can be integrated.
    3. Future work should incorporate more advanced algorithms to further optimize interest recognition and support additional languages and content types (e.g., image and video narratives).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/146640/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642320
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
8 related papers