PANDALens: Towards AI-Assisted In-Context Writing on OHMD During Travels
Authors
Generative AI (Text, Image, Music, Video)Human-LLM Collaboration
Title of the Paper
PANDALens: Towards AI-Assisted In-Context Writing on OHMD During Travels
Paper Information
- Research Area: Application of human-computer interaction and augmented reality technologies in travelogue content generation
- Keywords: HMD, smart glasses, artificial intelligence, large language models, multimodal information, human-AI collaborative writing, in-situ writing, travel blogs
Research Background and Problem
- Problem or Challenge: Traditional in-situ writing tools are relatively passive, offering limited intervention in users' experiences and observation behaviors, which leads to lower writing quality and issues with memory decay. Additionally, users' writing interest and attention are often disrupted by unintelligent or poorly designed devices.
- Importance: Writing effectively helps users record personal experiences and share unique feelings, especially during travels, which is crucial for fostering deeper reflection and enhancing travel experiences.
- Research Motivation and Related Work: Existing studies have explored the importance of automatically capturing key moments and reducing subsequent editing workload. However, current systems lack sufficient human-AI collaboration and context-based interaction mechanisms during document generation.
Solution
- Proposed Method: Introducing an active AI-assisted narrative recording system named PANDALens, which integrates optical transparent head-mounted display (OHMD) design and a multimodal information analysis framework. The system uses artificial intelligence to analyze user behavior and the environment in real time, supporting personalized recording and writing.
- Innovations:
- Integration of multimodal information (visual, audio, spatial, and temporal) for interest recognition and context-based writing.
- Utilization of large language models (LLM) to generate relevant questions in real time, encouraging user expression and gradually creating detailed travel narratives.
- Design of a hybrid interaction model, combining human and automated interactions to optimize the writing experience.
- Provision of low-distraction notification design and content generation strategies to enhance users' focus on primary tasks (e.g., traveling).
- Implementation Steps:
- Capturing Moments of Interest: Using OHMD multimodal information tools to detect user behavior, combined with voice input and user gestures to record situational moments.
- Generating Context-Relevant Questions: Using LLM to generate questions based on captured multimodal data, encouraging users to express themselves in depth.
- Draft Writing and Content Editing: Users can select content to include, and the LLM generates high-quality documents based on personalized writing styles.
Research Outcomes
- Specific Results:
- PANDALens increases the number of moments recorded (by 19.2%) and comments made (by 72.6%) during travels through proactive suggestions and real-time questioning, enriching the content.
- Compared to smartphone-based writing tools (LiveSnippets), PANDALens significantly improves language quality, creativity, appeal, and alignment with writing style. Overall user satisfaction scores for writing increased from 60.88 to 82.19 (out of 100).
- Post-writing editing time was reduced by 70.7%.
- Advantages:
- The hybrid interaction model reduces the burden of recording for users compared to tools requiring active user initiation, improving recording quality.
- Integration and utilization of multimodal information significantly enhance the expressiveness and personalization of writing content.
- Distraction management design ensures negligible impact on primary travel activities during the recording process.
- Experimental or Evaluation Results: Tested in real museum scenarios and compared with LiveSnippets, PANDALens demonstrated significant advantages in writing quality, travel experience enhancement, and user satisfaction.
- Limitations and Future Directions:
- Hardware limitations (e.g., current OHMD weight and transparency issues) may affect user experience.
- For more complex scenarios, such as social interactions or noisy environments, more advanced voice recognition and image analysis technologies can be integrated.
- Future work should incorporate more advanced algorithms to further optimize interest recognition and support additional languages and content types (e.g., image and video narratives).
Research Questions / Practical Problems
Question signals indexed for this paper.
help
Research Questions
3- How can optical head-mounted displays (OHMDs) combined with AI technology improve users' travel writing-while-recording experience?Category: Creative Inspiration and Divergent ThinkingSimilar questionsarrow_forward
- How can multimodal information (such as visual, audio, spatial, and temporal) optimize content generation in real-time travel recording?Category: Creative Inspiration and Divergent ThinkingSimilar questionsarrow_forward
- How can a human-AI interaction model be designed to improve user experience of writing-while-recording tools?Category: Creative Inspiration and Divergent ThinkingSimilar questionsarrow_forward
lightbulb
Practical Problems
1- During travel, users struggle to efficiently capture moments and write creative travel documentation.Category: Creative Inspiration and Divergent ThinkingSimilar questionsarrow_forward
- 100%
Exploring User Experiences with Generative AI-Reconstructed Daily Photos
DIS '25· Generative AI (Text, Image, Music, Video) +1
- 67%
Integrating Machine Learning Data with Symbolic Knowledge from Collaboration Practices of Curators to Improve Conversational Systems
CHI '21· Conversational Chatbots +2
- 67%
Rambler: Supporting Writing With Speech via LLM-Assisted Gist Manipulation
CHI '24· Voice User Interface (VUI) Design +2
- 67%
User Experience Design Professionals’ Perceptions of Generative Artificial Intelligence
CHI '24· Generative AI (Text, Image, Music, Video) +1
- 67%
Thoughtful, Confused, or Untrustworthy: How Text Presentation Influences Perceptions of AI Writing Tools
C&C '25· Generative AI (Text, Image, Music, Video) +2
- 67%
How People Prompt Generative AI to Create Interactive VR Scenes
DIS '24· Social & Collaborative VR +2
- 67%
Better Together? An Evaluation of AI-Supported Code Translation
IUI '22· Generative AI (Text, Image, Music, Video) +1
- 67%
Patchview: LLM-powered Worldbuilding with Generative Dust and Magnet Visualization
UIST '24· Generative AI (Text, Image, Music, Video) +1
Based on Jaccard similarity of research subtopics & professions (≥60%)
Quick Actions
AdRecommended
Learn AI Coding at CodeNow
open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642320
At a Glance
fact_checkPaper Snapshot
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Human-LLM Collaboration
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
8 related papers