RealityTalk: Real-time Speech-driven Augmented Presentation for AR Live Storytelling
Authors
Title of the Paper
RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live Storytelling
Paper Information
- Domain: Human-Computer Interaction (HCI), Augmented Reality (AR)
- Keywords: augmented reality, mixed reality, augmented presentation, natural language processing, gesture recognition, real-time video processing
Research Background and Problem
-
Identified Problems or Challenges:
- Traditional augmented presentations with embedded visual effects and animations require post-production tools (e.g., Adobe After Effects), which demand significant time and expertise.
- Even recent real-time presentation tools (e.g., Prezi Video and mmhmm.app) are often constrained by linear and scripted limitations, lacking necessary interactivity and improvisation.
- Existing work primarily focuses on gesture-driven or drawing-driven augmented presentations, while speech-driven interactive augmentation techniques remain underexplored.
-
Why It Matters:
- Augmented presentations hold great potential in education, business, and live e-commerce, enhancing audience experience through dynamic information delivery.
- Real-time speech-driven presentations can reduce preparation time and improve interactivity and improvisational capabilities.
-
Motivation and Related Work:
- Existing research mainly focuses on gesture-driven (e.g., Interactive Body-Driven Graphics) or drawing-based augmented presentations, making speech-driven improvisational augmented presentations a novel area.
- An analysis of 177 videos with augmented effects revealed that speech plays a key role in triggering diverse visual and textual elements, highlighting the necessity of further research on speech-driven methods.
Solution
-
Methods and Solutions: RealityTalk is a real-time speech-driven augmented presentation system that generates and manipulates interactive virtual elements during a presentation. Users can interact with augmented text and visual objects through speech and gestures.
-
Innovations:
- Proposes a speech-driven real-time augmented presentation method, offering an important improvement to existing linear presentation approaches.
- Extracts interaction techniques from existing video editing-based augmented presentations and designs a new interaction technique and design space.
- Provides a complete real-time augmented presentation tool from preparation to delivery.
-
Implementation Steps and Key Technologies:
- Interaction Design:
- The system uses speech recognition (WebSpeech API) and natural language processing (Transformer-based NLP with spaCy) to extract keywords from the speech.
- Matches user-defined keywords with visual elements and integrates real-time gesture detection (MediaPipe) for interaction.
- Visual Embedding:
- Implements real-time rendering of 2D and 3D images/videos using React.js, Konva.js, and A-Frame.
- Utilizes 8th Wall for 3D world and object tracking, enabling recognition of horizontal or vertical surfaces and object binding.
- User Interaction:
- Provides a simple keyword matching table, allowing users to pre-associate content and support low-preparation-cost improvisational presentations.
- Offers multiple gesture-based interaction operations, such as position specification, dragging, scaling, rotating, and deletion.
- Interaction Design:
Research Outcomes
-
Specific Outcomes:
- Proposes the RealityTalk system framework, supporting speech-triggered interaction with augmented text and visual elements.
- Extracts interaction behaviors from 177 video examples, defining four core dimensions of augmented presentations: text elements, visual elements, element positioning, and element interaction.
- User evaluations of the system indicate that RealityTalk effectively supports semi-structured improvisational augmented presentations.
-
Advantages:
- Enhances the flexibility and interactivity of real-time presentations.
- Does not require overly complex hardware and can run on standard computers or common webcam setups.
-
Experimental or Evaluation Results:
- User testing with 15 participants showed an average score of 5.4/7 for RealityTalk (usability score: 5.8/7, preparation phase score: 5.7/7).
- System error analysis:
- Gesture interaction errors: Gesture tracking failures occurred an average of 2.36 times, with sensitivity issues in text dragging or scaling interactions.
- Speech recognition errors: Keyword recognition failures (average 0.5 times), including misrecognitions or failure to trigger corresponding elements.
- Element display errors: Issues such as position offsets and duplicate displays.
- The system's average response delay was approximately 1.47 seconds, with keyword-triggered animations generally meeting real-time application requirements.
-
Limitations and Future Directions:
- Limitations:
- The accuracy of speech and gesture recognition is limited in certain cases, particularly with WebSpeech API's compatibility with accents.
- Highly customized presentation functionalities still require further development to support more complex use cases.
- Future Directions:
- Optimize keyword-visual association for more efficient semi-automated suggestions.
- Expand scenarios to support classroom teaching, live-streaming interactions, and even applications in HMDs (head-mounted displays) or spatial projections.
- Develop enhanced interaction features for fully improvisational discussions and collaborative scenarios (e.g., brainstorming).
- Limitations:
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can a voice-driven real-time augmented presentation system interact with augmented text and visual elements?Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
- How do voice-driven methods improve flexibility and improvisation in real-time augmented presentations?Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
- How can interaction patterns from existing video-edited augmented presentations be extracted and applied to real-time voice-driven system design?Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
Practical Problems
1- Traditional augmented presentation preparation is cumbersome, and real-time presentations lack flexibility and interactivity.Category: Multimodal Video Editing Expression and ControlSimilar questionsarrow_forward
- 60%
Location-Aware Adaptation of Augmented Reality Narratives
CHI '23· AR Navigation & Context Awareness +1
- 60%
Examining the Effects of Immersive and Non-Immersive Presenter Modalities on Engagement and Social Interaction in Co-located Augmented Presentations
CHI '25· AR Navigation & Context Awareness +1
- 60%
RealityCanvas: Augmented Reality Sketching for Embedded and Responsive Scribble Animation Effects
UIST '23· AR Navigation & Context Awareness +1
Based on Jaccard similarity of research subtopics & professions (≥60%)