RealityTalk: Real-time Speech-driven Augmented Presentation for AR Live Storytelling

AR Navigation & Context AwarenessInteractive Narrative & Immersive StorytellingUI/UX DesignersDancers & Performing Artists

Title of the Paper

RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live Storytelling

Paper Information

  • Domain: Human-Computer Interaction (HCI), Augmented Reality (AR)
  • Keywords: augmented reality, mixed reality, augmented presentation, natural language processing, gesture recognition, real-time video processing

Research Background and Problem

  • Identified Problems or Challenges:

    1. Traditional augmented presentations with embedded visual effects and animations require post-production tools (e.g., Adobe After Effects), which demand significant time and expertise.
    2. Even recent real-time presentation tools (e.g., Prezi Video and mmhmm.app) are often constrained by linear and scripted limitations, lacking necessary interactivity and improvisation.
    3. Existing work primarily focuses on gesture-driven or drawing-driven augmented presentations, while speech-driven interactive augmentation techniques remain underexplored.
  • Why It Matters:

    • Augmented presentations hold great potential in education, business, and live e-commerce, enhancing audience experience through dynamic information delivery.
    • Real-time speech-driven presentations can reduce preparation time and improve interactivity and improvisational capabilities.
  • Motivation and Related Work:

    • Existing research mainly focuses on gesture-driven (e.g., Interactive Body-Driven Graphics) or drawing-based augmented presentations, making speech-driven improvisational augmented presentations a novel area.
    • An analysis of 177 videos with augmented effects revealed that speech plays a key role in triggering diverse visual and textual elements, highlighting the necessity of further research on speech-driven methods.

Solution

  • Methods and Solutions: RealityTalk is a real-time speech-driven augmented presentation system that generates and manipulates interactive virtual elements during a presentation. Users can interact with augmented text and visual objects through speech and gestures.

  • Innovations:

    1. Proposes a speech-driven real-time augmented presentation method, offering an important improvement to existing linear presentation approaches.
    2. Extracts interaction techniques from existing video editing-based augmented presentations and designs a new interaction technique and design space.
    3. Provides a complete real-time augmented presentation tool from preparation to delivery.
  • Implementation Steps and Key Technologies:

    1. Interaction Design:
      • The system uses speech recognition (WebSpeech API) and natural language processing (Transformer-based NLP with spaCy) to extract keywords from the speech.
      • Matches user-defined keywords with visual elements and integrates real-time gesture detection (MediaPipe) for interaction.
    2. Visual Embedding:
      • Implements real-time rendering of 2D and 3D images/videos using React.js, Konva.js, and A-Frame.
      • Utilizes 8th Wall for 3D world and object tracking, enabling recognition of horizontal or vertical surfaces and object binding.
    3. User Interaction:
      • Provides a simple keyword matching table, allowing users to pre-associate content and support low-preparation-cost improvisational presentations.
      • Offers multiple gesture-based interaction operations, such as position specification, dragging, scaling, rotating, and deletion.

Research Outcomes

  • Specific Outcomes:

    1. Proposes the RealityTalk system framework, supporting speech-triggered interaction with augmented text and visual elements.
    2. Extracts interaction behaviors from 177 video examples, defining four core dimensions of augmented presentations: text elements, visual elements, element positioning, and element interaction.
    3. User evaluations of the system indicate that RealityTalk effectively supports semi-structured improvisational augmented presentations.
  • Advantages:

    • Enhances the flexibility and interactivity of real-time presentations.
    • Does not require overly complex hardware and can run on standard computers or common webcam setups.
  • Experimental or Evaluation Results:

    1. User testing with 15 participants showed an average score of 5.4/7 for RealityTalk (usability score: 5.8/7, preparation phase score: 5.7/7).
    2. System error analysis:
      • Gesture interaction errors: Gesture tracking failures occurred an average of 2.36 times, with sensitivity issues in text dragging or scaling interactions.
      • Speech recognition errors: Keyword recognition failures (average 0.5 times), including misrecognitions or failure to trigger corresponding elements.
      • Element display errors: Issues such as position offsets and duplicate displays.
    3. The system's average response delay was approximately 1.47 seconds, with keyword-triggered animations generally meeting real-time application requirements.
  • Limitations and Future Directions:

    1. Limitations:
      • The accuracy of speech and gesture recognition is limited in certain cases, particularly with WebSpeech API's compatibility with accents.
      • Highly customized presentation functionalities still require further development to support more complex use cases.
    2. Future Directions:
      • Optimize keyword-visual association for more efficient semi-automated suggestions.
      • Expand scenarios to support classroom teaching, live-streaming interactions, and even applications in HMDs (head-mounted displays) or spatial projections.
      • Develop enhanced interaction features for fully improvisational discussions and collaborative scenarios (e.g., brainstorming).

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/85010/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3526113.3545702
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
AR Navigation & Context Awareness, Interactive Narrative & Immersive Storytelling
work
Professions
UI/UX Designers, Dancers & Performing Artists
article
Content Status
Full text indexed
hub
Related Papers
3 related papers