“I Won't Go Speechless”: Design Exploration on a Real-Time Text-To-Speech Speaking Tool for Videoconferencing

Privacy by Design & User ControlHome Voice Assistant Experience

Title of the Paper

"I Won't Go Speechless": Design Exploration on a Real-Time Text-To-Speech Speaking Tool for Videoconferencing

Bibliographic Information

  • Subject Area: Human-Computer Interaction (HCI), Remote Collaboration, and Videoconferencing
  • Keywords: Videoconferencing, Text-To-Speech (TTS), Augmentative and Alternative Communication (AAC), User Experience, Real-Time Communication, Voice Interaction

Research Background and Problem

  • Problem Identification:

    1. With the widespread adoption of videoconferencing during COVID-19, it has become a key communication tool for remote work. However, environments unsuitable for voice communication (e.g., noisy cafes, quiet libraries, or shared spaces with family) limit user participation.
    2. In such settings, users often refrain from speaking due to unavailable microphones or concerns about noise and privacy issues.
    3. While text chat can serve as an alternative, existing studies indicate that text chat disrupts the flow of meetings and is less effective than voice-based communication in multi-party meetings.
  • Significance: In the post-COVID-19 era, as people participate in videoconferences from increasingly diverse environments, designing new tools to support users with limited voice communication becomes crucial.

  • Research Motivation and Related Work:

    1. Augmentative and Alternative Communication (AAC) systems are primarily designed for individuals with speech impairments or disabilities, with limited attention to the contextual needs of other users.
    2. Advances in Text-To-Speech (TTS) technology (e.g., naturalness of voice generation and real-time capabilities) have created potential for applications in various scenarios. However, there has been relatively little development and user research on its application in multi-party videoconferencing contexts.
    3. The authors aim to explore how TTS-based real-time communication tools can help videoconferencing users overcome barriers in voice-restricted environments.

Solution

  • Research Methods:

    1. Utilizing the Technology Probe method, the authors conducted two field studies to observe user behavior during videoconferences.
    2. Designed a TTS tool prototype and conducted user experiments to explore experiences in multi-party meeting scenarios.
  • Features of the TTS Tool:

    1. Converts user-inputted text into synthesized speech in real time and plays it for other meeting participants.
    2. Offers additional features such as word-by-word speech playback and adjustable speech rate and pitch.
  • Innovations:

    1. Positions TTS as a tool to support real-time participation, targeting not only traditional AAC users (e.g., individuals with speech impairments) but also ordinary users in contextually restricted environments.
    2. Emphasizes a user experience-driven design approach, exploring needs based on real-world usage scenarios and contextual constraints.
  • Implementation Steps:

    1. Observed user performance in videoconferences using the TTS tool during Study 1 and Study 2.
    2. Collected meeting content, user behavior videos, and interview records to analyze the tool's feasibility, advantages, and user feedback.

Research Outcomes

  • Key Results and Findings:

    1. Compared to text chat, the TTS tool significantly enhances users' sense of participation and presence in voice-restricted scenarios.
    2. Users generally recognized the value of the TTS tool but highlighted key areas for improvement in both technology and interaction design:
      • Voice naturalness: Better expression of emotions and tone is needed.
      • Real-time performance: Slight delays in voice generation affect the fluency of communication.
      • User interface: More intuitive operation processes are required, such as shortcut key support and simplified multitasking interfaces.
  • User Needs and Challenges:

    1. Users felt restricted when expressing attitudes, tones, and auxiliary language such as laughter using the TTS tool.
    2. Users desired quick access to frequently used phrases and voice control options (e.g., speech rate and pitch).
    3. It was challenging for recipients to identify the speaker, especially when multiple participants used TTS simultaneously.
  • Comparison with Existing Solutions:

    1. Compared to text chat, the TTS tool offers stronger synchronization and a greater sense of participation.
    2. While existing TTS technologies focus on one-way content (e.g., screen reading), this study expands its scope by exploring its potential in real-time interpersonal communication.
  • Experimental Data and User Feedback:

    1. In Study 1, experiments showed that participants using the TTS tool scored significantly higher in contribution and participation compared to those using only text.
    2. Study 2 emphasized the importance of customizability and interface intuitiveness in enhancing user experience.
  • Future Directions for Improvement:

    1. Enhance the real-time dynamics and expressiveness of voice generation, such as developing more natural speech synthesizers.
    2. Provide more automated interface designs to reduce user operational burden.
    3. Expand research scenarios to include more heterogeneous user groups and meeting formats (e.g., hybrid meetings involving more speakers).
    4. Explore potential privacy and ethical risks associated with TTS based on users' real voices.

Conclusion

This paper explores the value of real-time TTS tools in voice-restricted videoconferencing scenarios, demonstrating their significant potential and research implications. Future studies should delve deeper into the design considerations for specific scenarios to provide users with more personalized and efficient voice interaction solutions.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/96164/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581215
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Privacy by Design & User Control, Home Voice Assistant Experience
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
1 related papers