Choice of Voices: A Large-Scale Evaluation of Text-to-Speech Voice Quality for Long-Form Content

Intelligent Voice Assistants (Alexa, Siri, etc.)Multilingual & Cross-Cultural Voice InteractionVisualization Perception & Cognition

The advancement of text-to-speech (TTS) voices and a rise of commercial TTS platforms allow people to easily experience TTS voices across a variety of technologies, applications, and form factors. As such, we evaluated TTS voices for long-form content: not individual words or sentences, but voices that are pleasant to listen to for several minutes at a time. We introduce a method using a crowdsourcing platform and an online survey to evaluate voices based on listening experience, perception of clarity and quality, and comprehension. We evaluated 18 TTS voices, three human voices, and a text-only control condition. We found that TTS voices are close to rivaling human voices, yet no single voice outperforms the others across all evaluation dimensions. We conclude with considerations for selecting text-to-speech voices for long-form content.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/31906/2020

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3313831.3376789
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2020
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Multilingual & Cross-Cultural Voice Interaction, Visualization Perception & Cognition
work
Professions
—
article
Content Status
Abstract only
hub
Related Papers
0 related papers