"I Want to Publicize My Stutter'': Community-led Collection and Curation of Chinese Stuttered Speech Data

This paper documents the process undertaken by StammerTalk, a grassroots community of Chinese-speaking people who stutter, to autonomously collect and curate stuttered speech data for more inclusive speech AI models. While people with disabilities are often excluded or treated merely as the subjects of AI data collection, our work introduces a new model for disability data collection in which the disability community exerts agency and control over their personal data and data-driven experiences. Our ethnographic data show that community-led data collection not only produces data needed to represent the community in AI systems, but also empowers the community and its members, by embracing - rather than concealing - stuttering and stutterer identity, and strengthening the social bonds of the community. Recognizing the lack of adequate socio-technical infrastructure for community-led, grassroots data collection, we discuss practical challenges, as well as the strategies and factors for communities to succeed in similar endeavors.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/cscw/178900/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3687014
At a Glance

Paper Snapshot

fact_check
dataset
Source
CSCW
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
—
work
Professions
—
article
Content Status
Abstract only
hub
Related Papers
0 related papers