Research Background and Issues

  • What problems or challenges did the authors identify?
    Current research on "kawaii" (Japanese for "cute") primarily focuses on visual aspects, with limited exploration of sound. Specifically, in the field of Human-Computer Interaction (HCI), the domain of Voice User Experience (Voice UX) has yet to systematically investigate how manipulating sound characteristics can enhance the perception of "cuteness."

  • Why is this issue important?
    By integrating "kawaii" into the domain of sound, this research could provide novel design insights for voice user interfaces (e.g., voice assistants), game character design, and other socio-cultural interaction applications. Enhancing the perception of "cute" in sound may lead to interfaces that are more emotionally appealing and acceptable to users.

  • Research Motivation and Related Work
    Building on foundational studies (e.g., Seaborn et al.'s preliminary research on kawaii vocalics), the authors aim to examine whether sound characteristics (e.g., fundamental frequency and formant frequencies) can be consciously manipulated to enhance or reduce the perception of "cuteness" and explore its relationship with social identity (e.g., gender, age perception).

Solutions

  • What methods or solutions did the authors propose?
    The study proposed a multi-stage experimental approach using sound signal processing techniques to manually and automatically adjust sound characteristics related to "cuteness," aiming to validate and expand the kawaii vocalics model.

  • What is innovative about this solution?
    By adjusting the fundamental frequency (F0) and formant frequencies (F1, F2, F3) of speech, the authors systematically attempted to manipulate the perception of "cuteness" for the first time. These experiments not only validated the relationship between sound and social identity perception but also explored specific methods for adjusting sound characteristics.

  • What are the implementation steps and key technologies used?
    The research was conducted in four stages:

    1. Manual Processing: Using digital audio workstations (DAWs) to manually adjust text-to-speech (TTS) generated sounds.
    2. Automated Processing: Employing code libraries such as WORLD and Legacy-STRAIGHT to automatically adjust sound frequencies.
    3. Extension to Game Character Voices: Processing the voices of 18 different game characters and exploring their perceived "cuteness."
    4. Fine-Tuning Adjustments: Refining results through frequency adjustments of one to three semitones.

Research Outcomes

  • What specific results were achieved?

    • Core Findings: Increasing the fundamental frequency and the first formant frequency significantly enhanced the perception of "cuteness," particularly in TTS-generated voices.
    • Limitations: For game character voices, frequency manipulation's impact on "cuteness" varied by character and did not always linearly increase perceived cuteness.
    • User Perception Associations: Age perception showed a significant positive correlation with "cuteness," while gender ambiguity had limited correlation with "cuteness."
  • What advantages does this solution have compared to existing ones?

    • The study provides innovative methods for enriching acoustic and social identity perception applications in HCI.
    • It introduces a feasible approach for systematically manipulating sound characteristics related to "cuteness," opening new directions for voice user interface design.
  • What were the experimental or evaluation results?

    • Successfully amplified the perception of cuteness in certain TTS samples.
    • Inconsistent manipulation of game character voices: Some voices exhibited "sweet spots" after frequency adjustments, while others showed "ceiling effects."
    • Perception of "cuteness" was positively correlated with humanization and credibility.
  • Limitations and Future Directions

    • Limitations: The study did not examine the potential impact of voice content on perception, lacked validated tools for quantifying "cuteness" perception, and was limited to Japanese voice samples.
    • Future Research Suggestions:
      1. Conduct more precise temporal sequence analyses of sound.
      2. Investigate the influence of voice content on "cuteness."
      3. Expand cross-cultural studies to examine "cuteness" perception across different languages and cultural contexts.
      4. Explore potential ethical issues surrounding "cute" voice manipulation to prevent misuse in commercial or technological practices.
      5. Integrate the proposed manipulation methods into specific application domains (e.g., real-world scenarios for games or voice assistants).

Conclusion

The study innovatively explored methods for manipulating sound characteristics related to "cuteness," proposing a scientifically validated model that complements existing research on the concept of "kawaii." Despite certain technical and theoretical challenges, this research lays the groundwork for broad applications in HCI and cultural studies in the future.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/188595/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3713709
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Intelligent Voice Assistants (Alexa, Siri, etc.), Agent Personality & Anthropomorphism
work
Professions
—
article
Content Status
Full text indexed
hub
Related Papers
10 related papers