SEMOUR: Scripted EMOtional speech repository for URdu

Multilingual & Cross-Cultural Voice InteractionAgent Personality & AnthropomorphismVoice AccessibilityAI/ML Researchers & EngineersHCI ResearchersCognitive Scientists

Title of the Paper

SEMOUR: A Scripted Emotional Speech Repository for Urdu

Paper Information

  • Subject Area: Emotional speech recognition datasets and deep learning
  • Keywords: Emotional speech, speech dataset, digital recordings, speech emotion recognition, Urdu, human annotation, machine learning, deep learning

Research Background and Problem

  • Identified Issues or Challenges:

    • Current speech emotion recognition (SER) datasets are primarily focused on a few languages (e.g., English, German, and Italian), and a comprehensive dataset for widely spoken Urdu has not yet been developed.
    • Urdu is the 11th most spoken language in the world (used by approximately 171 million people), yet there is a lack of high-quality speech data to support emotion recognition.
    • Existing Urdu emotional datasets are small (400 samples) and have limited emotional categories, making them insufficient for complex emotion analysis or machine learning model requirements.
  • Significance:

    • Developing an Urdu emotional speech recognition system is crucial for applications like health diagnostics (psychological disorder detection), public sentiment analysis, and social equity research.
    • Providing a large, well-annotated emotional speech dataset can improve the accuracy of emotion recognition systems and broaden their application scope.
  • Research Motivation and Related Work:

    • Previous studies have shown that emotional speech recognition systems for mainstream languages benefit significantly from deep learning technologies, such as those using the IEMOCAP (English) and Emo-DB (German) databases. These systems rely on large, high-quality annotated datasets, and the lack of such data limits model training and generalization.
    • There is an urgent need for a phoneme- and gender-balanced emotional speech dataset in Urdu to support model training and application in this language.

Solution

  • Proposed Method or Solution:

    • Developed SEMOUR, the first emotional speech dataset for Urdu, covering eight emotions (e.g., happiness, sadness, surprise, etc.) and including 15,040 high-quality speech samples.
    • Data was recorded by professional actors in soundproof studios, with balanced phoneme composition and high consistency in annotations provided by human labelers.
    • Proposed a deep neural network (DNN)-based emotion recognition model and recorded benchmark experimental results on the dataset.
  • Innovations:

    • The dataset ensures balanced distribution of Urdu phonemes, grammatical complexity, and gender composition, enhancing model generalization.
    • Unlike previous corpora that only include natural emotions or limited content, SEMOUR uses scripted and manually annotated high-quality emotional speech samples.
    • The model was trained using multiple speech features (MFCC, Chroma, Mel spectrogram), achieving 92% emotion recognition accuracy.
  • Implementation Steps and Key Techniques:

    1. Script Design: Extracted 67 Urdu phonemes from two major data sources (high-frequency word sets and dictionaries) and designed diverse scripts containing 235 speech instances.
    2. Speech Recording: Hired professional actors to perform recordings in soundproof studios, covering eight emotional categories with gender balance.
    3. Data Processing: Performed noise reduction, segmentation, and standardization on the recorded files.
    4. Human Annotation: Selected 5,000 speech samples for scoring by 16 experts to ensure consistency in emotional labels.
    5. Feature Extraction and Modeling: Used the Librosa library to extract MFCC, Chroma, and Mel spectrogram features; trained a DNN classification model.

Research Outcomes

  • Specific Results:

    • The SEMOUR dataset successfully includes over 15,000 samples, with an average duration of 1.657 seconds per speech instance.
    • The developed DNN model achieved 92% accuracy on the speech emotion recognition task, significantly outperforming traditional machine learning models (e.g., SVM or Random Forest).
  • Comparison with Existing Solutions:

    • SEMOUR substantially expands the scope of emotional speech data in Urdu, while the largest previous Urdu dataset contained only 400 speech samples.
    • Compared to traditional feature extraction and classification methods, deep learning approaches demonstrated higher accuracy and flexibility.
  • Experimental or Evaluation Results:

    • Through 10-fold cross-validation, the model achieved an average accuracy of 90% on random test sets.
    • In gender-specific analysis, models trained separately on male and female actors achieved accuracies of 96% and 92%, respectively, though performance dropped in cross-gender testing.
    • For unseen speakers (Leave-One-Out experiments), accuracy dropped to 39%, indicating a significant impact of speaker individuality on model performance.
  • Limitations and Future Directions:

    1. Limitations:
      • Limited generalization to natural emotional speech: SEMOUR data is primarily based on scripted performances, which may not accurately simulate the complexity of natural emotional speech.
      • Insufficient representation of language and dialects: The current dataset only includes one Urdu dialect (Lahore accent).
      • Lack of exploration of more advanced deep learning models (e.g., LSTM or Transformer) to address semantic diversity in unseen speakers.
    2. Future Directions:
      • Expand the dataset to include all major Urdu dialects and accents.
      • Collect natural emotional speech data to validate the model's generalization capabilities.
      • Test more advanced model architectures to improve classification performance for unseen speakers.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47362/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445171
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Multilingual & Cross-Cultural Voice Interaction, Agent Personality & Anthropomorphism, Voice Accessibility
work
Professions
AI/ML Researchers & Engineers, HCI Researchers, Cognitive Scientists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers