Styling Words: A Simple and Natural Way to Increase Variability in Training Data Collection for Gesture Recognition

Hand Gesture RecognitionVisualization Perception & CognitionSoftware Engineers & DevelopersAI/ML Researchers & Engineers

Document Title

Styling Words: A Simple and Natural Way to Increase Variability in Training Data Collection for Gesture Recognition

Document Information

  • Domain: Human-Computer Interaction and Machine Learning, focusing on optimizing training data collection for gesture recognition.
  • Keywords: Gesture recognition, data collection, machine learning, styling words, human motion variability

Research Background and Problem

  • Challenges:

    1. Gestures are essential tools in human-computer interaction, but the accuracy of gesture recognition models is affected by individual differences, emotions, fatigue, as well as the quantity and quality of gesture samples.
    2. Data collection is time-consuming and expensive, while limited variability in controlled environments restricts data diversity.
    3. The quality and diversity of training data are crucial for improving the generalization performance of machine learning models.
  • Significance: By increasing variability in training data, gesture recognition models can better adapt to complex real-world applications, improving recognition accuracy.

  • Motivation and Related Work:

    1. Current research focuses more on optimizing deep learning models rather than improving data quality.
    2. Data variability is typically achieved by increasing the number of participants, which is costly and time-intensive.
    3. Data augmentation techniques have been applied to image data, but their application to gesture video data remains challenging.

Solution

  • Proposed Method: The authors propose using styling words (SWs) in instructions to guide participants in performing gestures, significantly increasing data diversity. For example, "Perform gesture #1 quickly" generates more diverse gesture variations compared to "Perform gesture #1."

  • Innovations:

    1. Utilizing simple and natural language instructions (styling words) to embed human-induced variability.
    2. Designing unique styling words categorized into intuitive and abstract types.
    3. Experimental validation of the direct relationship between data variability and gesture recognition accuracy.
  • Implementation Steps:

    1. Select 12 gestures and categorize them into single-hand or dual-hand types.
    2. Design two types of styling words: intuitive words (e.g., "slowly") and abstract words (e.g., "gracefully").
    3. Use OpenPose to extract skeletal key points from gesture videos.
    4. Test gesture recognition performance using the DD-Net model and compare datasets with and without SWs.

Research Findings

  • Specific Results:

    1. Datasets using SWs showed significant increases in variability, with standard deviation (std) rising by up to 419%.
    2. Gesture recognition models trained on high-variability datasets demonstrated significantly higher accuracy compared to models trained on traditional datasets.
  • Advantages:

    1. Models using SW datasets achieved a 9.26% improvement in average recognition accuracy on test sets.
    2. Even with reduced training data, models trained on SW datasets maintained stable accuracy.
    3. Intuitive SWs (e.g., "slowly") outperformed abstract SWs (e.g., "gracefully"), enhancing model robustness.
  • Experimental or Evaluation Results:

    1. Gesture recognition models trained on SW datasets performed better on higher variability data.
    2. Confusion matrices of the dataset revealed significant accuracy improvements for complex gestures (e.g., "bus stop gesture").
  • Limitations and Future Directions:

    1. Certain SWs may introduce excessive variability, increasing noise—for example, "hello gesture" could be misclassified.
    2. Styling words rely on participants' subjective interpretations, potentially leading to inconsistent performances.
    3. Future research could explore optimized styling words and extend this method to other modalities such as voice and touch data.

Conclusion and Significance

This study proposes a novel gesture data collection method based on styling words to enhance data diversity and optimize machine learning model performance. The findings demonstrate that high-quality, high-variability datasets can reduce dependency on training sample size and improve gesture recognition accuracy and generalization. The applicability and potential of styling words warrant further exploration in other interaction design domains.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47488/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445457
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Hand Gesture Recognition, Visualization Perception & Cognition
work
Professions
Software Engineers & Developers, AI/ML Researchers & Engineers
article
Content Status
Full text indexed
hub
Related Papers
2 related papers