EventAnchor: Reducing Human Interactions in Event Annotation of Racket Sports Videos

Human Pose & Activity RecognitionInteractive Data VisualizationAthletes & Fitness EnthusiastsSociologists & Anthropologists

Title of the Paper

EventAnchor: Reducing Human Interactions in Event Annotation of Racket Sports Videos

Paper Information

  • Subject Area: Sports video analysis and AI-assisted annotation
  • Keywords: sports video, visual analysis, human-computer interaction, data annotation, event detection, table tennis

Research Background and Problem Statement

  • Problems and Challenges:

    • Manually extracting key information from lengthy match videos is highly challenging.
    • Existing computer vision-based data collection systems perform poorly on low-quality videos (e.g., broadcast videos), making it difficult to accurately track objects.
    • These systems often focus on low-level object recognition (e.g., human actions) but are weak in extracting high-level event information (e.g., action outcomes).
    • Current annotation systems lack scalability and are inefficient when handling large-scale annotation tasks.
    • For fast-paced sports like table tennis, existing data annotation methods are not effectively applicable.
  • Significance of the Research:

    • Table tennis and other racket sports are globally popular, with a high demand for data analysis, particularly for match tactics and player performance evaluation.
  • Motivation and Related Work:

    • Experts from different fields require analysis of table tennis match videos, including ball and player positions, action types, and tactical styles. Existing systems fail to meet these complex analytical needs.
    • The paper reviews current interactive video annotation systems, model-based annotation methods, and their design challenges.

Proposed Solution

  • Proposed Solution:

    • The authors propose the EventAnchor framework, which uses computer vision algorithms to identify significant events as anchors to help users locate, analyze, and annotate match videos.
    • They designed a table tennis annotation system based on EventAnchor, named EventAnchor for Table Tennis (ETT).
  • Innovations:

    • Dividing annotation tasks into three levels:
      • Object Level: Identifying key objects in the video (ball, players, table, etc.) using computer vision.
      • Event Level: Capturing significant events (e.g., racket-ball contact, ball bouncing on the table) through object motion information and specific algorithms.
      • Context Level: Allowing users to add contextual information interactively (e.g., tactical types).
    • Enhancing the timeline tool for quick video event browsing, with the use of a "calibration box" and "annotation box" to assist users in efficient labeling.
  • Technical Implementation:

    • Applied the FOTS model for optical character recognition (OCR).
    • Used TrackNet to extract ball trajectories and OpenPose to estimate player poses.
    • Employed ResNet-50 and support vector machines to classify video scenes.
    • Provided interactive tools for time calibration and event annotation.

Research Outcomes

  • Specific Outcomes:

    • The framework's effectiveness was validated through two experiments:
      1. Event-Level Task: Annotating the time and location of ball-table contact in table tennis videos.
      2. Context-Level Task: Labeling the "serve and attack" tactical type used by players during matches.
  • Experimental Results:

    • Compared to traditional systems, ETT significantly improved annotation efficiency:
      • Event-Level Task: Average task completion time reduced by approximately 14%, with a 50% reduction in time errors.
      • Context-Level Task: Completion time for complex tactical annotations was reduced by over 50%.
    • Users provided positive feedback on ETT's design and functionality, noting that anchor hints and interactive tools greatly reduced cognitive and interaction burdens.
  • Comparison with Existing Solutions:

    • ETT outperformed baseline systems in annotation efficiency.
    • For complex context-level tasks, ETT demonstrated significant time advantages, though annotation accuracy was comparable to baseline systems.
  • Limitations and Future Directions:

    • Limitations:
      • Limited flexibility in event definitions, currently relying primarily on ball contact with the table or racket.
      • Insufficient utilization of audio information.
      • Users may lack sufficient validation of results provided by computer vision.
    • Future Work:
      • Develop functionality for users to define new event types.
      • Enhance audio information recognition to complement visual data.
      • Improve interaction design to increase user engagement and result validation accuracy.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/47811/2021

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3411764.3445431
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2021
emoji_events
Award
No award tagged
group
Authors
9 authors
sell
Subtopics
Human Pose & Activity Recognition, Interactive Data Visualization
work
Professions
Athletes & Fitness Enthusiasts, Sociologists & Anthropologists
article
Content Status
Full text indexed
hub
Related Papers
0 related papers