TmoTA: Simple, Highly Responsive Tool for Multiple Object Tracking Annotation

Customizable & Personalized ObjectsPrototyping & User TestingComputational Methods in HCISoftware Engineers & DevelopersHCI Researchers

Document Title

TmoTA: Simple, Highly Responsive Tool for Multiple Object Tracking Annotation

Document Information

  • Subject Area: Machine learning data annotation tools, multi-object tracking, video annotation
  • Keywords: manual annotation, data annotation, video sequence annotation, multi-object tracking, responsiveness, open-source tool, interaction design, time efficiency, user experience

Research Background and Issues

  • Identified Problems or Challenges:

    1. Data annotation in current machine learning projects is a time-consuming and costly process, especially when domain experts are required for annotation.
    2. Existing video annotation tools (e.g., CVAT and Label Studio) are powerful but lack efficiency and responsiveness in use.
    3. Semi-automated annotation tools, while improving efficiency in certain cases, often require domain-specific model training and exhibit slow responsiveness.
  • Importance: Data annotation is a critical prerequisite for training machine learning models. Designing efficient and user-friendly annotation tools can directly impact the model training cycle and reduce enterprise costs.

  • Research Motivation and Related Work:

    1. Previous annotation tools include manual tools (e.g., CVAT, Label Studio), which are more general but have low interaction efficiency, and semi-automated or automated tools (e.g., Supervisely, VATIC), which improve efficiency but face domain adaptation challenges.
    2. The authors propose a tool specifically designed to accelerate the manual multi-object tracking annotation process—TmoTA—and validate its effectiveness through comparative experiments with existing tools.

Solution

  • Proposed Method:

    1. TmoTA is an open-source, highly responsive tool designed for manual multi-object tracking annotation in 2D videos.
    2. The tool incorporates various optimization features, such as view centering, loop playback, occlusion handling, and efficient interaction mechanisms.
  • Innovations:

    1. By eliminating automated computation, the tool ensures immediate responsiveness during use.
    2. Enhances user interaction during the annotation process, such as allowing boundary box edge adjustments with a single click and supporting linear interpolation.
    3. Introduces occlusion region identification and object centralization features to simplify multi-object annotation in complex scenes.
  • Implementation Steps and Key Techniques:

    1. Efficient installation and startup: Users can begin annotation by simply downloading, extracting, and dragging videos into the tool.
    2. Multi-view layout design: Integrates video view, timeline view, and classic user interface components.
    3. Interface interaction design: Supports rapid frame switching, boundary box adjustments, and playback controls.
    4. Algorithm support: Linear interpolation reduces user workload, while real-time view centering and loop playback facilitate annotation quality verification.
    5. Utilizes OpenGL and efficient data structures to ensure fast rendering and data access operations.

Research Outcomes

  • Specific Results:

    1. Compared to other widely used manual annotation tools (e.g., CVAT and Label Studio), TmoTA reduces per-frame object annotation time by 20%-40%.
    2. While slightly slower than semi-automated tools (e.g., Supervisely and VATIC), it achieves comparable efficiency while maintaining high responsiveness.
    3. In System Usability Scale (SUS) evaluations, TmoTA is particularly favored by experienced users, though there is room for improvement for novices.
  • Advantages:

    1. Significant improvement in annotation speed, reducing overall annotation time by approximately 38%-61%.
    2. Equipped with comprehensive user guides and intuitive operation methods, lowering the learning curve.
  • Experimental Results:

    1. Time Efficiency: User studies show that TmoTA's average annotation time is significantly lower than other manual tools.
    2. Accuracy Assessment: Using multi-object tracking accuracy (MOTA) and tracking accuracy (HOTA), TmoTA demonstrates annotation quality comparable to or slightly better than existing tools.
  • Limitations and Future Directions:

    1. The current version requires videos to be fully loaded into memory, posing limitations for high-resolution, long-duration videos.
    2. SUS scores for novice users remain lower than those for CVAT, indicating a need for improved beginner-friendliness.
    3. Future work will explore applying TmoTA's technical concepts to semi-automated or fully automated tools and extend user studies to validate learning effects on a larger scale.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/95750/2023

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3544548.3581185
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2023
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Customizable & Personalized Objects, Prototyping & User Testing, Computational Methods in HCI
work
Professions
Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
8 related papers