Gesture-aware Interactive Machine Teaching with In-situ Object Annotations

Hand Gesture RecognitionHuman Pose & Activity RecognitionSoftware Engineers & DevelopersHCI Researchers

Title of the Paper

Gesture-aware Interactive Machine Teaching with In-situ Object Annotations

Paper Information

  • Domain: Interactive Machine Learning, Visual Interaction, User Interface Design
  • Keywords: Interactive Machine Teaching, Indicative Gestures, In-situ Annotation, Dataset, Object Segmentation

Research Background and Problem

  • What problems or challenges did the authors identify?

    • Existing visual interactive machine teaching systems often overlook the annotation of target objects or require users to annotate at a later stage, which may lead to models learning incorrect features unrelated to the target objects.
    • Annotation at later stages increases user burden and reduces the usability of teaching systems.
  • Why is this problem important?

    • Improving the accuracy and reliability of machine learning models is crucial for non-expert users, especially in practical applications, as avoiding incorrect recognition enhances user trust in the models.
    • Reducing user interaction burden can significantly improve the user experience of machine teaching systems.
  • Motivation and Related Work

    • Indicative gestures are highly relevant when users point to objects and can be utilized to guide object segmentation.
    • Existing studies have explored interactive machine teaching systems and annotation tools but have not deeply integrated object annotation into the teaching process.

Solution

  • What methods or solutions did the authors propose?

    • Proposed a visual interactive machine teaching system named “LookHere,” which integrates real-time object segmentation and annotation based on user indicative gestures.
    • Developed a dataset named “HuTics,” comprising 2040 annotated images of users pointing to target objects using various indicative gestures.
  • What is innovative about this solution?

    • Generates object highlight regions (segmentation masks) in real-time during the teaching process, guiding the model to focus on target object areas.
    • Incorporates real-time segmentation information into model training, achieving in-situ annotation.
    • Provides model evaluation visualization tools (e.g., saliency maps) to help users understand the model’s feature focus.
  • What key technologies were used in the implementation steps?

    • Real-time hand segmentation algorithms combined with a U-Net architecture to achieve object segmentation based on indicative gestures.
    • Utilized a re-optimized hand segmentation model (based on the LIP dataset) as auxiliary features.
    • Introduced joint classification and segmentation techniques to optimize saliency map generation.
    • Constructed and released the HuTics dataset for training object segmentation models.

Research Outcomes

  • What specific outcomes were achieved?

    • “LookHere” significantly reduced the time cost for users in the model creation process (approximately 14.3 times less compared to later-stage annotation).
    • Models created using “LookHere” demonstrated significantly improved object segmentation accuracy compared to traditional teaching systems (average IoU increased to 0.718).
    • Users expressed satisfaction with the efficiency and intuitiveness of “LookHere.”
  • What advantages does it have compared to existing solutions?

    • Achieved seamless integration of teaching and annotation processes, eliminating the burden of later-stage annotation.
    • Provided real-time feedback (object highlight regions), enhancing user control and confidence.
    • Performance evaluation showed comparable results to later-stage annotation but greatly simplified user interaction.
  • What were the experimental or evaluation results?

    • Users were able to create highly accurate models in a short time, with segmentation performance reaching an average IoU of 0.718.
    • NASA-TLX metrics indicated that compared to traditional annotation workflows, LookHere excelled in reducing mental workload, physical effort, and time pressure.
  • Limitations and Future Directions

    • The current system faces challenges in long-distance pointing or understanding 3D scenes; future work could incorporate depth sensing and 3D scene reconstruction technologies.
    • Lacks functionality for actively correcting erroneous segmentation regions; future integration of voice input or other active annotation features could address this.
    • Privacy concerns arise from capturing user activities; further research is needed to balance privacy protection with efficient annotation design.
    • The dataset has potential for expansion to other application scenarios, such as virtual backgrounds, intelligent portrait modes, and support for visually impaired users.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/uist/85000/2022

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3526113.3545648
At a Glance

Paper Snapshot

fact_check
dataset
Source
UIST
calendar_month
Year
2022
emoji_events
Award
No award tagged
group
Authors
2 authors
sell
Subtopics
Hand Gesture Recognition, Human Pose & Activity Recognition
work
Professions
Software Engineers & Developers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
9 related papers