Gesture-aware Interactive Machine Teaching with In-situ Object Annotations
Title of the Paper
Gesture-aware Interactive Machine Teaching with In-situ Object Annotations
Paper Information
- Domain: Interactive Machine Learning, Visual Interaction, User Interface Design
- Keywords: Interactive Machine Teaching, Indicative Gestures, In-situ Annotation, Dataset, Object Segmentation
Research Background and Problem
-
What problems or challenges did the authors identify?
- Existing visual interactive machine teaching systems often overlook the annotation of target objects or require users to annotate at a later stage, which may lead to models learning incorrect features unrelated to the target objects.
- Annotation at later stages increases user burden and reduces the usability of teaching systems.
-
Why is this problem important?
- Improving the accuracy and reliability of machine learning models is crucial for non-expert users, especially in practical applications, as avoiding incorrect recognition enhances user trust in the models.
- Reducing user interaction burden can significantly improve the user experience of machine teaching systems.
-
Motivation and Related Work
- Indicative gestures are highly relevant when users point to objects and can be utilized to guide object segmentation.
- Existing studies have explored interactive machine teaching systems and annotation tools but have not deeply integrated object annotation into the teaching process.
Solution
-
What methods or solutions did the authors propose?
- Proposed a visual interactive machine teaching system named “LookHere,” which integrates real-time object segmentation and annotation based on user indicative gestures.
- Developed a dataset named “HuTics,” comprising 2040 annotated images of users pointing to target objects using various indicative gestures.
-
What is innovative about this solution?
- Generates object highlight regions (segmentation masks) in real-time during the teaching process, guiding the model to focus on target object areas.
- Incorporates real-time segmentation information into model training, achieving in-situ annotation.
- Provides model evaluation visualization tools (e.g., saliency maps) to help users understand the model’s feature focus.
-
What key technologies were used in the implementation steps?
- Real-time hand segmentation algorithms combined with a U-Net architecture to achieve object segmentation based on indicative gestures.
- Utilized a re-optimized hand segmentation model (based on the LIP dataset) as auxiliary features.
- Introduced joint classification and segmentation techniques to optimize saliency map generation.
- Constructed and released the HuTics dataset for training object segmentation models.
Research Outcomes
-
What specific outcomes were achieved?
- “LookHere” significantly reduced the time cost for users in the model creation process (approximately 14.3 times less compared to later-stage annotation).
- Models created using “LookHere” demonstrated significantly improved object segmentation accuracy compared to traditional teaching systems (average IoU increased to 0.718).
- Users expressed satisfaction with the efficiency and intuitiveness of “LookHere.”
-
What advantages does it have compared to existing solutions?
- Achieved seamless integration of teaching and annotation processes, eliminating the burden of later-stage annotation.
- Provided real-time feedback (object highlight regions), enhancing user control and confidence.
- Performance evaluation showed comparable results to later-stage annotation but greatly simplified user interaction.
-
What were the experimental or evaluation results?
- Users were able to create highly accurate models in a short time, with segmentation performance reaching an average IoU of 0.718.
- NASA-TLX metrics indicated that compared to traditional annotation workflows, LookHere excelled in reducing mental workload, physical effort, and time pressure.
-
Limitations and Future Directions
- The current system faces challenges in long-distance pointing or understanding 3D scenes; future work could incorporate depth sensing and 3D scene reconstruction technologies.
- Lacks functionality for actively correcting erroneous segmentation regions; future integration of voice input or other active annotation features could address this.
- Privacy concerns arise from capturing user activities; further research is needed to balance privacy protection with efficient annotation design.
- The dataset has potential for expansion to other application scenarios, such as virtual backgrounds, intelligent portrait modes, and support for visually impaired users.
Research Questions / Practical Problems
Question signals indexed for this paper.
Research Questions
3- How can interactive machine teaching accuracy be improved by combining real-time object segmentation and indicative user gestures?Category: AI Trust in Education and Learning ContextsSimilar questionsarrow_forward
- How does real-time segmentation information reduce users' annotation burden in interactive machine teaching?Category: AI Trust in Education and Learning ContextsSimilar questionsarrow_forward
- How does user feedback (e.g., saliency maps) enhance user trust and sense of control in model teaching?Category: AI Trust in Education and Learning ContextsSimilar questionsarrow_forward
Practical Problems
1- Users must spend significant time on post-hoc object annotation, increasing usage burden.Category: AI Trust in Education and Learning ContextsSimilar questionsarrow_forward
- 75%
Grasping Microgestures: Eliciting Single-hand Microgestures for Handheld Objects
CHI '19· Hand Gesture Recognition +1
- 75%
µGlyph: a Microgesture Notation
CHI '23· Hand Gesture Recognition
- 67%
Log2Motion: Biomechanical Motion Synthesis from Touch Logs
CHI '26· Hand Gesture Recognition +2
- 60%
Pentelligence: Combining Pen Tip Motion and Writing Sounds for Handwritten Digit Recognition
CHI '18· Hand Gesture Recognition +1
- 60%
The Voight-Kampff Machine for Automatic Custom Gesture Rejection Threshold Selection
CHI '22· Hand Gesture Recognition +1
- 60%
ReflecTouch: Detecting Grasp Posture of Smartphone Using Corneal Reflection Images
CHI '22· Human Pose & Activity Recognition +1
- 60%
RestfulRaycast: Exploring Ergonomic Rigging and Joint Amplification for Precise Hand Ray Selection in XR
DIS '25· Hand Gesture Recognition +1
- 60%
Abacus Gestures: A Large Set of Math-Based Usable Finger-Counting Gestures for Mid-Air Interactions
UbiComp '23· Hand Gesture Recognition +1
- 60%
TouchPose: Hand Pose Prediction, Depth Estimation, and Touch Classification from Capacitive Images
UIST '21· Hand Gesture Recognition +1
Based on Jaccard similarity of research subtopics & professions (≥60%)