Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery Learning

Surgical Assistance & Medical TrainingPrototyping & User TestingSurgeons (Surgical Assistance Systems)

Title of the Paper

Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery Learning

Paper Information

  • Subject Area: Medical education technology, video-based learning, scene segmentation, and interaction design
  • Keywords: Video learning, question generation, scene segmentation, video navigation, surgical learning

Research Background and Issues

  • Identified Challenges:
    1. Surgical video learning is often passive, lacking interactivity and feedback.
    2. Existing technologies primarily focus on extracting task steps but provide insufficient support for learning critical details in single frames.
    3. Automated question generation technologies are predominantly text-based, making them less applicable to visually intensive surgical learning scenarios.
    4. Extracting keyframes suitable for teaching from videos is complex and time-consuming.
  • Research Significance:
    1. Surgical learning is a highly visual process involving anatomical structures, tool operations, and surgical decision-making.
    2. Embedding interactive questions and visual feedback into video viewing has been proven to enhance learning outcomes.
  • Research Motivation: This study aims to develop a system that enables surgeons to efficiently create interactive exercises with feedback based on surgical videos, while reducing the burden of annotating keyframes and generating questions.

Solution

  • Proposed Methods and Solutions:
    1. Developed a web-based system named Surgment, leveraging scene segmentation pipelines SegGPT + SAM for visual support.
    2. Integrated two key functionalities: "Search-by-Mask" and "Question Generation Tool."
  • Innovations:
    1. Combined the Segment Anything Model (SAM) and SegGPT models to significantly improve surgical scene segmentation accuracy (F1-score reaching 92%).
    2. Provided interactive features allowing users to quickly locate video frames by adjusting scene masks.
    3. Supported the generation of diverse question types (e.g., multiple-choice questions, path drawing) and scene-based visual feedback.
  • Implementation Steps and Key Technologies:
    1. Scene Segmentation: Extracted video frames, predicted labels using SegGPT, defined segmentation regions with SAM, and optimized outputs using a majority voting algorithm.
    2. Video Navigation: Identified keyframes and enabled quick retrieval of frames matching user-defined masks.
    3. Question Generation and Feedback: Created questions relevant to real surgical scenarios and provided high-quality visual feedback.

Research Outcomes

  • Specific Results:
    1. Segmentation Accuracy: The SegGPT+SAM scene segmentation framework outperformed UNet and SegGPT on public datasets, producing reasonable segmentation results with only 22 annotated images (0.15% of the total).
    2. Successful Application of the Question Generation Tool: All participants (including 11 surgeons) successfully used Surgment to create high-educational-value exercises and feedback.
    3. Navigation Performance: Surgment's search functionality achieved an image retrieval accuracy of 88%, far surpassing baseline methods (31.1%).
  • Comparison with Existing Solutions:
    1. Compared to traditional video browsing methods (e.g., frame-by-frame scrolling), Surgment's search tool significantly saved time.
    2. Visual feedback outperformed text/manual annotations, aiding in training details and teaching critical operations.
  • Experimental and Evaluation Results:
    1. Surgeons rated the system's usability highly, and visual feedback enhanced spatial awareness and learning specificity.
    2. Innovative interactive questions based on surgical scenarios (e.g., path drawing) received positive responses.
  • Limitations and Future Directions:
    1. User Interface Improvement: Current mask adjustment interactions are time-consuming; introducing voice commands is recommended to improve efficiency.
    2. Fine-grained Segmentation Enhancement: Especially for precise annotation of small anatomical features (e.g., gallbladder neck or arteries).
    3. 3D Scene Support: Surgery involves highly three-dimensional processes, necessitating the introduction of depth perception and dynamic scene tracking technologies.
    4. Function Expansion: Extend the system to complex surgical types and explore its application in augmented reality environments to support real-time surgical teaching and feedback.

This study combines advanced segmentation technologies with user-driven interactive design to provide an efficient and innovative tool for surgical teaching and learning.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/147894/2024

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3613904.3642587
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2024
emoji_events
Award
No award tagged
group
Authors
6 authors
sell
Subtopics
Surgical Assistance & Medical Training, Prototyping & User Testing
work
Professions
Surgeons (Surgical Assistance Systems)
article
Content Status
Full text indexed
hub
Related Papers
1 related papers