FuzzySeek: Multimodal Refinement of Imprecise Video Queries for Moment Retrieval
Recent AI advances have made it possible to retrieve specific moments from long-form videos using natural language queries. However, existing systems can struggle to align retrieval results with user intent due to the lack of means for users to express their intents in simple natural language text. Moreover, there is limited support for helping users express or refine their intents interactively. We present FuzzySeek, a video moment retrieval interface that supports the expression and specification of imprecise or broad exploratory queries through multimodal interaction. FuzzySeek proposes three key components (1) Multimodality-blended text querying to improve expressivity, enabling users to directly anchor multimodal content within their textual queries, (2) Proactive Multimodal Guidance, which identifies imprecise/broad terms and phrases and surfaces targeted clarifications across modalities to improve query specificity and, (3) Query rollback to enable iterative back and forth exploration to enable direct or exploratory searches. Through a technical evaluation, multiple illustrative use cases and a user study with 11 participants, we show that FuzzySeek improves clarification efficiency, reduces cognitive load, and better supports video moment retrieval for imprecise queries compared to a baseline system without such support.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)