ComVi: Context-Aware Optimized Comment Display in Video Playback

Social & Collaborative VRImmersion & Presence ResearchLive Streaming & Content CreatorsContent Creators (YouTubers, Podcasters)Game Developers & DesignersHCI Researchers

Paper Title

ComVi: Context-Aware Optimized Comment Display in Video Playback

Publication Info

  • Topic area: Video playback interfaces with synchronized comment display.
  • Keywords: Video playback, comment synchronization, semantic alignment, user interface, dynamic programming, audio-visual correlation, user study, contextual relevance, engagement, Danmaku.

Background and Problem

  • Problem / challenge: Existing video platforms either display comments independently of video playback or use Danmaku interfaces that clutter the screen and lack synchronization for general comments without predefined timestamps.
  • Significance: Misaligned or overwhelming comment displays can disrupt viewer immersion, cause spoilers, and reduce engagement during video playback.
  • Motivation and related work: Previous research has explored filtering relevant comments or synchronizing textual content like transcripts with video timestamps. However, these approaches do not address the temporal alignment of general comments with specific video scenes.

Solution

  • Proposed approach: ComVi, a system that synchronizes comments with semantically relevant video timestamps and optimizes their display sequence based on relevance, popularity, and readability.
  • Novelty:
    1. Mapping general comments to video timestamps using audio-visual semantic correlation.
    2. Optimizing comment sequences with dynamic programming to balance relevance, popularity, and display duration.
    3. Introducing personalized comment curation features, such as query-based filtering and adjustable display settings.
  • Procedure and key techniques:
    • Compute audio-visual correlations using Sentence-BERT embeddings for subtitles and video captions.
    • Generate candidate comment sequences ensuring adequate reading time and non-overlapping display.
    • Score comments based on semantic relevance and normalized popularity (like counts).
    • Select the optimal sequence using dynamic programming.
    • Extend functionality for user-driven customization, including concurrent comment display and query-based filtering.

Results

  • Concrete findings:
    • ComVi selected 46 comments from 14,880 in a 3-minute 34-second video, with an average display duration of 4.65 seconds and normalized like count of 1,069.24.
    • Semantic correlation scores: ComVi outperformed random selection and achieved scores comparable to ground-truth timestamped comments.
    • Popularity scores: ComVi prioritized comments with higher normalized like counts compared to baselines.
    • Reading speed adjustments: Faster reading speeds resulted in shorter display durations and more comments, while slower speeds reduced the total number of comments.
    • User-specified queries and concurrent display settings modified comment sequences effectively.
  • Advantage over baselines:
    • ComVi significantly reduced mental and physical demand compared to YouTube and Danmaku interfaces.
    • Achieved higher contextual alignment and overall engagement ratings in user studies.
    • Preferred by 71.9% of participants over other interfaces.
  • Experiments / evaluation:
    • User study with 32 participants comparing five interface conditions (ComVi, YouTube, Danmaku, and single-comment variants).
    • Metrics: Mental demand, physical demand, contextual alignment, engagement, and preference ratings.
  • Limitations and future work:
    • Difficulty displaying overall comments unrelated to specific timestamps.
    • Bias toward comments quoting video content due to Sentence-BERT’s lexical similarity scoring.
    • Potential cognitive overload during high-density video scenes.
    • Static comment placement may obscure subtitles or key visual elements.

Summary

ComVi is a novel system that synchronizes comments with semantically relevant video timestamps, optimizing their display sequence based on relevance, popularity, and readability. Using audio-visual correlations and dynamic programming, ComVi enhances viewer engagement and reduces cognitive demand compared to existing interfaces like YouTube and Danmaku. User studies confirm its effectiveness, with 71.9% of participants preferring ComVi. Future work aims to improve handling of overall comments, balance novelty and quoting, and adapt comment placement dynamically to video content.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222817/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791018
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
3 authors
sell
Subtopics
Social & Collaborative VR, Immersion & Presence Research, Live Streaming & Content Creators
work
Professions
Content Creators (YouTubers, Podcasters), Game Developers & Designers, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers