Counting How the Seconds Count: Understanding TikTok Behavior via ML-driven Analysis of Video Content

Generative AI (Text, Image, Music, Video)AI-Assisted Decision-Making & AutomationRecommender System UXSocial Platform Design & User BehaviorSoftware Engineers & DevelopersData Scientists & AnalystsHCI Researchers

Paper Title

Counting How the Seconds Count: Understanding TikTok Behavior via ML-driven Analysis of Video Content

Publication Info

  • Topic area: Human-computer interaction and algorithm-user interplay in short video platforms.
  • Keywords: TikTok, recommendation systems, short video platforms, Vision Language Models, user behavior, video content analysis, algorithm-user interaction, personalization, HCI measurement, multimodal analysis.

Background and Problem

  • Problem / challenge: Existing studies on TikTok’s recommendation algorithms rely heavily on manual analysis, textual data, or user interviews, which are limited in scale and depth. There is a lack of automated, multimodal approaches to analyze video content and its impact on user experience.
  • Significance: Understanding TikTok’s recommendation algorithms is crucial for improving user experience, designing better algorithms, and addressing concerns about personalization, filter bubbles, and algorithmic transparency.
  • Motivation and related work: Prior research has explored folk theories, user perceptions, and statistical analyses of TikTok’s algorithms, but has not deeply analyzed video content or its temporal dynamics at scale. This paper introduces a scalable, multimodal measurement approach to fill this gap.

Solution

  • Proposed approach: Video Content Analysis (VCA), an automated tool leveraging Vision Language Models (VLMs) to analyze video content at scale, combined with user studies and data donation.
  • Novelty:
    1. Application of VLMs for analyzing user behavior in short video platforms.
    2. Integration of multimodal data sources (video embeddings, user studies, browsing histories).
    3. Quantitative validation of folk theories and discovery of new insights about TikTok’s recommendation system.
    4. Identification of temporal dynamics in user-algorithm interactions.
  • Procedure and key techniques:
    • Develop VCA using Video-LLaMA embeddings to represent video content.
    • Cluster videos into 100 abstract categories using KMeans.
    • Conduct user studies with modified and unmodified TikTok feeds.
    • Analyze browsing history files donated by users to study long-term recommendation trends.
    • Validate findings through statistical tests and regression analyses.

Results

  • Concrete findings:
    1. Users spend 50% of daily watch time in their top 5 clusters, but these clusters change frequently over time.
    2. Recommendations become increasingly individualized, moving away from trending content.
    3. User interactions (likes/shares) correlate more with past recommendations than future ones, contradicting folk theories.
    4. Users prefer sessions with novel content over those aligned with their historical preferences.
    5. Dropping 30% of videos from the recommended sequence significantly reduces user engagement, while a 5% drop has negligible impact.
    6. VCA-based prediction of user engagement achieves 69.65% accuracy, highlighting the limits of predictability in user behavior.
  • Advantage over baselines: VCA enables scalable, content-aware analysis of video recommendations, overcoming the limitations of manual and hashtag-based methods.
  • Experiments / evaluation:
    • User studies with 68 participants across in-person and online platforms.
    • Analysis of 2.65 million videos from donated browsing histories.
    • Validation of VCA clusters through a separate study with 22 participants.
    • Statistical tests (e.g., t-tests, Wilcoxon signed-rank test) and regression analyses.
  • Limitations and future work:
    • Skewed participant demographics (18–34 years old).
    • Use of TikTok’s web browser version instead of the mobile app.
    • Dependence on pretrained VLMs and unsupervised clustering.
    • Need for more diverse user studies and exploration of soft clustering techniques.

Summary

This paper introduces Video Content Analysis (VCA), a scalable tool leveraging Vision Language Models to analyze TikTok’s recommendation algorithms and their impact on user behavior. Key findings include the dynamic nature of recommendations, limited influence of user interactions on future recommendations, and user preference for novel content. The study highlights the importance of sequence continuity in recommendations and demonstrates the predictive value of content-based representations. VCA offers a reusable HCI tool for multimodal data analysis, with implications for algorithm design, user experience, and broader video platforms.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/223095/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3790311
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
5 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), AI-Assisted Decision-Making & Automation, Recommender System UX, Social Platform Design & User Behavior
work
Professions
Software Engineers & Developers, Data Scientists & Analysts, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
1 related papers