Where is the Boundary? Understanding How People Recognize and Evaluate Generative AI-extended Videos

Generative AI (Text, Image, Music, Video)Explainable AI (XAI)Misinformation & Fact-CheckingContent Creators (YouTubers, Podcasters)Film & Animation ProducersFact-Checkers

Research Background and Issues

  • What problems or challenges did the authors identify?
    With the rapid development of video generation models, it is becoming increasingly difficult for people to distinguish the authenticity of video content. Specifically, AI-augmented videos blend real footage with generated content, making the boundary between real and fake segments a significant challenge to discern. Additionally, AI-augmented videos have the potential to mislead the public, while also offering new experiences for video creators. However, there is currently limited research on how people identify and evaluate AI-augmented videos.

  • Why is this issue important?
    Inability to accurately identify generated content could lead to the spread of fake news, public opinion crises, and mislead viewers' understanding of video authenticity. Addressing this issue can not only improve video content verification but also promote better utilization of generative technologies in the video creation field.

  • Research Motivation and Related Work
    Current methods primarily rely on video data features to detect AI-generated content, but these approaches overlook the role of human cognition and lack exploration of mixed-type videos (real + generated). Furthermore, as the quality of AI models improves, current detection methods are prone to becoming obsolete. This study aims to fill this research gap through human-computer interaction experiments, exploring human recognition and evaluation of boundaries in mixed content.

Solutions

  • What methods or solutions did the authors propose?
    The authors designed and implemented a web-based experimental system to collect participants' feedback and ratings on boundary recognition in AI-augmented videos. Additionally, they conducted quantitative and qualitative analyses through questionnaires and personal interviews to gain deeper insights into people's perceptual differences, evaluation criteria, and attitudes.

  • What are the innovative aspects of this solution?

    • Combined multiple generative models and language models (e.g., Tarsier and GPT-4) to produce high-quality AI-augmented videos for experiments.
    • Systematically studied human capability to recognize boundaries in mixed video content, the influencing factors, and evaluation criteria.
    • Used large-scale user experiments and interviews to reveal the impact of video style, consistency, and narrative coherence on human cognition.
  • What are the implementation steps and key technologies used?

    1. Video Data Preparation
      Collected original videos and extended their content using AI models (e.g., Kling).
    2. Experimental Design
      Designed a web-based experimental system where participants selected boundaries using a progress bar and completed questionnaires on influencing factors and video ratings.
    3. Data Collection and Analysis
      Extracted participants' boundary time deviations, factor selection rates, and rating data for quantitative analysis; supplemented with qualitative analysis through semi-structured interviews.
    4. Review and Discussion
      Combined quantitative data and participant feedback to summarize key factors influencing cognition and evaluation.

Research Findings

  • What specific findings were achieved?

    1. Boundary Recognition Ability
      Human recognition of boundaries between original and generated video content showed time deviations concentrated within [-0.1961 seconds, 1.9785 seconds], indicating relatively accurate overall recognition.
    2. Analysis of Influencing Factors
      Dynamics (e.g., motion and geometric changes) had the greatest impact on boundary recognition, while scene changes were less noticeable. Several factors showed significant positive correlations, such as motion and object geometry.
    3. Video Rating Analysis
      Narrative coherence was the most important factor influencing overall evaluation, closely related to video style consistency and perceived quality.
    4. User Attitudes
      The majority of participants held an optimistic attitude toward AI-augmented videos, believing they could enhance creative efficiency and inspire imagination, while also expressing concerns about potential risks such as legal and copyright issues.
  • What are the advantages compared to existing solutions?

    • Considered user experience and human cognitive factors, addressing a research gap overlooked by existing methods.
    • Proposed an experimental framework combining qualitative and quantitative approaches, enriching user research methods for AI-augmented videos.
    • Explored the potential impact of future technological developments on the field of video generation, offering both theoretical and practical value.
  • What were the experimental or evaluation results?

    • In 95% of cases, participants' boundary recognition time deviations were small, indicating that video content generation technology has not yet fully deceived human cognition.
    • Quantitative data revealed strong correlations among influencing factors (e.g., motion and geometric changes), while qualitative data supplemented potential reasons for recognition difficulties.
  • Limitations and Future Directions

    1. Video selection was overly narrow; future research should extend to diverse content types such as animation and films.
    2. Experimental video lengths were relatively short, limiting the study of complex scenes or narrative designs.
    3. Limited cultural and geographical diversity among participants affected the generalizability of conclusions.
    4. Did not explore the role of audio in video recognition and evaluation, which could be addressed in future studies.
    5. The current experiment explicitly informed participants that the videos were AI-augmented, without testing recognition ability in natural scenarios.

Through systematic research, this study revealed the characteristics of human recognition and evaluation of mixed video content and proposed recommendations for improving video generation and augmentation technologies. These findings contribute to the future development of AI-augmented video technologies and provide theoretical support for user research in related fields.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/189654/2025

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://dl.acm.org/doi/10.1145/3706598.3714061
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2025
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Generative AI (Text, Image, Music, Video), Explainable AI (XAI), Misinformation & Fact-Checking
work
Professions
Content Creators (YouTubers, Podcasters), Film & Animation Producers, Fact-Checkers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers