How Well do LLMs Assist Parents in Assessing Child Appropriateness of Videos?

Human-LLM CollaborationExplainable AI (XAI)Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)Early Childhood EducatorsHCI Researchers

Paper Title

How Well do LLMs Assist Parents in Assessing Child Appropriateness of Videos?

Publication Info

  • Topic area: Evaluating the role of Large Language Models (LLMs) in aiding parental decision-making for child-safe video content.
  • Keywords: LLMs, child safety, video moderation, parental decision-making, explainable AI, GPT-4, content appropriateness, transparency, human-in-the-loop, child development.

Background and Problem

  • Problem / challenge: The vast volume of online video content makes manual curation for child-appropriateness impractical. Automated systems lack transparency, fail to capture diverse parental values, and struggle with nuanced content moderation.
  • Significance: Exposure to inappropriate content can have lasting negative effects on children, including psychological harm and unsafe behaviors. Parents need tools to make informed decisions without co-viewing all content.
  • Motivation and related work: Existing systems like YouTube Kids fail to filter inappropriate content 27% of the time. State-of-the-art models achieve high accuracy but lack explainability. Recent research suggests LLMs could bridge the gap between content moderation and explainability, but their effectiveness in multimedia content moderation remains unexplored.

Solution

  • Proposed approach: Evaluate the ability of GPT-4 to assess and describe the appropriateness of YouTube videos for children under 7, focusing on its alignment with parental reasoning and its potential as a decision-support tool.
  • Novelty:
    1. Analysis of GPT-4's reasoning about child-appropriateness based on engagement, safety, and educational value.
    2. Comparison of LLM-generated decisions with parental judgments.
    3. Development and evaluation of refined prompts to align LLM-generated descriptions with parental preferences.
    4. Exploration of LLMs as tools for aiding parental decision-making rather than making final moderation decisions.
  • Procedure and key techniques:
    • Two studies with 370 parents (120 in Study 1, 250 in Study 2).
    • Study 1: Evaluate GPT-4's classification and explanation capabilities using 110 video transcripts.
    • Study 2: Test refined prompts for generating video descriptions and compare parental decisions across four groups (baseline, control, two LLM-generated description groups).
    • Mixed-method analysis combining qualitative coding, quantitative accuracy metrics, and satisfaction/trust scales.

Results

  • Concrete findings:
    • GPT-4 achieved 72% accuracy on the Samba dataset, lower than prior models (88–95% accuracy).
    • Parents rated LLM-generated explanations as satisfying (92.1% satisfaction), but trust in the model was low (25.1% trustful).
    • Refined prompts improved alignment with parental preferences, with one prompt achieving statistically equivalent decisions to parents who watched full videos (within a 0.3 error margin).
    • Root Mean Square Error (RMSE) for parental decisions using refined prompts was 0.418, compared to 0.826 for metadata-only decisions.
  • Advantage over baselines:
    • Descriptions generated by refined prompts enabled parents to make decisions comparable to watching full videos, significantly outperforming metadata-based decisions.
    • Refined prompts captured nuanced parental concerns, such as safety and potential future consequences, better than unrefined prompts.
  • Experiments / evaluation:
    • Study 1: Compared GPT-4's classifications with Samba dataset labels and parental judgments.
    • Study 2: Evaluated refined prompts using four participant groups (baseline, control, two LLM-generated description groups) with independent t-tests and RMSE analysis.
    • Metrics: Accuracy, precision, recall, Explanation Satisfaction (ES) scale, Trust/Reliance Measurement (TM) scale.
  • Limitations and future work:
    • Limited to GPT-4; other LLMs were not evaluated.
    • Contextual limitations of transcripts; future studies could use audio-visual inputs.
    • Participants were exclusively from the US; results may not generalize globally.
    • Future work could explore hybrid or in-person studies, larger datasets, and integration with conversational agents on platforms like YouTube.

Summary

This study evaluated GPT-4's ability to assist parents in assessing the appropriateness of YouTube videos for children under 7. While GPT-4's classification accuracy was subpar (72%), its explanations were highly satisfying to parents (92.1% satisfaction). Refined prompts improved alignment with parental preferences, enabling decisions comparable to watching full videos. The findings suggest that LLMs should not make final moderation decisions but can serve as effective tools for aiding parental decision-making. Future work could explore broader applications, improved LLM models, and real-world integration to support child safety online.

Quick Actions

Share

Share this page

ios_share

https://hci.top/en/papers/chi/222255/2026

AdRecommended

Learn AI Coding at CodeNow

open_in_newOpen DOI Link
DOI: https://doi.org/10.1145/3772318.3791610
At a Glance

Paper Snapshot

fact_check
dataset
Source
CHI
calendar_month
Year
2026
emoji_events
Award
No award tagged
group
Authors
4 authors
sell
Subtopics
Human-LLM Collaboration, Explainable AI (XAI), Cognitive Impairment & Neurodiversity (Autism, ADHD, Dyslexia)
work
Professions
Early Childhood Educators, HCI Researchers
article
Content Status
Full text indexed
hub
Related Papers
0 related papers