How Well do LLMs Assist Parents in Assessing Child Appropriateness of Videos?
Authors
Paper Title
How Well do LLMs Assist Parents in Assessing Child Appropriateness of Videos?
Publication Info
- Topic area: Evaluating the role of Large Language Models (LLMs) in aiding parental decision-making for child-safe video content.
- Keywords: LLMs, child safety, video moderation, parental decision-making, explainable AI, GPT-4, content appropriateness, transparency, human-in-the-loop, child development.
Background and Problem
- Problem / challenge: The vast volume of online video content makes manual curation for child-appropriateness impractical. Automated systems lack transparency, fail to capture diverse parental values, and struggle with nuanced content moderation.
- Significance: Exposure to inappropriate content can have lasting negative effects on children, including psychological harm and unsafe behaviors. Parents need tools to make informed decisions without co-viewing all content.
- Motivation and related work: Existing systems like YouTube Kids fail to filter inappropriate content 27% of the time. State-of-the-art models achieve high accuracy but lack explainability. Recent research suggests LLMs could bridge the gap between content moderation and explainability, but their effectiveness in multimedia content moderation remains unexplored.
Solution
- Proposed approach: Evaluate the ability of GPT-4 to assess and describe the appropriateness of YouTube videos for children under 7, focusing on its alignment with parental reasoning and its potential as a decision-support tool.
- Novelty:
- Analysis of GPT-4's reasoning about child-appropriateness based on engagement, safety, and educational value.
- Comparison of LLM-generated decisions with parental judgments.
- Development and evaluation of refined prompts to align LLM-generated descriptions with parental preferences.
- Exploration of LLMs as tools for aiding parental decision-making rather than making final moderation decisions.
- Procedure and key techniques:
- Two studies with 370 parents (120 in Study 1, 250 in Study 2).
- Study 1: Evaluate GPT-4's classification and explanation capabilities using 110 video transcripts.
- Study 2: Test refined prompts for generating video descriptions and compare parental decisions across four groups (baseline, control, two LLM-generated description groups).
- Mixed-method analysis combining qualitative coding, quantitative accuracy metrics, and satisfaction/trust scales.
Results
- Concrete findings:
- GPT-4 achieved 72% accuracy on the Samba dataset, lower than prior models (88–95% accuracy).
- Parents rated LLM-generated explanations as satisfying (92.1% satisfaction), but trust in the model was low (25.1% trustful).
- Refined prompts improved alignment with parental preferences, with one prompt achieving statistically equivalent decisions to parents who watched full videos (within a 0.3 error margin).
- Root Mean Square Error (RMSE) for parental decisions using refined prompts was 0.418, compared to 0.826 for metadata-only decisions.
- Advantage over baselines:
- Descriptions generated by refined prompts enabled parents to make decisions comparable to watching full videos, significantly outperforming metadata-based decisions.
- Refined prompts captured nuanced parental concerns, such as safety and potential future consequences, better than unrefined prompts.
- Experiments / evaluation:
- Study 1: Compared GPT-4's classifications with Samba dataset labels and parental judgments.
- Study 2: Evaluated refined prompts using four participant groups (baseline, control, two LLM-generated description groups) with independent t-tests and RMSE analysis.
- Metrics: Accuracy, precision, recall, Explanation Satisfaction (ES) scale, Trust/Reliance Measurement (TM) scale.
- Limitations and future work:
- Limited to GPT-4; other LLMs were not evaluated.
- Contextual limitations of transcripts; future studies could use audio-visual inputs.
- Participants were exclusively from the US; results may not generalize globally.
- Future work could explore hybrid or in-person studies, larger datasets, and integration with conversational agents on platforms like YouTube.
Summary
This study evaluated GPT-4's ability to assist parents in assessing the appropriateness of YouTube videos for children under 7. While GPT-4's classification accuracy was subpar (72%), its explanations were highly satisfying to parents (92.1% satisfaction). Refined prompts improved alignment with parental preferences, enabling decisions comparable to watching full videos. The findings suggest that LLMs should not make final moderation decisions but can serve as effective tools for aiding parental decision-making. Future work could explore broader applications, improved LLM models, and real-world integration to support child safety online.
Research Questions / Practical Problems
Question signals indexed for this paper.
No related papers with ≥60% similarity
Based on Jaccard similarity of research subtopics & professions (≥60%)